Compare commits

..
47 Commits
Author SHA1 Message Date
petrbalvin cf6bc6987e fix(ci): pass the upload file to curl, not its interpolation
Test / test (push) Successful in 2m26s
Release / gates (push) Successful in 2m25s
Release / build (amd64, linux) (push) Successful in 1m15s
Release / build (arm64, linux) (push) Successful in 1m16s
Release / build (loong64, linux) (push) Successful in 1m16s
Release / build (riscv64, linux) (push) Successful in 1m16s
Release / release (push) Successful in 34s
2026-09-22 01:31:49 +02:00
petrbalvin ff7b1452b1 docs: name 0.35.0 as the supported release
Test / test (push) Successful in 2m33s
Release / gates (push) Successful in 2m29s
Release / build (amd64, linux) (push) Successful in 1m18s
Release / build (arm64, linux) (push) Successful in 1m20s
Release / build (loong64, linux) (push) Successful in 1m17s
Release / build (riscv64, linux) (push) Successful in 1m26s
Release / release (push) Failing after 35s
2026-09-22 00:52:56 +02:00
petrbalvin 517c1cea25 chore: prepare release v0.35.0
Test / test (push) Successful in 2m33s
Release / gates (push) Failing after 46s
Release / build (amd64, linux) (push) Skipped
Release / build (arm64, linux) (push) Skipped
Release / build (loong64, linux) (push) Skipped
Release / build (riscv64, linux) (push) Skipped
Release / release (push) Skipped
2026-09-22 00:44:10 +02:00
petrbalvin a3e3010e0f fix(cmd): resolve the runtime header test GOROOT from the go command
Test / test (push) Successful in 2m39s
2026-09-21 22:46:07 +02:00
petrbalvin 057c4eb545 docs: complete the release delta in the changelog and readme 2026-09-21 22:45:56 +02:00
petrbalvin f720381d43 feat(asm): the segment-absolute and crash-store forms GOROOT writes
Test / test (push) Failing after 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-21 22:19:53 +02:00
petrbalvin 2c9042d62c feat(asm): PCALIGN alignment on amd64
Assisted-by: GLM 5.3 Flash
2026-09-21 22:00:30 +02:00
petrbalvin 82ef289d3a feat(asm): the immediate multiply and arm64 indirect branches GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:50:11 +02:00
petrbalvin 7246b0e002 feat(asm): the TLS access pair in the toolchain's one-instruction form
Assisted-by: GLM 5.3 Flash
2026-09-21 21:35:15 +02:00
petrbalvin 8cfd40aac8 feat(asm): the operand forms and defines GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:17:34 +02:00
petrbalvin 5382c9a8e4 feat(audit): list every corpus failure per architecture 2026-09-21 21:17:34 +02:00
petrbalvin 53de91b2df docs(asm): describe the four target architectures
Test / test (push) Failing after 2m23s
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin 8a36af7c7d docs(asm): generate the instruction appendices
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin e9789ce3f4 chore(arch): regenerate the instruction tables 2026-09-21 20:15:55 +02:00
petrbalvin 837231c068 docs(asm): open the assembly language reference
Assisted-by: GLM 5.3 Flash
2026-09-21 19:49:04 +02:00
petrbalvin 95025be1bc docs(changelog): describe the encoder entries by content
Test / test (push) Failing after 2m33s
2026-09-21 19:20:01 +02:00
petrbalvin 03a964bb2d docs(goobj): document the GOOBJ object file format 2026-09-21 19:19:53 +02:00
petrbalvin 123a16e346 docs(readme): state the documentation goal 2026-09-21 18:35:27 +02:00
petrbalvin 9701812bee docs: changelog for the completeness waves
Test / test (push) Failing after 3m6s
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin 29ac03468e feat(amd64): floating-point immediates through a synthesised pool
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin bfb7701db1 feat(amd64): emit the quad-register EVEX families
Assisted-by: GLM 5.3 Flash
2026-09-21 02:02:19 +02:00
petrbalvin e8b6ff5d7c fix(parser): fold a signed parenthesised displacement expression
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-21 00:45:33 +02:00
petrbalvin 1456907000 feat(riscv64,loong64): operand tail, float DATA and honest port classification
Assisted-by: GLM 5.3 Flash
2026-09-21 00:44:47 +02:00
petrbalvin ec1c521187 feat(cmd): GOOS-aware headers, audit battery shapes and semicolon spacing
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:46 +02:00
petrbalvin a7744c24bd fix(parser): substitute macro parameters behind element selectors
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 522e6f2ae8 feat(parser): bracket register ranges, index-only VSIB and bare trailing immediates
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 81d4bd81e4 test(verify): register the wave kernels
Test / test (push) Failing after 2m20s
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:31 +02:00
petrbalvin 687678a2ea feat(elf): emit data relocations on arm64, riscv64 and loong64
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin b0f9071bf5 feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 81e2673923 feat(amd64): encode the AVX-512 and BMI corpus families
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 75e9fd771b feat(parser): split plain statements on semicolons in the raw parse
Test / test (push) Failing after 2m30s
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:35 +02:00
petrbalvin 863926abd6 test(verify): register the loong64 vector kernels
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 241e7256f6 fix(arm64): reject bare BTI with a diagnostic and accept the full family
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 6556b85abf feat(asm): symbol-valued DATA, division slash in symbols and plain semicolons
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 289cabe993 feat(loong64): encode the full LSX and LASX table
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin d6cf7cfa44 fix(format): keep statement separators and canonical macro bodies
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin 4cc2f0eba5 feat(cmd): generate go_asm.h for package-context assembly
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin 97dfaa7526 docs: changelog for macro expansion and the corrected corpus audit
Test / test (push) Successful in 2m14s
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 66aa4dbc8b test(verify): register the campaign kernels in the ground-truth suites
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin dce5d31462 feat(amd64): LOCK and REP prefixes, literal data pseudo-ops and ADJSP
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9dc3987e02 feat(riscv64,loong64): PCALIGN, branch relaxation and operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9b238a525a feat(arm64): wide immediates, SIMD compare and system operand forms
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin ad82aac663 feat(parser): macro expansion, conditionals and include splicing with -I
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 0629f5e2df feat(arm64): assemble PCALIGN padding and BYTE literal bytes
Test / test (push) Successful in 2m16s
Assisted-by: GLM 5.3 Flash
2026-09-20 11:49:05 +02:00
petrbalvin ecb203dcf5 fix(lexer): treat trailing CR as line end so comment text is idempotent
Test / test (push) Successful in 2m13s
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
petrbalvin 6c672567f3 feat(amd64): assemble the double-shift and static-SB operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
petrbalvin cc6e416c59 fix(lint): exempt shift counts, SETcc and ABIInternal from false positives
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
137 changed files with 24962 additions and 604 deletions
+4 -1
View File
@@ -342,7 +342,10 @@ jobs:
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
q{-H}, q{Content-Type: application/octet-stream},
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
# The @ must not sit inside a qq{} string: there it starts an
# array interpolation and the upload body collapses to empty,
# which Gitea stores as a 201-created zero-byte attachment.
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
my $code = <$curl>;
+126 -9
View File
@@ -9,7 +9,43 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
- **The GOROOT instruction wave, part 1.** The encoder now covers the
-
## [0.35.0] - 2026-09-22
### Added
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
itself when an assembly file includes it: the Go files beside the source
are type-checked for the target architecture and the constants and field
offsets become assembler defines, so package-context files assemble with
no compiler and no `go build` in the loop. `-GOOS` selects the
type-checking GOOS for GOOS-specific files, and the corpus audit derives
the GOOS from the file name.
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
every architecture (amd64 carried them already), so standalone ELF
objects link on all four targets.
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
prints every failing file with its failure reason, per architecture,
instead of one representative file per reason.
- **DATA with symbol values and relaxed symbol spellings.** DATA
initialisers accept `$symbol(SB)` values, laid down as an absolute
relocation at the data field (GOOBJ on all four architectures and ELF
on all four as of this release), and U+2215 is accepted inside symbol
package paths.
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
`gasm audit-instructions` now preprocess assembly the way the
toolchain does: object and parameterised `#define` macros expand at
the point of use, `#undef` and the `#ifdef`/`#ifndef`/`#else`/
`#endif` family select branches, `#include` splices headers resolved
through the source directory and the new repeatable `-I` flag, `;`
separates statements, and constant expressions left in operands
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
language server keep reading the raw file.
- **Encoder coverage: the instruction families GOROOT's real code
uses.** The encoder now covers the
instruction families GOROOT's real code uses that gasm lacked,
byte-verified against `go tool asm`: on amd64 the carry ALU, the
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
@@ -25,14 +61,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
riscv64 `VSETVLI` with an immediate length now canonicalises to
`vsetivli` as the toolchain does.
- **The corpus audit measures honestly.** Files named for Go ports gasm
does not target (arm, 386, s390x, ...) are no longer attempted for the
four supported architectures (no supported build compiles them), and
the headline rate is reported over attemptable files: 136 of 433 on
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
127 that the previous release measured. The probe battery that
decides encodability gained the operand shapes the new families use.
-
- **Encoder coverage: quad-register AVX-512 and floating-point
immediates.** The encoder gains the
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
floating-point immediates on the SSE scalar moves and arithmetic
(the constant lands in a synthesised read-only pool, a positive zero
collapses to XORPS exactly as the toolchain does), accept-and-ignore
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
operands for the legacy SSE moves, and the pooled 64-bit immediate
materialisation on riscv64. The parser carries bracketed register
ranges, index-only VSIB memory operands and bare trailing immediates;
macro substitution reaches parameters used with element suffixes
(`A.S4`), and `;` separates statements in plain files.
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
roles the ABI fixes, addressing, operand order with every special form,
constants and materialisation, alignment, fences and the relocations
each target emits. An instruction inventory appendix per architecture
is generated from the toolchain's own tables by `just gen`, and the
regenerated tables recognise 147 more mnemonics than the previous
release carried (arm64 107, riscv64 31, loong64 9).
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
opens the complete language reference with its common core: the lexicon,
statement structure and constant expressions, the operand grammar with
the pseudo-registers and symbol naming, the directives and the function
flag vocabulary, preprocessing with `#define` and `#include`, and the
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
runtime contract). Every claim is verified against `go tool asm` of
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
generated instruction appendices follow.
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
documents the Go object file format in full: both containers, the 96
byte header and all 19 blocks, every structure with its byte
offsets, symbol kinds and flag bits, all 106 relocation types with
the weak variants, aux symbols, the FuncInfo payload, the pc-value
table encoding, the content hashes and the builtin table, all
verified byte for byte against objects produced by Go 1.27.1's own
tools.
### Changed
- **The corpus audit measures like a build.** Files named for a Go port
gasm does not target (arm, 386, s390x, ...) are never attempted, because
no supported build compiles them; the GOOS comes from the file name; and
each target's go_asm.h is generated on the fly. The headline is reported
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
for every target architecture and 295 of 303 on real code (97.4 %),
against 108 of 627 over all files (17.2 %) that the previous release
measured.
### Fixed
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
own instruction counting and fold jump-to-jump chains exactly as its
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
assemble to the toolchain's RIP-relative LEA with an R_PCREL
relocation; negated constant expressions in operands (`ADJSP
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
one-instruction form (the bare `MOVQ TLS, r` load nops out and
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
the zero-immediate store (`MOVD $0, mem` through the zero register,
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
aligns on amd64, padding with the toolchain's greedy
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
AX` and the store direction) and the absolute crash-store
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
GOROOT corpus measure moves to 291 of 353 files assembling for every
target architecture (82.4 %), 97.4 % of the real-code corpus, from
70.8 % and 82.2 %.
- **Tool corrections across the pipeline.** The formatter keeps square
brackets in SIMD operands, statement separators and canonical macro
bodies; the linter drops false positives on shift counts, SETcc
spellings and ABIInternal references; the lexer treats a trailing
carriage return as a line end so comment text stays idempotent; and
arm64 rejects bare BTI with a diagnostic while accepting the full
family.
## [0.34.0] - 2026-09-20
@@ -129,6 +241,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- **The corpus audit attempts fewer files that no build would compile.**
Files named for Go ports gasm does not target (arm, 386, s390x, ...)
are reported as other-port and never attempted, the headline rate is
computed over attemptable files, and the audit searches the
toolchain's shipped headers (funcdata.h and friends) automatically.
- **riscv64 JALR silently jumped to the wrong register.** The trampoline
form `JALR X0, 0(X5)` read the memory operand's base as the destination,
encoding a jump to X0 with no diagnostic; the destination is the first
+44 -6
View File
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
GOOBJ format, which needs the installed toolchain and which `go build`
consumes in place of the toolchain's output. Framed functions get the
stack-split guard and the morestack block, byte-identical to the
toolchain's, so split functions link too.
toolchain's, so split functions link too. The assembler preprocesses
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
offsets after assembling, or disassembles raw bytes from a file or stdin.
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
@@ -110,9 +113,9 @@ Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
"Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
@@ -124,8 +127,8 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
(31.4 %) assembling for every target architecture today (files named for
`gasm audit-instructions --corpus` reports 291 of 353 attemptable files
(82.4 %) assembling for every target architecture today (files named for
other Go ports are counted but never attempted), with the top failure
reasons per architecture; the number moves with every release.
@@ -157,6 +160,39 @@ been compiled and read, never executed. Its architecture-neutral units
run under `go test ./...`, which the race workflow and a manual run
perform; the default `just test` gate does not sweep `./debug/...`.
## The documentation goal
The toolkit is the primary goal. The secondary one is documentation: a
specification of the Plan 9 assembly language and of the GOOBJ object
format that is 100 % complete, detailed enough to implement against,
and written to a professional standard. These are the two subjects this
project works with every day, and they are the two for which no usable
documentation exists.
Go documents the language on a single page, "A Quick Guide to Go's
Assembler", which carries no section for loong64, one of the four
architectures gasm supports, and covers a fraction of what each
assembler accepts. What exists beyond it lives as comments inside the
toolchain's internal source: per-architecture reference manuals for
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
maintainers rather than for an outside reader, and none at all for
amd64. GOOBJ fares worst of all. The format that `go build` consumes
has no specification anywhere: it is described by a comment in an
internal package, it is not a stable interface, and it can change with
any toolchain release.
The gap is therefore filled the only way it can be filled: by reverse
engineering the toolchain itself, the same work the encoders already
perform. Most of the documentation can come from nowhere else, and it
is written as that knowledge is produced during development. It is
verified the way the code is verified: an encoding documented here is
one that differential tests against `go tool asm` confirm
byte-for-byte, and a format field documented here is one the linker
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
specifies the object file format completely, and
[docs/asm/README.md](docs/asm/README.md) opens the language reference
with its common core. The per-architecture pages follow.
## Direction
The plan, in the order it is being worked:
@@ -282,6 +318,8 @@ recipe.
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [CHANGELOG.md](CHANGELOG.md): release history
+1 -1
View File
@@ -7,7 +7,7 @@ releases do not receive them.
| Version | Supported |
|---|---|
| 0.34.0 | yes |
| 0.35.0 | yes |
| older releases | no |
## Reporting a vulnerability
+101 -3
View File
@@ -8,6 +8,10 @@
// names so gasm-devkit supports every instruction the real assembler does,
// with no hand-maintained (and therefore inevitably incomplete) lists.
//
// The same data feeds the generated instruction appendices of the assembly
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
// cannot drift from the tables it documents.
//
// Usage (via the justfile):
//
// just gen
@@ -26,6 +30,9 @@ import (
"path/filepath"
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
)
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
@@ -39,11 +46,30 @@ var archDirs = []struct {
{"loong64", "loong64"},
}
// docPages maps an architecture to its generated appendix in the language
// reference. The amd64 page carries a per-mnemonic encodability column,
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
// other targets have no single cheap predicate, so their pages carry the
// inventory and point at the live measurement instead.
var docPages = []struct {
arch arch.Arch
title string
file string
anames string
encodable bool
}{
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
}
func main() {
goroot := strings.TrimSpace(runGoEnvGOROOT())
if goroot == "" {
fatal("could not determine GOROOT")
}
version := strings.TrimSpace(runGoEnv("GOVERSION"))
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
@@ -57,16 +83,24 @@ func main() {
}
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
names := map[string][]string{}
for _, a := range archDirs {
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
names, err := extractInstrs(path)
names[a.arch], err = extractInstrs(path)
if err != nil {
fatal("extract %s: %v", a.arch, err)
}
if err := writeGen(a.arch, a.sub, names); err != nil {
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
fatal("write %s: %v", a.arch, err)
}
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
}
for _, p := range docPages {
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
fatal("write %s: %v", p.file, err)
}
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
}
}
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
}
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
// the language reference for one architecture: every mnemonic the toolchain
// accepts, with the curated summary where the architecture table carries one
// and, on amd64, a per-mnemonic encodability column.
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
table := arch.ForArch(a)
instrs := table.Instructions()
var b strings.Builder
b.WriteString("# " + title + ": instruction inventory\n\n")
b.WriteString("Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table\n")
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
encodableCount := 0
if encodable {
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
b.WriteString("`gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
b.WriteString("|---|---|---|\n")
for _, in := range instrs {
ok := asm.Encodable(in.Name)
if ok {
encodableCount++
}
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
} else {
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
b.WriteString("encodability is decided per operand shape, and the live measured\n")
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | Notes |\n")
b.WriteString("|---|---|\n")
for _, in := range instrs {
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
}
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
}
// yesNo renders a boolean as the word the appendix tables use.
func yesNo(v bool) string {
if v {
return "yes"
}
return "no"
}
func runGoEnvGOROOT() string {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
return string(out)
}
// runGoEnv runs `go env` for a single variable.
func runGoEnv(name string) string {
out, err := exec.Command("go", "env", name).Output()
if err != nil {
return ""
}
return string(out)
}
func fatal(format string, args ...any) {
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
os.Exit(1)
+1
View File
@@ -33,6 +33,7 @@ func arm64Registers() []Register {
for i := 0; i <= 30; i++ {
add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register")
}
add("R18_PLATFORM", GPR, "R18 under its toolchain-reserved Windows name (an alias of R18)")
add("ZR", Special, "zero register (reads as 0)")
add("SP", Special, "stack pointer")
add("LR", Special, "link register (alias of R30)")
+107
View File
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
"REVW",
"ROR",
"RORW",
"RPRFM",
"SB",
"SBC",
"SBCS",
"SBCSW",
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
"UXTH",
"UXTHW",
"UXTW",
"VABS",
"VADD",
"VADDP",
"VADDV",
"VAND",
"VBCAX",
"VBIC",
"VBIF",
"VBIT",
"VBSL",
"VCLS",
"VCLZ",
"VCMEQ",
"VCMGE",
"VCMGT",
"VCMHI",
"VCMHS",
"VCMLE",
"VCMLT",
"VCMTST",
"VCNT",
"VDUP",
"VEOR",
"VEOR3",
"VEXT",
"VFABS",
"VFADD",
"VFADDP",
"VFCMEQ",
"VFCMGE",
"VFCMGT",
"VFCMLE",
"VFCMLT",
"VFCVTL",
"VFCVTL2",
"VFCVTN",
"VFCVTN2",
"VFCVTZS",
"VFCVTZU",
"VFDIV",
"VFMAX",
"VFMAXNM",
"VFMAXNMP",
"VFMAXNMV",
"VFMAXP",
"VFMAXV",
"VFMIN",
"VFMINNM",
"VFMINNMP",
"VFMINNMV",
"VFMINP",
"VFMINV",
"VFMLA",
"VFMLS",
"VFMUL",
"VFNEG",
"VFRINTM",
"VFRINTN",
"VFRINTP",
"VFRINTZ",
"VFSQRT",
"VFSUB",
"VLD1",
"VLD1R",
"VLD2",
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
"VLD3R",
"VLD4",
"VLD4R",
"VMLA",
"VMLS",
"VMOV",
"VMOVD",
"VMOVI",
"VMOVQ",
"VMOVS",
"VMUL",
"VNEG",
"VNOT",
"VORN",
"VORR",
"VPMULL",
"VPMULL2",
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
"VREV16",
"VREV32",
"VREV64",
"VSCVTF",
"VSHADD",
"VSHL",
"VSHRN",
"VSHRN2",
"VSLI",
"VSMAX",
"VSMAXP",
"VSMAXV",
"VSMIN",
"VSMINP",
"VSMINV",
"VSMLAL",
"VSMLAL2",
"VSMLSL",
"VSMLSL2",
"VSMULL",
"VSMULL2",
"VSQABS",
"VSQADD",
"VSQNEG",
"VSQSHL",
"VSQSUB",
"VSQXTN",
"VSQXTN2",
"VSQXTUN",
"VSQXTUN2",
"VSRHADD",
"VSRI",
"VSRSHR",
"VSSHL",
"VSSHLL",
"VSSHLL2",
"VSSHR",
"VST1",
"VST2",
"VST3",
"VST4",
"VSUB",
"VSXTL",
"VSXTL2",
"VTBL",
"VTBX",
"VTRN1",
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
"VUADDLV",
"VUADDW",
"VUADDW2",
"VUCVTF",
"VUHADD",
"VUMAX",
"VUMAXP",
"VUMAXV",
"VUMIN",
"VUMINP",
"VUMINV",
"VUMLAL",
"VUMLAL2",
"VUMLSL",
"VUMLSL2",
"VUMULL",
"VUMULL2",
"VUQADD",
"VUQSHL",
"VUQSUB",
"VUQXTN",
"VUQXTN2",
"VURHADD",
"VUSHL",
"VUSHLL",
"VUSHLL2",
"VUSHR",
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
"VUZP1",
"VUZP2",
"VXAR",
"VXTN",
"VXTN2",
"VZIP1",
"VZIP2",
"WFE",
+9
View File
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
"FNMADDF",
"FNMSUBD",
"FNMSUBF",
"FRINTD",
"FRINTF",
"FSCALEBD",
"FSCALEBF",
"FSEL",
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
"FTINTWF",
"JIRL",
"LL",
"LLACQV",
"LLACQW",
"LLV",
"LLW",
"LU12IW",
"LU32ID",
"LU52ID",
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
"ROTR",
"ROTRV",
"SC",
"SCQ",
"SCRELV",
"SCRELW",
"SCV",
"SCW",
"SGT",
"SGTU",
"SLL",
+31
View File
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
"CLD",
"CLDSP",
"CLI",
"CLMUL",
"CLMULH",
"CLMULR",
"CLUI",
"CLW",
"CLWSP",
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
"CSDSP",
"CSLLI",
"CSRAI",
"CSRC",
"CSRCI",
"CSRLI",
"CSRR",
"CSRRC",
"CSRRCI",
"CSRRS",
"CSRRSI",
"CSRRW",
"CSRRWI",
"CSRS",
"CSRSI",
"CSRW",
"CSRWI",
"CSUB",
"CSUBW",
"CSW",
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
"ORCB",
"ORI",
"ORN",
"PAUSE",
"RDCYCLE",
"RDINSTRET",
"RDTIME",
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
"VADDVI",
"VADDVV",
"VADDVX",
"VANDNVV",
"VANDNVX",
"VANDVI",
"VANDVV",
"VANDVX",
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
"VASUBUVX",
"VASUBVV",
"VASUBVX",
"VBREV8V",
"VBREVV",
"VCLMULHVV",
"VCLMULHVX",
"VCLMULVV",
"VCLMULVX",
"VCLZV",
"VCOMPRESSVM",
"VCPOPM",
"VCPOPV",
"VCTZV",
"VDIVUVV",
"VDIVUVX",
"VDIVVV",
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
"VREMUVX",
"VREMVV",
"VREMVX",
"VREV8V",
"VRGATHEREI16VV",
"VRGATHERVI",
"VRGATHERVV",
"VRGATHERVX",
"VROLVV",
"VROLVX",
"VRORVI",
"VRORVV",
"VRORVX",
"VRSUBVI",
"VRSUBVX",
"VS1RV",
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
"VWMULVX",
"VWREDSUMUVS",
"VWREDSUMVS",
"VWSLLVI",
"VWSLLVV",
"VWSLLVX",
"VWSUBUVV",
"VWSUBUVX",
"VWSUBUWV",
+101
View File
@@ -245,3 +245,104 @@ func main() {
t.Error("binary does not contain expected symbol")
}
}
// TestGOObjectAARCH64DataSymbolLink does for symbol-valued DATA fields what
// the rt0 files do ("DATA _rt0…lib+0(SB)/8, $_rt0…lib(SB)"): the gasm object
// carries an R_ADDR against the file's own TEXT symbol, the toolchain links
// it, and the binary is checked for the symbol (no arm64 host to run it).
func TestGOObjectAARCH64DataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVD entry+0(SB), R4
MOVD R4, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
func keepme()
func entryptr() uintptr
func main() {
if entryptr() == 0 {
panic("the entry word is empty")
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module a64dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
build.Env = append(os.Environ(), "GOARCH=arm64")
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_arm64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
src, err := os.ReadFile(filepath.Join(dir, "main_arm64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_arm64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
gasmObj, err := img.GOObjectAARCH64("a64dlink", "main_arm64.s")
if err != nil {
t.Fatalf("GOObjectAARCH64: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
linkCmd.Env = append(os.Environ(), "GOARCH=arm64")
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
binData, err := os.ReadFile(filepath.Join(dir, "prog"))
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(binData), "keepme") {
t.Error("binary does not contain the keepme symbol")
}
}
+1698 -129
View File
File diff suppressed because it is too large Load Diff
+369 -55
View File
@@ -29,6 +29,7 @@ package asm
import (
"maps"
"math/bits"
"strconv"
"strings"
@@ -78,6 +79,11 @@ func arm64RegNum(name string) int {
return 17
case "R18":
return 18
case "R18_PLATFORM":
// The toolchain's Windows spelling: R18 is renamed R18_PLATFORM in
// cmd/asm/internal/arch so assembly cannot use it by accident, and
// sys_windows_arm64.s references it only through this name.
return 18
case "R19":
return 19
case "R20":
@@ -159,6 +165,67 @@ func a64MoveWide(sf, opc, hw, imm16, rd uint32) uint32 {
return sf<<31 | opc<<29 | 0x25<<23 | hw<<21 | imm16<<5 | rd
}
// ---- logical immediate ----
// a64LogicalImm encodes v as the AArch64 logical (bitmask) immediate for the
// given lane width (32 or 64): it returns the N, immr and imms fields of the
// imm13 encoding. The algorithm mirrors cmd/internal/obj/arm64's
// encodeLogicalImmArrEncoding: replicate the value, shrink it to the smallest
// repeating element, find the run of ones and its rotation. ok is false when
// v is not expressible (all zeros, all ones, or not a single cyclic run).
func a64LogicalImm(v int64, width int) (n, immr, imms uint32, ok bool) {
u := uint64(v)
if width == 32 {
u &= 0xFFFFFFFF
}
size := uint64(width)
mask := ^uint64(0)
if size < 64 {
mask = uint64(1)<<size - 1
}
u &= mask
// All zeros and all ones are MOV territory, not bitmask immediates.
if u == 0 || u == mask {
return 0, 0, 0, false
}
// Shrink to the smallest repeating element.
for size > 2 {
half := size / 2
hm := uint64(1)<<half - 1
if u&hm == u>>half&hm {
size = half
u &= hm
} else {
break
}
}
ones := bits.OnesCount64(u)
// Find the right-rotation that lays the ones out contiguously at the
// bottom of the element; the hardware applies the inverse rotation.
em := uint64(1)<<size - 1
expected := uint64(1)<<ones - 1
rot := -1
for r := 0; r < int(size); r++ {
rotated := u>>r | u<<(int(size)-r)
if size < 64 {
rotated &= em
}
if rotated == expected {
rot = r
break
}
}
if rot < 0 {
return 0, 0, 0, false
}
if size == 64 {
n = 1
}
immr = uint32((int(size) - rot) % int(size))
imms = ^uint32(uint32(size*2-1))&0x3F | uint32(ones-1)
return n, immr, imms, true
}
// ---- load/store (unsigned immediate, scaled) ----
// a64LSU encodes a load/store register (unsigned immediate, scaled):
@@ -246,6 +313,8 @@ const (
a64CondLT = 0xb
a64CondGT = 0xc
a64CondLE = 0xd
a64CondAL = 0xe
a64CondNV = 0xf
)
// arm64CondMap maps Go assembler condition mnemonics to AArch64 condition codes.
@@ -266,6 +335,8 @@ var arm64CondMap = map[string]uint32{
"LT": a64CondLT,
"GT": a64CondGT,
"LE": a64CondLE,
"AL": a64CondAL,
"NV": a64CondNV,
}
// ---- instruction format tags ----
@@ -273,46 +344,47 @@ var arm64CondMap = map[string]uint32{
type a64Format uint8
const (
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
a64FMovWide // move wide: MOVZ, MOVN, MOVK
a64FBranch // unconditional branch (B/BL)
a64FBranchCond // conditional branch (B.cond)
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
a64FADR // ADR/ADRP
a64FEXTR // EXTR
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FDP1 // data-processing (1 source): RBIT, REV, CLZ, CLS
a64FBitfield2 // bitfield extract: UBFX, SBFX and the W forms
a64FCondCmp // conditional compare: CCMP, CCMN
a64FBranch19 // compare-and-branch: CBZ, CBNZ and the W forms
a64FTestBranch // test-and-branch: TBZ, TBNZ and the W forms
a64FPair // load/store pair: LDP, STP, LDPW, STPW, FLDPD, FSTPD
a64FAcqRel // acquire/release: LDAR family, STLR family
a64FSys // system: BRK, SVC, DMB, DSB, ISB, DC, MRS, MSR, PRFM
a64FCrypto2 // crypto 2-register: AESD, AESE, AESIMC, AESMC, SHA1H, ...
a64FCrypto3 // crypto 3-register: SHA1C, SHA256H, SHA512SU1, ...
a64FSIMDV // SIMD 3-register with arrangement: VADD, VAND, VCMEQ, VZIP1, ...
a64FSIMDVZero // SIMD compare against zero: VCMEQ $0, Vn, Vd
a64FSIMDV2 // SIMD 2-register with arrangement: VREV32, VREV64, VUADDLV, VMOV
a64FSIMDV4 // SIMD 4-register / imm 3-register: VEOR3, VBCAX, VXAR, VEXT
a64FVTBL // SIMD table lookup: VTBL
a64FDUP // SIMD element moves: VDUP, VMOV with element indices
a64FVLDST // SIMD structure loads/stores: VLD1, VST1, VLD1R, VLD4R
a64FShiftImm // SIMD shift by immediate: VSHL, VUSHR, VSRI
a64FMoviLit // VMOVS/VMOVD/VMOVQ with a large constant (literal pool)
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
a64FMovWide // move wide: MOVZ, MOVN, MOVK
a64FBranch // unconditional branch (B/BL)
a64FBranchCond // conditional branch (B.cond)
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
a64FADR // ADR/ADRP
a64FEXTR // EXTR
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
a64FBitfieldAlias // bitfield alias: BFI/BFXIL/SBFIZ/UBFIZ, ($lsb, Rn, $width, Rd)
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FDP1 // data-processing (1 source): RBIT, REV, CLZ, CLS
a64FBitfield2 // bitfield extract: UBFX, SBFX and the W forms
a64FCondCmp // conditional compare: CCMP, CCMN
a64FBranch19 // compare-and-branch: CBZ, CBNZ and the W forms
a64FTestBranch // test-and-branch: TBZ, TBNZ and the W forms
a64FPair // load/store pair: LDP, STP, LDPW, STPW, FLDPD, FSTPD
a64FAcqRel // acquire/release: LDAR family, STLR family
a64FSys // system: BRK, SVC, DMB, DSB, ISB, DC, MRS, MSR, PRFM
a64FCrypto2 // crypto 2-register: AESD, AESE, AESIMC, AESMC, SHA1H, ...
a64FCrypto3 // crypto 3-register: SHA1C, SHA256H, SHA512SU1, ...
a64FSIMDV // SIMD 3-register with arrangement: VADD, VAND, VCMEQ, VZIP1, ...
a64FSIMDVZero // SIMD compare against zero: VCMEQ $0, Vn, Vd
a64FSIMDV2 // SIMD 2-register with arrangement: VREV32, VREV64, VUADDLV, VMOV
a64FSIMDV4 // SIMD 4-register / imm 3-register: VEOR3, VBCAX, VXAR, VEXT
a64FVTBL // SIMD table lookup: VTBL
a64FDUP // SIMD element moves: VDUP, VMOV with element indices
a64FVLDST // SIMD structure loads/stores: VLD1, VST1, VLD1R, VLD4R
a64FShiftImm // SIMD shift by immediate: VSHL, VUSHR, VSRI
a64FMoviLit // VMOVS/VMOVD/VMOVQ with a large constant (literal pool)
)
// a64Enc is one instruction's encoding: its bit layout (format) and the
@@ -425,6 +497,16 @@ func init() {
a64InstrTable["MADDW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24}
a64InstrTable["MSUB"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<15}
a64InstrTable["MSUBW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24 | 1<<15}
// The widening multiplies: a 64-bit result riding the same layout, the
// three-operand forms reading the accumulate register as ZR.
a64InstrTable["SMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21}
a64InstrTable["UMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23}
a64InstrTable["SMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15}
a64InstrTable["UMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15}
a64InstrTable["SMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 31<<10}
a64InstrTable["UMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 31<<10}
a64InstrTable["SMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15 | 31<<10}
a64InstrTable["UMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15 | 31<<10}
// ---- move wide ----
// MOVZ/MOVN/MOVK
@@ -474,6 +556,15 @@ func init() {
// ---- bitfield ----
a64InstrTable["BFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 1<<29 | 0x26<<23 | 0<<22}
// The four-operand bitfield aliases: ($lsb, Rn, $width, Rd).
a64InstrTable["BFI"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFIW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
a64InstrTable["BFXIL"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFXILW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
a64InstrTable["SBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x93400000}
a64InstrTable["SBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x13000000}
a64InstrTable["UBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x53000000}
a64InstrTable["UBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x33000000}
a64InstrTable["SBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 0<<29 | 0x26<<23 | 1<<22}
a64InstrTable["SBFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 0<<29 | 0x26<<23 | 0<<22}
a64InstrTable["UBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 2<<29 | 0x26<<23 | 1<<22}
@@ -654,6 +745,13 @@ func init() {
"RBIT": 0xdac00000, "REV16": 0xdac00400, "REV32": 0xdac00800,
"REV": 0xdac00c00, "CLZ": 0xdac01000, "CLS": 0xdac01400,
"RBITW": 0x5ac00000, "REVW": 0x5ac00800, "CLZW": 0x5ac01000, "CLSW": 0x5ac01400,
// Extend and byte-reverse: the UBFM/SBFM aliases with imms fixing
// the source width.
"SXTB": 0x93401c00, "SXTBW": 0x13001c00, "SXTH": 0x93403c00,
"SXTHW": 0x13003c00, "SXTW": 0x93407c00,
"UXTB": 0x53001c00, "UXTBW": 0x53001c00, "UXTH": 0x53403c00,
"UXTHW": 0x53003c00, "UXTW": 0x53407c00,
"REV16W": 0x5ac00400,
}
for m, op := range dp1 {
a64InstrTable[m] = a64Enc{format: a64FDP1, op: op}
@@ -672,7 +770,7 @@ func init() {
a64InstrTable["CCMNW"] = a64Enc{format: a64FCondCmp, op: 0x3a400000}
// ---- system operations ----
for _, m := range []string{"BRK", "SVC", "DMB", "DSB", "ISB", "DC", "MRS", "MSR", "PRFM"} {
for _, m := range []string{"BRK", "SVC", "DMB", "DSB", "ISB", "CLREX", "HINT", "BTI", "HLT", "SMC", "HVC", "DCPS1", "DCPS2", "DCPS3", "DRPS", "ERET", "AUTIASP", "AUTIBSP", "AUTIA1716", "AUTIB1716", "SEVL", "SEV", "WFE", "WFI", "YIELD", "DC", "MRS", "MSR", "PRFM"} {
a64InstrTable[m] = a64Enc{format: a64FSys}
}
@@ -723,6 +821,73 @@ func init() {
for m, op := range lse {
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
}
// The remaining width and ordering spellings of the same shapes, and the
// CAS compare-and-swap family, word-verified against go tool asm.
lseMore := map[string]uint32{
"LDADDAB": 0x38a00000,
"LDADDAH": 0x78a00000,
"LDADDALB": 0x38e00000,
"LDADDALH": 0x78e00000,
"LDADDLB": 0x38600000,
"LDADDLD": 0xf8600000,
"LDADDLH": 0x78600000,
"LDADDLW": 0xb8600000,
"LDCLRAB": 0x38a01000,
"LDCLRAH": 0x78a01000,
"LDCLRALH": 0x78e01000,
"LDCLRB": 0x38201000,
"LDCLRD": 0xf8201000,
"LDCLRH": 0x78201000,
"LDCLRLB": 0x38601000,
"LDCLRLD": 0xf8601000,
"LDCLRLH": 0x78601000,
"LDCLRLW": 0xb8601000,
"LDCLRW": 0xb8201000,
"LDEORAB": 0x38a02000,
"LDEORAD": 0xf8a02000,
"LDEORAH": 0x78a02000,
"LDEORALB": 0x38e02000,
"LDEORALH": 0x78e02000,
"LDEORAW": 0xb8a02000,
"LDEORB": 0x38202000,
"LDEORD": 0xf8202000,
"LDEORH": 0x78202000,
"LDEORLB": 0x38602000,
"LDEORLD": 0xf8602000,
"LDEORLH": 0x78602000,
"LDEORLW": 0xb8602000,
"LDEORW": 0xb8202000,
"LDORAB": 0x38a03000,
"LDORAD": 0xf8a03000,
"LDORAH": 0x78a03000,
"LDORALH": 0x78e03000,
"LDORAW": 0xb8a03000,
"LDORB": 0x38203000,
"LDORD": 0xf8203000,
"LDORH": 0x78203000,
"LDORLB": 0x38603000,
"LDORLD": 0xf8603000,
"LDORLH": 0x78603000,
"LDORLW": 0xb8603000,
"LDORW": 0xb8203000,
"SWPAB": 0x38a08000,
"SWPAD": 0xf8a08000,
"SWPAH": 0x78a08000,
"SWPALH": 0x78e08000,
"SWPAW": 0xb8a08000,
"SWPB": 0x38208000,
"SWPH": 0x78208000,
"SWPLB": 0x38608000,
"SWPLD": 0xf8608000,
"SWPLH": 0x78608000,
"SWPLW": 0xb8608000,
"CASAD": 0xc8e07c00,
"CASALB": 0x08e0fc00,
"CASLW": 0x88a0fc00,
}
for m, op := range lseMore {
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
}
// ---- carry-setting/carry-using arithmetic and widening multiply ----
// MUL and SMULH/UMULH are the MADD/MSUB layout with the accumulate
@@ -732,7 +897,12 @@ func init() {
"ADCS": 0xba000000, "ADCSW": 0x3a000000,
"SBC": 0xda000000, "SBCW": 0x5a000000,
"SBCS": 0xfa000000, "SBCSW": 0x7a000000,
"MUL": 0x9b007c00, "MULW": 0x1b007c00,
// MNEG/MSUB and NGC/SBC with the complementing register preset to ZR.
"MNEG": 0x9b00fc00, "MNEGW": 0x1b00fc00,
"NGC": 0xda000000, "NGCW": 0x5a000000,
"NGCS": 0xfa000000, "NGCSW": 0x7a000000,
"NEGSW": 0x6b000000,
"MUL": 0x9b007c00, "MULW": 0x1b007c00,
"SMULH": 0x9b407c00, "UMULH": 0x9bc07c00,
}
for m, op := range dpsrExtra {
@@ -773,12 +943,20 @@ func init() {
a64InstrTable["VSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 21<<10}
a64InstrTable["VUSHR"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 1<<10}
a64InstrTable["VSRI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 17<<10}
a64InstrTable["VSSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 1<<10}
a64InstrTable["VSRA"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 17<<10}
a64InstrTable["VSRSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 9<<10}
a64InstrTable["VSLI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 21<<10}
a64InstrTable["VSQSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 29<<10}
a64InstrTable["VUQSHL"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 29<<10}
a64InstrTable["VLD1"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD1.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VST1"] = a64Enc{format: a64FVLDST}
a64InstrTable["VST1.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VLD1R"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD1R.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VLD4R"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD4R.P"] = a64Enc{format: a64FVLDST, op: 1}
}
// a64SimdVSpec is one arrangement-aware SIMD instruction: the 8B base word,
@@ -835,6 +1013,23 @@ func a64ElemLetter(s string) bool {
return false
}
// fpSimdArrs and fpAcrossArrs bound the arrangements the FP SIMD forms
// accept: H, S and D widths for the pairwise data-processing, H and S for
// the across-vector reductions.
var fpSimdArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D)
var fpAcrossArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S)
// a64SimdQOnly names the forms whose arrangement contributes the 128-bit
// flag alone, without the size bits: the FP converts, the FP round-to-integral
// and pairwise compares among them. Word-verified against go tool asm.
var a64SimdQOnly = map[string]bool{
"VSCVTF": true, "VUCVTF": true, "VFCVTZS": true, "VFCVTZU": true,
"VFABS": true, "VFNEG": true, "VFSQRT": true,
"VFRINTN": true, "VFRINTP": true, "VFRINTM": true, "VFRINTZ": true,
"VFADDP": true, "VFMAXP": true, "VFMAXNMP": true,
"VFMAXV": true, "VFMAXNMV": true,
}
// a64ArrBits carries the fixed bits an arrangement contributes to the
// three-same word shape: the element size at bits 23:22 and the 128-bit
// flag at bit 30. Bit 29 belongs to the instruction's own base.
@@ -854,19 +1049,96 @@ var a64ArrBits = [a64ArrCount]uint32{
// instructions (word = base | arrBits | Rm<<16 | Rn<<5 | Rd). Every base
// word and arrangement bit was read off go tool asm.
var a64SimdVTable = map[string]a64SimdVSpec{
"VADD": {0x0e208400, 0x7f, false},
"VSUB": {0x2e208400, 0x7f, false},
"VMUL": {0x0e209c00, 0x3f, false}, // no 2D: integer multiply stops at 4S
"VAND": {0x0e201c00, 0x03, false}, // logical ops accept 8B and 16B only
"VEOR": {0x2e201c00, 0x03, false},
"VORR": {0x0ea01c00, 0x03, false},
"VADDP": {0x0e20bc00, 0x7f, false},
"VZIP1": {0x0e003800, 0x7f, false},
"VZIP2": {0x0e007800, 0x7f, false},
"VCMEQ": {0x2e208c00, 0x7f, false},
"VRAX1": {0xce608c00, 1 << a64Arr2D, true}, // SHA3 group, D2 only
"VPMULL": {0x0e20e000, 1<<a64Arr8B | 1<<a64ArrD1, false},
"VPMULL2": {0x0e20e000, 1<<a64Arr16B | 1<<a64Arr2D, false},
"VADD": {0x0e208400, 0x7f, false},
"VSUB": {0x2e208400, 0x7f, false},
"VMUL": {0x0e209c00, 0x3f, false}, // no 2D: integer multiply stops at 4S
"VAND": {0x0e201c00, 0x03, false}, // logical ops accept 8B and 16B only
"VEOR": {0x2e201c00, 0x03, false},
"VORR": {0x0ea01c00, 0x03, false},
"VADDP": {0x0e20bc00, 0x7f, false},
"VZIP1": {0x0e003800, 0x7f, false},
"VZIP2": {0x0e007800, 0x7f, false},
"VCMEQ": {0x2e208c00, 0x7f, false},
"VCMGE": {0x0e203c00, 0x7f, false},
"VCMGT": {0x0e203400, 0x7f, false},
"VCMHI": {0x2e203400, 0x7f, false},
"VCMHS": {0x2e203c00, 0x7f, false},
// FP compares take H, S and D arrangements only (the toolchain rejects
// the byte forms), and VFCMLE/VFCMLT have no register form at all.
"VFCMEQ": {0x0e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFCMGE": {0x2e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFCMGT": {0x2ea0e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
// FP arithmetic shares the same arrangement restriction.
"VFADD": {0x0e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFSUB": {0x0ea0d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMUL": {0x2e20dc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFDIV": {0x2e20fc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAX": {0x0e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMIN": {0x0ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXNM": {0x0e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINNM": {0x0ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMLA": {0x0e20cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMLS": {0x0ea0cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
// Saturating, halving, polynomial and pairwise arithmetic, the logical
// VBIT/VBSL family and the FP pairwise forms: word-verified against go
// tool asm.
"VBIC": {0x0e601c00, 0x7f, false},
"VBIF": {0x2ee01c00, 0x7f, false},
"VBIT": {0x6ea01c00, 0x7f, false},
"VBSL": {0x6e601c00, 0x7f, false},
"VCMTST": {0x0e208c00, 0x7f, false},
"VFADDP": {0x2e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXP": {0x2e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINP": {0x6ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXNMP": {0x2e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINNMP": {0x6ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VMLA": {0x4ea09400, 0x7f, false},
"VMLS": {0x6ea09400, 0x7f, false},
"VORN": {0x4ee01c00, 0x7f, false},
"VSHADD": {0x4ea00400, 0x7f, false},
"VSRHADD": {0x4ea01400, 0x7f, false},
"VUHADD": {0x6ea00400, 0x7f, false},
"VURHADD": {0x6ea01400, 0x7f, false},
"VSMAX": {0x4ea06400, 0x7f, false},
"VSMIN": {0x4ea06c00, 0x7f, false},
"VSMAXP": {0x4ea0a400, 0x7f, false},
"VSMINP": {0x4ea0ac00, 0x7f, false},
"VUMAX": {0x2e206400, 0x7f, false},
"VUMIN": {0x2e206c00, 0x7f, false},
"VUMAXP": {0x6ea0a400, 0x7f, false},
"VUMINP": {0x6ea0ac00, 0x7f, false},
"VSQADD": {0x4ea00c00, 0x7f, false},
"VUQADD": {0x6ea00c00, 0x7f, false},
"VSQSUB": {0x4ea02c00, 0x7f, false},
"VUQSUB": {0x6ea02c00, 0x7f, false},
"VSSHL": {0x4ee04400, 0x7f, false},
"VUSHL": {0x6ee04400, 0x7f, false},
"VUZP1": {0x0e001800, 0x7f, false},
"VUZP2": {0x4ec05800, 0x7f, false},
"VTRN1": {0x4ec02800, 0x7f, false},
"VTRN2": {0x4ec06800, 0x7f, false},
"VRAX1": {0xce608c00, 1 << a64Arr2D, true}, // SHA3 group, D2 only
"VPMULL": {0x0e20e000, 1<<a64Arr8B | 1<<a64ArrD1, false},
"VPMULL2": {0x0e20e000, 1<<a64Arr16B | 1<<a64Arr2D, false},
}
// a64SimdVZero holds the compare-against-zero words of the SIMD compares
// spelled with a $0 first operand (word = base | arrBits | Rn<<5 | Rd).
// VCMHI and VCMHS have no zero form: the toolchain reports an illegal
// combination for them, so they stay out and the encoder rejects the shape.
var a64SimdVZero = map[string]uint32{
"VCMEQ": 0x0e209800,
"VCMGT": 0x0e208800,
"VCMGE": 0x2e208800,
"VCMLT": 0x0e20a800,
"VCMLE": 0x2e209800,
// FP compares against (0.0): the register forms above carry the U and op
// bits; the zero forms reshape them.
"VFCMEQ": 0x0ea0d800,
"VFCMGE": 0x2ea0c800,
"VFCMGT": 0x0ea0c800,
"VFCMLE": 0x2ea0d800,
"VFCMLT": 0x0ea0e800,
}
// a64SimdV2Table holds the arrangement-aware two-register SIMD instructions
@@ -875,8 +1147,41 @@ var a64SimdVTable = map[string]a64SimdVSpec{
var a64SimdV2Table = map[string]a64SimdVSpec{
"VREV32": {0x2e200800, 1<<a64Arr8B | 1<<a64Arr16B | 1<<a64Arr4H | 1<<a64Arr8H, false},
"VREV64": {0x0e200800, 0x3f, false},
"VREV16": {0x0e201800, 1<<a64Arr8B | 1<<a64Arr16B, false},
"VUADDLV": {0x2e303800, 0x3f, false},
"VMOV": {0x0ea01c00, 1<<a64Arr8B | 1<<a64Arr16B, false},
// Two-register data-processing across one arrangement.
"VABS": {0x0e20b800, 0x7f, false},
"VNEG": {0x2e20b800, 0x7f, false},
"VCLS": {0x0e204800, 0x7f, false},
"VCLZ": {0x2e204800, 0x7f, false},
"VCNT": {0x0e205800, 0x7f, false},
"VNOT": {0x2e205800, 0x7f, false},
"VSQABS": {0x0e207800, 0x7f, false},
"VSQNEG": {0x2e207800, 0x7f, false},
"VRBIT": {0x6e605800, 0x7f, false},
"VSCVTF": {0x4e21d800, fpSimdArrs, false},
"VUCVTF": {0x6e21d800, fpSimdArrs, false},
"VFCVTZS": {0x4ea1b800, fpSimdArrs, false},
"VFCVTZU": {0x6ea1b800, fpSimdArrs, false},
"VFABS": {0x0ea0f800, fpSimdArrs, false},
"VFNEG": {0x2ea0f800, fpSimdArrs, false},
"VFSQRT": {0x2ea1f800, fpSimdArrs, false},
"VFRINTN": {0x0e218800, fpSimdArrs, false},
"VFRINTP": {0x0ea18800, fpSimdArrs, false},
"VFRINTM": {0x0e219800, fpSimdArrs, false},
"VFRINTZ": {0x0ea19800, fpSimdArrs, false},
// Across-vector reductions: the operand arrangement rides as usual and
// the destination stays a bare V register.
"VADDV": {0x0e31b800, 0x3f, false},
"VSMAXV": {0x0e30a800, 0x3f, false},
"VSMINV": {0x0e31a800, 0x3f, false},
"VUMAXV": {0x2e30a800, 0x3f, false},
"VUMINV": {0x2e31a800, 0x3f, false},
"VFMAXV": {0x2e30f800, fpAcrossArrs, false},
"VFMINV": {0x2eb0f800, fpAcrossArrs, false},
"VFMAXNMV": {0x2e30c800, fpAcrossArrs, false},
"VFMINNMV": {0x2eb0c800, fpAcrossArrs, false},
}
// a64CryptoArr is the arrangement each crypto instruction's operands must
@@ -904,6 +1209,15 @@ var a64MRSOps = map[string]uint32{
"ID_AA64ISAR1_EL1": 0xd5380620, "CNTFRQ_EL0": 0xd53be000,
"CNTPCT_EL0": 0xd53be020, "CNTVCT_EL0": 0xd53be040,
"DCZID_EL0": 0xd53b00e0, "DIT": 0xd53b42a0, "ID_AA64ZFR0_EL1": 0xd5380480,
"NZCV": 0xd53b4200, "FPCR": 0xd53b4400, "FPSR": 0xd53b4420,
}
// a64MSRRegOps maps the system register names GOROOT writes through the
// MSR (register) form, spelled in Go assembly as MOVD Rn, <sysreg> or
// MSR Rn, <sysreg>; the source register rides bits 4:0.
var a64MSRRegOps = map[string]uint32{
"NZCV": 0xd51b4200, "FPCR": 0xd51b4400, "FPSR": 0xd51b4420,
"ELR_EL1": 0xd5184020,
}
// a64MSROps maps the system register names GOROOT writes to their fixed
+288 -13
View File
@@ -4,6 +4,7 @@
package asm
import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -104,6 +105,7 @@ func TestArm64RegNum(t *testing.T) {
}{
{"R0", 0}, {"R4", 4}, {"R29", 29}, {"R30", 30}, {"R31", 31},
{"FP", 29}, {"LR", 30}, {"LINK", 30}, {"SP", 31}, {"ZR", 31},
{"R18_PLATFORM", 18},
{"F0", 0}, {"F4", 4}, {"F31", 31},
{"INVALID", -1}, {"X0", -1}, {"", -1},
}
@@ -674,6 +676,36 @@ func TestArm64AcquireRelease(t *testing.T) {
}
}
// TestArm64BTI pins the landing-pad family against the toolchain words:
// only the uppercase C/J/JC spellings assemble, and bare BTI is a
// diagnostic, never a panic.
func TestArm64BTI(t *testing.T) {
got := arm64Words(t, "\tBTI C\n\tBTI J\n\tBTI JC\n")
want := []uint32{
0xd503245f, // BTI C
0xd503249f, // BTI J
0xd50324df, // BTI JC
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("got %d words, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %#x, want %#x", i, got[i], want[i])
}
}
for _, src := range []string{"\tBTI\n", "\tBTI c\n", "\tBTI B\n"} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+src+"\tRET\n")
if len(errs) > 0 {
continue
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("BTI spelling %q should be rejected, as go tool asm rejects it", src)
}
}
}
// TestArm64System pins BRK, SVC, the barriers, cache maintenance and the
// system register accesses.
func TestArm64System(t *testing.T) {
@@ -830,6 +862,99 @@ func TestArm64SIMDElement(t *testing.T) {
}
}
// TestArm64GPIntoVector pins the whole-vector moves VMOV/VDUP Rs, Vd.<T>
// against `go tool asm -S` output (Go 1.27, arm64): word = Q | 7<<25 |
// imm5<<16 | 3<<10 | rs<<5 | rd, shared by both mnemonics, the form
// sys_windows_arm64.s and the bytealg loops use. The D1 destination is
// rejected, as the toolchain rejects it.
func TestArm64GPIntoVector(t *testing.T) {
got := arm64Words(t, "\tVMOV R5, V5.B16\n\tVMOV R1, V2.B8\n\tVMOV R3, V4.H4\n"+
"\tVMOV R9, V10.S4\n\tVMOV R7, V31.H8\n\tVMOV R11, V12.D2\n"+
"\tVDUP R5, V5.B16\n\tVDUP R9, V10.H8\n\tVMOV V4.B16, V20.B16\n")
want := []uint32{
0x4e010ca5, // VMOV R5, V5.B16
0x0e010c22, // VMOV R1, V2.B8
0x0e020c64, // VMOV R3, V4.H4
0x4e040d2a, // VMOV R9, V10.S4
0x4e020cff, // VMOV R7, V31.H8
0x4e080d6c, // VMOV R11, V12.D2
0x4e010ca5, // VDUP R5, V5.B16 (same word as VMOV)
0x4e020d2a, // VDUP R9, V10.H8
0x4ea41c94, // VMOV V4.B16, V20.B16 (vector to vector stays ORR)
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tVMOV R7, V8.D1\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("VMOV R7, V8.D1 assembled, want an arrangement error")
}
}
// TestArm64SimdTwoOperand pins the two-operand accumulate spellings
// VADD/VSUB Vm, Vn against `go tool asm -S` output (Go 1.27, arm64):
// word = 5<<28|7<<25|7<<21|1<<15|1<<10 for VADD (7<<28 for VSUB) with
// rf<<16 | rn<<5 | rn, bare V registers only (asm7.go case 89).
func TestArm64SimdTwoOperand(t *testing.T) {
got := arm64Words(t, "\tVADD V7, V8\n\tVSUB V7, V8\n\tVADD V1, V2\n\tVADD V0.B16, V1.B16, V2.B16\n")
want := []uint32{
0x5ee78508, // VADD V7, V8
0x7ee78508, // VSUB V7, V8
0x5ee18442, // VADD V1, V2
0x4e208422, // VADD arranged: the ordinary three-register path
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64TruncMove pins the truncating register moves against
// `go tool asm -S` output (Go 1.27, arm64): the signed forms lower to SXTB,
// SXTH and SXTW (SBFM), the unsigned byte and halfword forms to UXTB and
// UXTH (UBFM), MOVWU to a W ORR, and a narrow move out of the zero register
// drops to the W ORR too (asm7.go case 45).
func TestArm64TruncMove(t *testing.T) {
got := arm64Words(t, "\tMOVB R3, R4\n\tMOVH R5, R6\n\tMOVW R9, R10\n"+
"\tMOVBU R3, R4\n\tMOVHU R3, R4\n\tMOVWU R3, R4\n\tMOVD R3, R4\n"+
"\tMOVD ZR, R4\n\tMOVB ZR, R4\n\tMOVWU ZR, R5\n")
want := []uint32{
0x93401c64, // MOVB = SXTB
0x93403ca6, // MOVH = SXTH
0x93407d2a, // MOVW = SXTW
0xd3401c64, // MOVBU = UXTB
0xd3403c64, // MOVHU = UXTH
0x2a0303e4, // MOVWU = ORR W
0xaa0303e4, // MOVD = ORR X
0xaa1f03e4, // MOVD ZR, R4 keeps the X form
0x2a1f03e4, // MOVB ZR, R4 drops to the W form
0x2a1f03e5, // MOVWU ZR, R5
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDLoadStore pins the structure loads and stores.
func TestArm64SIMDLoadStore(t *testing.T) {
got := arm64Words(t, "\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tVLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]\n"+
@@ -1295,20 +1420,170 @@ func TestArm64ExclNoOffset(t *testing.T) {
}
}
// TestArm64AddSubImmRange: immediates that cannot ride the imm12 field are
// rejected instead of wrapping through int32.
func TestArm64AddSubImmRange(t *testing.T) {
for _, body := range []string{
"\tADD $0x100000000, R0, R1\n",
"\tSUB $-0x100000000, R0, R1\n",
"\tCMP $0x100000000, R0\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
// TestArm64AddSubImmWide pins the wide-immediate classification the toolchain
// applies to the ADD/SUB family (asm7.go cases 48, 62, 13): the ADDCON2 split
// into two imm12 instructions for plain ADD/SUB, the bitmask ORR into REGTMP,
// and the MOVZ/MOVN/MOVK materialisations followed by the register form.
// Comparisons never split, and the W forms classify the 32-bit value. Every
// word is go tool asm's own for the same source.
func TestArm64AddSubImmWide(t *testing.T) {
got := arm64Words(t, strings.Join([]string{
"\tADD $0xaaaaaa, R2, R3",
"\tSUB $0xaaaaaa, R2",
"\tADD $0x186a0, R2, R5",
"\tADD $0x1ffe00, R2, R3",
"\tADD $0x3fffffffc000, R5",
"\tADD $-100000, R2, R3",
"\tADD $-2048, R2, R3",
"\tCMP $0xaaaaaa, R2",
"\tCMP $0xffffffffffa0, R3",
"\tCMPW $27745, R2",
"\tCMPW $0x60060, R2",
"\tADDS $0xaaaaaa, R2, R3",
"\tADD $0x12345678, R2, R3",
"\tADDW $0x60060, R2",
"\tSUB $0xe7791f700, R3, R1",
"\tADDW $0x12345678, R2, R3",
"\tCMN $0x1000000, R2",
}, "\n")+"\n")
want := []uint32{
0x912aa843, 0x916aa863, // ADD $0xaaaaaa, R2, R3: ADDCON2 split
0xd12aa842, 0xd16aa842, // SUB $0xaaaaaa, R2: split with Rd = Rn
0x911a8045, 0x914060a5, // ADD $0x186a0, R2, R5: split
0xb2772ffb, 0x8b1b0043, // ADD $0x1ffe00: bitmask beats the split
0xb2727ffb, 0x8b1b00a5, // ADD $0x3fffffffc000: bitmask into REGTMP
0x9290d3fb, 0xf2bfffdb, 0x8b1b0043, // ADD $-100000: MOVN + MOVK
0x9280fffb, 0x8b1b0043, // ADD $-2048: single MOVN + ADD
0xd295555b, 0xf2a0155b, 0xeb1b005f, // CMP: never split, MOVZ + MOVK
0x92800bfb, 0xf2e0001b, 0xeb1b007f, // CMP $0xffffffffffa0: MOVN + fixup
0x528d8c3b, 0x6b1b005f, // CMPW $27745: W movcon, single MOVZW
0x52800c1b, 0x72a000db, 0x6b1b005f, // CMPW $0x60060: S form skips the split
0xd295555b, 0xf2a0155b, 0xab1b0043, // ADDS $0xaaaaaa: MOVZ + MOVK + ADDS
0xd28acf1b, 0xf2a2469b, 0x8b1b0043, // ADD $0x12345678: MOVZ + MOVK
0x11018042, 0x11418042, // ADDW $0x60060: W split
0xd29ee01b, 0xf2aef23b, 0xf2c001db, 0xcb1b0061, // SUB $0xe7791f700
0x528acf1b, 0x72a2469b, 0x0b1b0043, // ADDW $0x12345678: MOVZW + MOVKW
0xd2a0201b, 0xab1b005f, // CMN $0x1000000: single MOVZ + CMN
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("wide word %d = %08x, want %08x", i, got[i], want[i])
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
// TestArm64CarryImmWide pins the carry family's $0 spellings in two and
// three operands, the ROR shift on the logical group (and its rejection for
// the arithmetic forms), the NGC/MNEG zero-register aliases and the vector
// alias with an element selector. Words are go tool asm's own.
func TestArm64CarryShiftAlias(t *testing.T) {
got := arm64Words(t, "\tADC $0, R20\n\tADC $0, R20, R4\n\tSBCS $0, R4, R12\n"+
"\tSBCS R15, R4, R12\n\tANDW R9@>7, R19, R26\n\tAND R1@>33, R2, R3\n"+
"\tNEGSW R23<<1, R30\n\tNGC R2, R7\n\tMNEG R14, R27, R23\n")
want := []uint32{
0x9a1f0294, // ADC ZR, R20, R20
0x9a1f0284, // ADC ZR, R20, R4
0xfa1f008c, // SBCS ZR, R4, R12
0xfa0f008c, // SBCS R15, R4, R12
0x0ac91e7a, // ANDW R9 ROR 7, R19, R26
0x8ac18443, // AND R1 ROR 33, R2, R3
0x6b1707fe, // SUBSW ZR, R30, R23 LSL 1
0xda0203e7, // SBC ZR, R7, R2
0x9b0eff77, // MSUB ZR, R27, R14, R23
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("carry word %d = %08x, want %08x", i, got[i], want[i])
}
}
// ROR on an arithmetic form is unallocated: the toolchain reports an
// unsupported shift operator.
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tADD R1@>33, R2, R3\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Error("ADD R1@>33: expected an error, got none")
}
}
// TestArm64VecAliasElement pins the register-alias rewrite inside a vector
// operand with an element selector and inside a split register list: the
// aliases resolve textually where the parser carries the selector apart from
// the name. Words are go tool asm's own.
func TestArm64VecAliasElement(t *testing.T) {
src := `#include "textflag.h"
#define POLY V15
#define ACC0 V8
#define ACC1 V9
TEXT ·f(SB), NOSPLIT, $0-0
VMOV R1, POLY.D[0]
VEOR POLY.B16, POLY.B16, POLY.B16
VLD1 (R0), [ACC0.B16]
VLD1.P (R0), [ACC0.B16, ACC1.B16]
VST1.P [ACC0.B16, ACC1.B16], 32(R1)
RET
`
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
got := leWords(img.Code)
want := []uint32{
0x4e081c2f, // INS V15.D[0], R1
0x6e2f1def, // VEOR V15.B16, V15.B16, V15.B16
0x4c407008, // VLD1 (R0), [V8.B16]
0x4cdfa008, // VLD1.P (R0), [V8.B16, V9.B16]
0x4c9fa028, // VST1.P [V8.B16, V9.B16], 32(R1)
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("vecalias word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64AddSubImmBeyond32 pins the materialisation the toolchain applies
// once the value leaves every imm12 form: a constant sequence into REGTMP
// (R27) followed by the register form. SUB $-0x100000000 is a bitmask
// immediate, so it rides the ORR form; the others take MOVZ. Words are go
// tool asm's own.
func TestArm64AddSubImmBeyond32(t *testing.T) {
got := arm64Words(t, "\tADD $0x100000000, R0, R1\n\tSUB $-0x100000000, R0, R1\n\tCMP $0x100000000, R0\n")
want := []uint32{
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
0x8b1b0001, // ADD R27, R0, R1
0xb2607ffb, // ORR $-4294967296, ZR, R27 (bitmask)
0xcb1b0001, // SUB R27, R0, R1
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
0xeb1b001f, // CMP R27, R0
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
+551 -38
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -28,7 +29,7 @@ import (
// emitted: the bytes match go tool asm only for NOSPLIT functions or
// zero-frame leaves, where the toolchain emits no guard either.
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
code, _, labels, _, _, err := assemble(t, nil)
code, _, labels, _, _, _, err := assemble(t, nil)
return code, labels, err
}
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
// rejects SB operands outright (single-function assembly cannot resolve
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
// file defines is recorded as an external relocation instead of failing
// the object-file emitters resolve it at link time.
// the object-file emitters resolve it at link time. goos selects the TLS
// access form: the empty default behaves as linux.
type linkInfo struct {
symbols map[string]bool
allowExternal bool
goos string
}
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
// keep the two-instruction form; shared linux does too, which gasm's raw
// path does not model and therefore does not select.
func (l *linkInfo) tlsOneInsn() bool {
switch l.goos {
case "", "linux", "freebsd":
return true
}
return false
}
// sbPatch is a function-relative static-symbol relocation: the disp32 field
@@ -66,7 +82,10 @@ type spadjStep struct {
// assemble encodes a TEXT body, returning the machine code, the static-symbol
// patch sites (for the file-level layout to resolve), the label table and the
// stack-adjustment boundaries.
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
if err := checkAdjspBalance(t); err != nil {
return nil, nil, nil, nil, nil, nil, err
}
fi := computeFrame(t)
chain := jumpChain(t)
resolve := func(name string) string {
@@ -82,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
// outgrows the short form.
long := make([]bool, len(t.Body))
sizes := make([]int, len(t.Body))
numTargets := make([]int, len(t.Body))
for i := range numTargets {
numTargets[i] = -1
}
offsets := map[string]int{}
pcs := make([]int, len(t.Body))
var guardJBlong, guardJBElong, moreJMPlong bool
poolSeen := map[string]bool{}
var poolList []floatPoolEntry
for {
guard := fi.guardLen(guardJBlong, guardJBElong)
pos := guard + len(fi.prologue)
for i := range numTargets {
numTargets[i] = -1
}
idxAtPc := map[int]int{}
for i, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
offsets[s.Name.Text] = pos
case *ast.Instr:
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
// The alignment pseudo-statement: its size is the
// padding to the next boundary at this very position,
// filled with NOPs at emission.
pad, err := pcAlignPad(pcAlignValue(s), pos)
if err != nil {
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
}
sizes[i] = pad
pcs[i] = pos
pos += pad
continue
}
sz, err := instrSize(s, fi, long[i], link)
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
sizes[i] = sz
pcs[i] = pos
idxAtPc[pos] = i
pos += sz
}
}
bodyLen := pos - (guard + len(fi.prologue))
// Expand any short jump whose displacement no longer fits rel8.
changed := false
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
// emission pass reads the same table after the loop converges. A
// target that is itself an unconditional local JMP is chased to the
// ultimate target: the toolchain's brloop pass collapses branch-to-
// branch chains before it encodes, so matching its bytes requires
// the same redirection.
for i := range numTargets {
numTargets[i] = -1
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if len(s.Operands) == 1 {
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
numTargets[i] = target
}
}
}
}
for i := range numTargets {
if numTargets[i] < 0 {
continue
}
tgt := numTargets[i]
for hop := 0; hop < len(t.Body); hop++ {
idx, ok := idxAtPc[tgt]
if !ok {
break
}
in, ok := t.Body[idx].(*ast.Instr)
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
break
}
if name, isLabel := labelName(in.Operands[0]); isLabel {
tgt = offsets[resolve(name)]
continue
}
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
next, okT := pcJumpTarget(t, idx, n, pcs)
if !okT {
break
}
tgt = next
continue
}
break // JMP through a register or memory: the chain ends
}
numTargets[i] = tgt
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if numTargets[i] >= 0 && !long[i] {
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
if !fits8(rel) {
long[i] = true
changed = true
}
}
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
@@ -203,6 +311,14 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{guardLen + len(fi.prologue), 8 + fi.size},
)
}
// frameBase is the SP delta the prologue leaves: 8 for the saved base
// pointer plus the frame, 0 frameless. bodyDelta tracks the ADJSP
// statements' straight-line sum, so a mid-body step's value is the
// frame base plus what the body has opened so far.
frameBase, bodyDelta := 0, 0
if fi.useFP {
frameBase = 8 + fi.size
}
pos := guardLen + len(fi.prologue)
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
@@ -218,18 +334,34 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{pos + epi, 0},
)
}
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
for _, entry := range pool {
if !poolSeen[entry.name] {
poolSeen[entry.name] = true
poolList = append(poolList, entry)
}
}
if len(code) != sizes[i] {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
}
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
for k := range ps {
ps[k].kind = RelCall
}
}
if strings.ToUpper(s.Mnemonic.Text) == "ADJSP" && len(s.Operands) == 1 && s.Operands[0].Imm.HasVal {
// The statement shifted SP mid-body: record the new running
// delta as the value in effect from just past the instruction.
v := s.Operands[0].Imm.Val
if s.Operands[0].Imm.Neg {
v = -v
}
bodyDelta += int(v)
steps = append(steps, spadjStep{pos + len(code), frameBase + bodyDelta})
}
patches = append(patches, ps...)
lines = append(lines, LineEntry{Offset: pos, Line: s.Pos().Line})
out = append(out, code...)
@@ -251,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
pos += len(suffix)
}
_ = pos
return out, patches, offsets, steps, lines, nil
return out, patches, offsets, steps, lines, poolList, nil
}
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
@@ -398,6 +530,83 @@ func computeFrame(t *ast.Text) frameInfo {
return fi
}
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
// the second instruction boundary after the branch.
func pcJumpOffset(op *ast.Operand) (int, bool) {
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
return 0, false
}
return int(op.Addr.Offset), true
}
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
// instruction statements after the jump itself (N = 0 is the jump's own
// address, the classic park loop), and the target is the start of the Nth
// one. It reports false when the count runs past the end of the function.
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
if n == 0 {
return pcs[j], true
}
seen := 0
for k := j + 1; k < len(t.Body); k++ {
if _, ok := t.Body[k].(*ast.Instr); !ok {
continue
}
seen++
if seen == n {
return pcs[k], true
}
}
return 0, false
}
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
// toolchain's asm6.go nop table); longer padding repeats the largest that
// fits, greedy from the end.
var x86Nops = [][]byte{
{0x90},
{0x66, 0x90},
{0x0F, 0x1F, 0x00},
{0x0F, 0x1F, 0x40, 0x00},
{0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
}
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
// the toolchain's fillnop.
func fillNOPs(p []byte) {
for len(p) > 0 {
m := min(len(p), len(x86Nops))
copy(p[:m], x86Nops[m-1])
p = p[m:]
}
}
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
// alignment must be a power of two in [8, 2048] and the padding runs to the
// next boundary (zero when the position is already aligned).
func pcAlignPad(align, pos int) (int, error) {
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
}
if lob := pos & (align - 1); lob != 0 {
return align - lob, nil
}
return 0, nil
}
// pcAlignValue reads a PCALIGN statement's alignment operand.
func pcAlignValue(s *ast.Instr) int {
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
return int(s.Operands[0].Imm.Val)
}
return 0 // rejected by pcAlignPad's range check
}
// hasCall reports whether the function body contains a CALL instruction.
func hasCall(t *ast.Text) bool {
for _, stmt := range t.Body {
@@ -412,6 +621,40 @@ func hasCall(t *ast.Text) bool {
return false
}
// checkAdjspBalance mirrors the toolchain's push/pop walk: every ADJSP
// shifts SP away from the entry state and every RET must see the shifts
// closed. The assembler's own prologue and epilogue contribute matching
// deltas on both sides, so the statements' straight-line sum must be zero
// at each RET; branches do not reset the walk, which runs over the program
// list in source order. go tool asm reports an offender as "unbalanced
// PUSH/POP" (verified against ADJSP $16 before a RET, accepted as a
// $16/$-16 pair, per-RET rather than per-function).
func checkAdjspBalance(t *ast.Text) error {
delta := 0
for _, stmt := range t.Body {
in, ok := stmt.(*ast.Instr)
if !ok {
continue
}
switch strings.ToUpper(in.Mnemonic.Text) {
case "ADJSP":
if len(in.Operands) != 1 || !in.Operands[0].Imm.HasVal {
continue // reported during emission
}
v := in.Operands[0].Imm.Val
if in.Operands[0].Imm.Neg {
v = -v
}
delta += int(v)
case "RET":
if delta != 0 {
return fmt.Errorf("unbalanced PUSH/POP")
}
}
}
return nil
}
// guardLen returns the byte length of the stack-split guard prefix. The
// final conditional branch (JBE, and JB in the big class) is 2 bytes in the
// short form and 6 in the long form.
@@ -538,7 +781,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
}
return jumpSize(mnem, long), nil
}
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
if err != nil {
return 0, err
}
@@ -573,9 +816,21 @@ func jumpSize(mnem string, long bool) int {
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
// jump label through the jump-to-jump chain before the offset lookup.
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnem := strings.ToUpper(s.Mnemonic.Text)
if mnem == "PCALIGN" {
// The layout pass already accounted the padding; emit the same
// amount of NOP bytes for the statement's own position.
pad, err := pcAlignPad(pcAlignValue(s), pc)
if err != nil {
return nil, nil, nil, err
}
out := make([]byte, pad)
fillNOPs(out)
return out, nil, nil, nil
}
var prefix []byte
if mnem == "RET" && fi.useFP {
prefix = fi.epilogue
@@ -583,6 +838,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
var code []byte
var ps []sbPatch
var pool []floatPoolEntry
var err error
if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
@@ -591,7 +847,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
// or the linker.
code, ps, err = encodeSBCall(s, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
for i := range ps {
ps[i].kind = RelCall
@@ -601,23 +857,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, nil, nil
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
return append(prefix, code...), nil, nil
return append(prefix, code...), nil, nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
} else {
code, ps, err = encodeNormal(s, fi, link)
code, ps, pool, err = encodeNormal(s, fi, link)
}
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
// Anchor the patch fields at function-relative positions: off indexes the
// disp32 field, after is the address just past the instruction.
@@ -626,49 +882,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, pool, nil
}
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
code, err := encodeBookkeeping(mnemUpper, s)
if err != nil {
return nil, nil, nil, err
}
return code, nil, nil, nil
}
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
// never as a 64-bit absolute immediate (verified against go tool asm).
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
s.Operands[0].Imm.Sym.Pseudo == "SB" {
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
e := &enc{}
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
}
return e.out, ps, nil, nil
}
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
// it out on the one-instruction TLS systems (linux and freebsd, not
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
// the macro GOROOT's go_tls.h defines, expands to exactly this
// statement, and the toolchain's pairing pass removes it whenever the
// following instruction's (TLS*1) index folds.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
return encodeTLSBaseLoad(s, fi, link)
}
_, size := splitSize(mnemUpper)
if size == 0 {
size = 8
}
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, size, fi, link)
o, err := operandFromAST(mnemUpper, op, size, fi, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
if p.tls {
ps[i].kind = RelTLSLE
}
}
return e.out, ps, nil
return e.out, ps, e.floatPoolList(), nil
}
// isBareTLS reports whether the operand is the bare TLS pseudo-register
// load source, the expansion of go_tls.h's get_tls(r) macro.
func isBareTLS(op *ast.Operand) bool {
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
op.Addr.Base == "" && op.Addr.Index == ""
}
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
// statement nops out: the following (TLS*1) access folds to a direct
// segment-absolute load. The two-instruction systems keep the segment load,
// nine bytes with the R_TLSLE patch site at the disp32.
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
if size == 0 {
size = 8
}
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
reg, ok := dst.(Reg)
if !ok || reg.isVec() {
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
}
if link == nil || link.tlsOneInsn() {
return nil, nil, nil, nil // noped out
}
seg := byte(0x64) // FS
if link.goos == "windows" {
seg = 0x65 // GS
}
e := &enc{}
i := &instr{
prefix: seg,
rexW: size == 8,
rexR: reg.idx >= 8,
opcode: []byte{0x8B},
modrm: 0x04 | (reg.idx&7)<<3,
sib: 0x25,
disp: le32(0),
tls: true,
}
if err := e.emit(i); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
}
return e.out, ps, nil, nil
}
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
// level, before operand conversion: the toolchain's shapes are FUNCDATA
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
// function body. The symbol reference must not run through the SB-operand
// path, which demands file-level resolution the statement never needs.
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
if len(s.Operands) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
}
a, b := s.Operands[0], s.Operands[1]
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil, nil
}
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
// target label, in the short (rel8) or long (rel32) form.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
// target label or from a numeric ±N(PC) instruction count, in the short
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
// numeric operand, negative when the operand is not one.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
if len(s.Operands) != 1 {
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
}
name, ok := labelName(s.Operands[0])
if !ok {
name, isLabel := labelName(s.Operands[0])
if !isLabel && numTarget < 0 {
return nil, fmt.Errorf("jump target must be a local label")
}
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
target, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
var target int
if isLabel {
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
t, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
}
target = t
} else {
target = numTarget
}
rel := int64(target - (pc + jumpSize(mnem, long)))
@@ -701,7 +1091,7 @@ func isSBCall(s *ast.Instr) bool {
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
if err != nil {
return nil, nil, err
}
@@ -742,6 +1132,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
return false
}
a := s.Operands[0].Addr
// ±N(PC) is the numeric relative form, the PC counts instructions from
// the branch: relative, not indirect.
if a.Base == "PC" || a.Index == "PC" {
return false
}
if a.Base != "" || a.Index != "" {
return true
}
@@ -758,7 +1153,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
@@ -775,8 +1170,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
var spReg = Reg{idx: 4, size: 8}
// operandFromAST converts a parsed operand into an encoder Operand, applying
// the frame translation to FP/SP pseudo-register operands.
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
// instruction's upper-case mnemonic, which the floating-point immediate gate
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
// accept one.
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
switch op.Kind {
case ast.OpImmediate:
if op.Imm.HasVal {
@@ -786,11 +1184,45 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Imm(v), nil
}
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
// $(-1.0) spelling (the constant-expression folder only folds
// integers, so that shape arrives with an empty Immediate and only
// the raw spelling carries the value). The toolchain rewrites it
// into a pooled-constant read on the SSE scalar paths and rejects
// it everywhere else.
if text, neg, ok := floatImmText(op); ok {
if !sseFloatImm[mnemUpper] {
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
}
return FloatImm{Text: text, Neg: neg}, nil
}
return nil, fmt.Errorf("non-integer immediate not supported")
case ast.OpAddr:
a := op.Addr
// A bracketed register range, [Z0-Z3]: the four-register source of
// the 4FMAPS/4VNNIW families. The range must span four consecutive
// same-width vector registers, exactly what the toolchain's parser
// takes; the EVEX quad-register emit path reads the low end.
if a.Range != nil {
lo, ok := ParseReg(a.Range.Lo)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
}
hi, ok := ParseReg(a.Range.Hi)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
}
if !lo.isVec() || lo.size != hi.size {
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
}
if hi.idx != lo.idx+3 {
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
}
return RegList{Lo: lo, Hi: hi}, nil
}
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
// symbol, not the address displacement.
if a.Sym != nil && a.Sym.Pseudo == "FP" {
@@ -822,12 +1254,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
// Memory with a real base register: (base), off(base), (base)(index*scale).
if a.Base != "" {
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
// spellings. The segment override prefixes a disp32 absolute
// reference with no relocation.
if a.Base == "GS" || a.Base == "FS" {
seg := byte(0x64)
if a.Base == "GS" {
seg = 0x65
}
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
}
base, ok := ParseReg(a.Base)
if !ok {
return nil, fmt.Errorf("unknown base register %q", a.Base)
}
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
if a.Index != "" {
if a.Index == "TLS" {
// off(base)(TLS*1): the thread-local annotation. The
// one-instruction TLS form folds it to off(TLS), the
// segment-prefixed absolute whose disp32 carries an
// R_TLS_LE patch site; the base register disappears
// from the encoding, exactly as the toolchain's
// progedit rewrites the address.
seg := byte(0x64) // FS on linux, freebsd, plan9
if link != nil && link.goos == "windows" {
seg = 0x65 // GS
}
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
}
if a.Index == "GS" || a.Index == "FS" {
// 0(CX)(GS): the segment annotation rides the base
// access as the override prefix.
m.Seg = 0x64
if a.Index == "GS" {
m.Seg = 0x65
}
return m, nil
}
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
@@ -838,6 +1302,21 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return m, nil
}
// Index-only memory: the VSIB form the gather/scatter families
// read, 8(X4*1). A scaled vector index addresses memory with no
// base register; the mod=00 SIB with base field 101 carries it.
if a.Index != "" {
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
}
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
}
// A bare displacement with no base: the absolute address form,
// MOVL $0xf1, 0xf1. No segment and no relocation.
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
return SegAbs{Disp: a.Offset, Size: size}, nil
}
// Bare register.
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if r, ok := ParseReg(a.Sym.Name); ok {
@@ -848,3 +1327,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return nil, fmt.Errorf("unsupported operand")
}
// floatImmText recovers a floating-point immediate's magnitude and sign from
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
// folder cannot read it, and only the verbatim operand text still carries
// the value. Anything that is not a number a float parser accepts reports
// not-ok, so every other shape keeps its existing diagnostic.
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
if op.Imm.Float != "" {
return op.Imm.Float, op.Imm.Neg, true
}
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
return "", false, false
}
// joinRaw spaced the token texts; the compact spelling is what matters.
compact := strings.ReplaceAll(op.Raw, " ", "")
inner, ok := strings.CutPrefix(compact, "$(")
if !ok || !strings.HasSuffix(inner, ")") {
return "", false, false
}
inner = strings.TrimSuffix(inner, ")")
inner = strings.TrimPrefix(inner, "+")
if s, ok := strings.CutPrefix(inner, "-"); ok {
neg = true
inner = s
}
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
return "", false, false
}
if _, err := strconv.ParseFloat(inner, 64); err != nil {
return "", false, false
}
return inner, neg, true
}
+154
View File
@@ -439,3 +439,157 @@ func TestSubSPEncodings(t *testing.T) {
}
}
}
// TestAssemblePseudoStatements runs LOCK/REP, BYTE/WORD and END through the
// full statement pipeline, pinned against go tool asm (Go 1.27, amd64). It
// asserts the three behaviours the toolchain shows: each prefix statement is
// a standalone byte with a PC of its own (so a label placed on the LOCK
// points at the F0), the data pseudo-ops write their literal bytes inline,
// and END terminates nothing (the statements after it still belong to the
// function and carry no trace of it).
func TestAssemblePseudoStatements(t *testing.T) {
fn := firstText(t, `
#include "textflag.h"
TEXT ·pseudo(SB), NOSPLIT, $0-0
pfx:
LOCK
CMPXCHGQ AX, (BX)
REP
MOVSQ
BYTE $0x0f
BYTE $0x1f
WORD $0x1234
END
BYTE $0x02
RET
`)
code, labels, err := Assemble(fn)
if err != nil {
t.Fatalf("Assemble: %v", err)
}
// go tool asm: f0 480fb103 f3 48a5 0f 1f 3412 02 c3
want := []byte{
0xf0,
0x48, 0x0f, 0xb1, 0x03,
0xf3, 0x48, 0xa5,
0x0f, 0x1f, 0x34, 0x12,
0x02, 0xc3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("pseudo statements:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
// The label sits on the LOCK byte, exactly where the toolchain's PC
// listing puts it.
if off := labels["pfx"]; off != 0 {
t.Errorf("label pfx = %d, want 0 (the LOCK's own byte)", off)
}
// The trailing BYTE lands where the layout says: after the 8 bytes of
// LOCK, CMPXCHGQ, REP and MOVSQ plus the 4 data bytes, END contributing
// none.
if code[12] != 0x02 {
t.Errorf("byte at 12 = %02x, want 02 (the BYTE after END)", code[12])
}
}
// TestAssembleAdjspBalance pins the toolchain's push/pop balance rule over
// ADJSP: the straight-line sum of the adjustments must be zero at each
// RET, branches in between counting for nothing (verified against go tool
// asm: ADJSP $16 before a RET is reported as "unbalanced PUSH/POP", a
// $16/$-16 pair with a JMP in between assembles).
func TestAssembleAdjspBalance(t *testing.T) {
// Balanced pair with a branch in between, bytes pinned from go tool asm.
fn := firstText(t, `
#include "textflag.h"
TEXT ·adjsp(SB), NOSPLIT, $0-0
ADJSP $16
JMP body
body:
ADJSP $-16
RET
`)
code, _, err := Assemble(fn)
if err != nil {
t.Fatalf("Assemble: %v", err)
}
want := []byte{0x48, 0x83, 0xEC, 0x10, 0xEB, 0x00, 0x48, 0x83, 0xC4, 0x10, 0xC3}
if hexBytes(code) != hexBytes(want) {
t.Errorf("adjsp pair:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
// Unbalanced at the RET: the toolchain diagnoses, so must we.
_, _, err = Assemble(firstText(t, `
#include "textflag.h"
TEXT ·unbalanced(SB), NOSPLIT, $0-0
ADJSP $16
RET
`))
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
t.Errorf("unbalanced ADJSP: err = %v, want unbalanced PUSH/POP", err)
}
// The check runs per RET: a closed pair before the first RET does not
// excuse an open adjustment before the second.
_, _, err = Assemble(firstText(t, `
#include "textflag.h"
TEXT ·tworet(SB), NOSPLIT, $0-0
ADJSP $8
ADJSP $-8
RET
mid:
ADJSP $8
RET
`))
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
t.Errorf("second RET with open ADJSP: err = %v, want unbalanced PUSH/POP", err)
}
// A framed function: the assembler's own prologue and epilogue
// contribute matching deltas, so the pair in the body still balances,
// and the bytes match go tool asm end to end.
fn = firstText(t, `
#include "textflag.h"
TEXT ·framed(SB), $16-8
ADJSP $8
ADJSP $-8
RET
`)
code, _, err = Assemble(fn)
if err != nil {
t.Fatalf("Assemble framed: %v", err)
}
want = []byte{
0x55, 0x48, 0x89, 0xE5, 0x48, 0x83, 0xEC, 0x10, // prologue
0x48, 0x83, 0xEC, 0x08, // ADJSP $8
0x48, 0x83, 0xC4, 0x08, // ADJSP $-8
0x48, 0x83, 0xC4, 0x10, 0x5D, // epilogue
0xC3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
// TestAssembleRegRange pins the bracketed register range at the statement
// level: exactly four consecutive same-width vector registers assemble, the
// toolchain's rejected shapes all report an error.
func TestAssembleRegRange(t *testing.T) {
asm := func(t *testing.T, op string) ([]byte, error) {
t.Helper()
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse %s: %v", op, errs)
}
code, _, err := Assemble(f.Decls[0].(*ast.Text))
return code, err
}
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
if _, err := asm(t, op); err != nil {
t.Errorf("%s: %v", op, err)
}
}
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
if _, err := asm(t, op); err == nil {
t.Errorf("%s: assembled, want an error", op)
}
}
}
+65 -10
View File
@@ -43,6 +43,9 @@ const (
stInfoShift = 4
rX8664PC32 = 2
// R_X86_64_32 (debug/elf): the absolute 32-bit address of a symbol, the
// R_ADDR shape a 4-byte DATA field carries.
rX8664Abs32 = 10
// R_X86_64_TPOFF32 (debug/elf): the local-exec TLS offset the stack
// guard loads from FS. 20 is R_X86_64_TLSLD, a different relocation.
rX8664TPOFF32 = 23
@@ -158,6 +161,50 @@ func (img *Image) ELFObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rX8664Abs64
case 4:
typ = rX8664Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Serialise the string tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -167,19 +214,13 @@ func (img *Image) ELFObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
// Section presence: .rela.text only when there are relocations.
hasRela := len(relas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Lay the file out: header, section data, section headers.
var out []byte
out = append(out, make([]byte, 64)...) // ELF header, filled last
@@ -214,7 +255,7 @@ func (img *Image) ELFObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -226,6 +267,17 @@ func (img *Image) ELFObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -284,6 +336,9 @@ func (img *Image) ELFObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
+108
View File
@@ -750,3 +750,111 @@ func readFormSkip(t *testing.T, r *ulebIter, form uint64) {
t.Fatalf("unsupported form %#x", form)
}
}
// TestELFObjectDataRelocation checks that a symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)") reaches the ELF object as a .rela.data entry: an
// absolute 64-bit relocation at the field's offset within .data, against
// the named symbol, external targets included.
func TestELFObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $24
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_X86_64_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_X86_64_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_X86_64_64), addend: 0, target: "extvar"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -18,12 +18,16 @@ const (
rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD page offset)
rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction)
rArm64Ldst64Lo12NC = 286 // R_AARCH64_LDST64_ABS_LO12_NC (64-bit LDR/STR page offset)
// R_AARCH64_ABS32 (debug/elf 258): the absolute 32-bit address of a
// symbol, the R_ADDR shape a 4-byte DATA field carries. ABS64 (257)
// lives with the DWARF fixup constants as rAARCH64Abs64.
rArm64Abs32 = 258
)
// ELFAARCH64Object returns the image as an ELF64 relocatable object file for
// AArch64 (EM_AARCH64, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFAARCH64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -133,6 +137,50 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rAARCH64Abs64
case 4:
typ = rArm64Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -142,18 +190,13 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -188,7 +231,7 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -200,6 +243,17 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -257,6 +311,9 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -197,3 +197,118 @@ TEXT ·add(SB), NOSPLIT, $0-24
t.Error("unexpected .rela.text section when there are no relocations")
}
}
// TestELFAARCH64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the AArch64 ELF object as a
// .rela.data entry: an R_AARCH64_ABS64 (ABS32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFAARCH64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_arm64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
obj, err := img.ELFAARCH64Object()
if err != nil {
t.Fatalf("ELFAARCH64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_AARCH64_ABS64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_AARCH64_ABS32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -23,12 +23,16 @@ const (
rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i)
rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st)
rLarchB26 = 66 // R_LARCH_B26 (b/bl, matches the Go linker's mapping)
// R_LARCH_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_LARCH_64 (2) lives
// with the DWARF fixup constants as rLarchAbs64.
rLarchAbs32 = 1
)
// ELFLOONG64Object returns the image as an ELF64 relocatable object file for
// LoongArch (EM_LOONGARCH, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFLOONG64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -117,6 +121,50 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rLarchAbs64
case 4:
typ = rLarchAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -126,18 +174,13 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -172,7 +215,7 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -184,6 +227,17 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -239,6 +293,9 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -245,3 +245,118 @@ func TestELFLOONG64BranchRelocation(t *testing.T) {
}
}
}
// TestELFLOONG64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the LoongArch ELF object as a
// .rela.data entry: an R_LARCH_64 (R_LARCH_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFLOONG64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_loong64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("AssembleFileLOONG64: %v", err)
}
obj, err := img.ELFLOONG64Object()
if err != nil {
t.Fatalf("ELFLOONG64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_LARCH_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_LARCH_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_LARCH_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_LARCH_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -10
View File
@@ -24,11 +24,16 @@ const (
rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20
rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I
rRISCVPCRELLO12S = 25 // R_RISCV_PCREL_LO12_S
// R_RISCV_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_RISCV_64 (2) lives
// with the DWARF fixup constants as rRISCVAbs64.
rRISVCAbs32 = 1
)
// ELFRISCVObject returns the image as an ELF64 relocatable object file for
// RISC-V (EM_RISCV, 64-bit, little-endian). The structure mirrors the amd64
// ELF emission: .text, .data, .symtab, .strtab and optional .rela.text.
// ELF emission: .text, .data, .symtab, .strtab, an optional .rela.text and
// an optional .rela.data.
func (img *Image) ELFRISCVObject() ([]byte, error) {
le := binary.LittleEndian
@@ -129,6 +134,50 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rRISCVAbs64
case 4:
typ = rRISVCAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -138,18 +187,13 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -184,7 +228,7 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -196,6 +240,17 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -251,6 +306,9 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+128
View File
@@ -0,0 +1,128 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"debug/elf"
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the RISC-V ELF object as a
// .rela.data entry: an R_RISCV_64 (R_RISCV_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFRISCVObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_riscv64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
obj, err := img.ELFRISCVObject()
if err != nil {
t.Fatalf("ELFRISCVObject: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_RISCV_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_RISCV_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_RISCV_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_RISCV_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+4 -1
View File
@@ -19,7 +19,10 @@ func Encodable(mnemonic string) bool {
// Fixed-name instructions (no size suffix).
switch upper {
case "RET", "NOP", "CALL", "JMP",
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2":
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
// The literal-data pseudo-ops, the accepted-and-ignored END and
// bookkeeping statements, and the SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
return true
}
if _, ok := noOperandTable[upper]; ok {
+307 -3
View File
@@ -5,6 +5,8 @@ package asm
import (
"fmt"
"math"
"strconv"
"strings"
)
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
type enc struct {
out []byte
patches []encPatch // disp32 fields awaiting static-symbol resolution
// FloatPool collects the pooled constants the floating-point
// immediates reference, in first-use order.
floatPool []floatPoolEntry
floatPoolSeen map[string]bool
}
// floatPoolEntry is one pooled floating-point constant: the symbol name
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
type floatPoolEntry struct {
name string
data []byte
}
// addFloatPool records a pooled constant, deduplicated by symbol name.
func (e *enc) addFloatPool(name string, bits uint64, width int) {
if e.floatPoolSeen == nil {
e.floatPoolSeen = map[string]bool{}
}
if e.floatPoolSeen[name] {
return
}
e.floatPoolSeen[name] = true
data := make([]byte, width)
for i := range width {
data[i] = byte(bits >> (8 * i))
}
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
}
// floatPoolList returns the pooled constants in first-use order.
func (e *enc) floatPoolList() []floatPoolEntry {
return e.floatPool
}
// encPatch marks a 4-byte displacement field in enc.out that must receive the
@@ -29,6 +64,7 @@ type encPatch struct {
off int
name string
addend int64
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
}
func (e *enc) encode(mnem string, ops []Operand) error {
@@ -92,6 +128,21 @@ func (e *enc) encode(mnem string, ops []Operand) error {
// SHA256RNDS2 carries the round constant in a literal X0 first operand.
case "SHA256RNDS2":
return e.encodeSha256rnds2(ops)
// BYTE, WORD, LONG and QUAD write the immediate into the text stream
// itself: 1, 2, 4 or 8 literal bytes, little-endian. END is accepted
// and ignored. ADJSP adjusts SP by the immediate, sign-chosen between
// the SUBQ and ADDQ forms.
case "BYTE", "WORD", "LONG", "QUAD":
return e.encodeData(upper, ops)
case "END":
return e.encodeEnd(ops)
case "ADJSP":
return e.encodeAdjsp(ops)
// The runtime's bookkeeping statements carry no text bytes: go tool asm
// records FUNCDATA and PCDATA in the program list only, so the encoded
// body shows nothing, on every architecture.
case "FUNCDATA", "PCDATA":
return e.encodeFuncdata(upper, ops)
}
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
@@ -103,6 +154,7 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return err
}
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) ||
isEvexPrefGather(base) ||
base == "KMOVW" || base == "KMOVQ" || base == "KMOVB" || base == "KMOVD" {
return e.encodeVec(base, ops, sfx)
}
@@ -130,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
}
// Legacy SSE packed binaries dispatch on the full name: the packed
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
// which the size split must not eat.
// which the size split must not eat. A floating-point immediate
// rewrites into a pooled-constant read on the scalar members.
if m, ok := sseBinTable[upper]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
if m, ok := sseBinTable[base]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
// The imm8-controlled legacy instructions, the lane extracts and inserts
@@ -173,7 +232,7 @@ func (e *enc) encode(mnem string, ops []Operand) error {
case "INC", "DEC", "NEG", "NOT", "MUL", "DIV", "IDIV":
return e.encodeUnary(unaryOp[base], ops, size)
case "SHL", "SHR", "SAR", "SAL", "ROL", "ROR", "RCL", "RCR":
return e.encodeShift(shiftOp[base], ops, size)
return e.encodeShift(base, ops, size)
case "BT", "BTS", "BTR", "BTC":
return e.encodeBitTest(base, ops, size)
case "XCHG":
@@ -211,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeCvtInt(base, ops, size)
case "FMOVD":
return e.encodeFmov(ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
case "MOVSD", "MOVSS":
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatMove(upper, f, ops)
}
return e.encodeSSEMove(sseMoveTable[base], ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
return e.encodeSSEMove(sseMoveTable[base], ops)
}
return fmt.Errorf("unsupported instruction %q", mnem)
@@ -241,6 +305,214 @@ var prefetchVariant = map[string]int{
"PREFETCHT2": 3,
}
// dataWidth is the literal byte count of each data-emission pseudo-op.
var dataWidth = map[string]int{
"BYTE": 1,
"WORD": 2,
"LONG": 4,
"QUAD": 8,
}
// encodeData emits the literal-data pseudo-ops: BYTE, WORD, LONG and QUAD
// write the immediate into the text stream as 1, 2, 4 or 8 bytes,
// little-endian, with no opcode lookup. The value is truncated to the
// width rather than range-checked, exactly as go tool asm behaves (BYTE
// $0x1FF emits FF, WORD $0x12345 emits 45 23, both without an error), and
// exactly one immediate is accepted: the toolchain rejects a list such as
// BYTE $1, $2, $3.
func (e *enc) encodeData(mnem string, ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 1 immediate operand, got %d", mnem, len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("%s requires an integer immediate", mnem)
}
width := dataWidth[mnem]
out := make([]byte, width)
u := uint64(imm)
for i := range width {
out[i] = byte(u >> (8 * i))
}
e.out = append(e.out, out...)
return nil
}
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
// for either (the entries live in the object's ancillary tables, not the
// function body), and the operand shapes it takes are exactly these: an
// integer count first, then a symbol reference for FUNCDATA and an integer
// value for PCDATA. The other architectures accept-and-ignore the same
// statements; amd64 now matches.
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
}
if _, ok := ops[0].(Imm); !ok {
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if _, ok := ops[1].(sbMem); !ok {
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if _, ok := ops[1].(Imm); !ok {
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil
}
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
// entirely: the AEND Prog is skipped when the program list is flushed, so
// the statements after an END still belong to the same function and the
// encoded body carries no trace of it, whatever operands follow the name
// (the toolchain takes END $0 and END AX alike). Zero bytes, no effect.
func (e *enc) encodeEnd(ops []Operand) error {
return nil
}
// encodeAdjsp emits ADJSP $imm: a positive value is SUBQ $imm, SP, a
// negative one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude
// picks (the same selection subSP and addSP make for the frame). go tool
// asm refuses ADJSP $0 outright, so a zero value is an error here too; the
// statement's effect on the SP balance is checked by the function-level
// assembly (checkAdjspBalance), as the toolchain's push/pop walk does.
func (e *enc) encodeAdjsp(ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("ADJSP expects 1 immediate operand, got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("ADJSP requires an integer immediate")
}
switch v := int(imm); {
case v > 0:
e.out = append(e.out, subSP(v)...)
case v < 0:
e.out = append(e.out, addSP(-v)...)
default:
return fmt.Errorf("ADJSP $0 has no encoding")
}
return nil
}
// --- floating-point immediates ----------------------------------------------
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
// immediate, the set go tool asm rewrites into a pooled-constant read: the
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
// reject the immediate in the toolchain and are absent here on purpose.
var sseFloatImm = map[string]bool{
"MOVSD": true, "MOVSS": true,
"ADDSD": true, "ADDSS": true,
"SUBSD": true, "SUBSS": true,
"MULSD": true, "MULSS": true,
"DIVSD": true, "DIVSS": true,
"COMISD": true, "COMISS": true,
"UCOMISD": true, "UCOMISS": true,
}
// floatImmOperand reports whether the operand list opens with a
// floating-point immediate in the two-operand spelling (imm, dst).
func floatImmOperand(ops []Operand) (FloatImm, bool) {
if len(ops) != 2 {
return FloatImm{}, false
}
f, ok := ops[0].(FloatImm)
return f, ok
}
// floatPoolValue evaluates a floating-point immediate at the width its
// mnemonic encodes and names the pool constant the toolchain synthesises:
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
// rounding of the parsed value). The name carries the IEEE-754 bits; the
// section holds them little-endian.
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
v, err := strconv.ParseFloat(f.Text, 64)
if err != nil {
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
}
if f.Neg {
v = -v
}
if strings.HasSuffix(mnem, "D") {
bits = math.Float64bits(v)
return bits, fmt.Sprintf("$f64.%016x", bits), nil
}
bits = uint64(math.Float32bits(float32(v)))
return bits, fmt.Sprintf("$f32.%08x", bits), nil
}
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
// source. A positive zero needs no memory read: the toolchain emits
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
// file-level layout or the linker resolves.
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
if bits == 0 {
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
if err := setRM(i, dst, dst, 8); err != nil {
return err
}
return e.emit(i)
}
m := sseMoveTable[mnem]
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
// a floating-point immediate source: the constant is read from the pool into
// the instruction's r/m side (reg = destination), the rewrite go tool asm
// performs at the source level.
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
// spellings end in D, the single spellings in S.
func mwidth(mnem string) int {
if strings.HasSuffix(mnem, "D") {
return 8
}
return 4
}
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
func splitSize(upper string) (base string, size int) {
if upper == "" {
@@ -307,6 +579,7 @@ type instr struct {
disp []byte
imm []byte
sb *sbRef // static-symbol displacement in disp, awaiting resolution
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
}
// sbRef records that an instruction's displacement refers to a static symbol
@@ -349,6 +622,9 @@ func (e *enc) emit(i *instr) error {
if i.sb != nil {
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
}
if i.tls {
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
}
e.out = append(e.out, i.disp...)
e.out = append(e.out, i.imm...)
return nil
@@ -403,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
i.disp = le32(0)
i.sb = &sbRef{name: r.name, addend: r.addend}
return nil
case TLSMem:
// off(TLS): the segment-prefixed absolute access, mod=00 with the
// SIB escape's disp32 absolute form. The displacement is the TLS
// slot offset, patched by the linker's TLS relocation.
i.prefix = r.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(r.Disp)
i.tls = true
return nil
case SegAbs:
// 0x30(GS): the segment override with the SIB escape's disp32
// absolute form, no relocation.
setSegAbs(i, regField, r)
return nil
default:
return fmt.Errorf("invalid r/m operand %T", rm)
}
}
func setMem(i *instr, regField int, m Mem) error {
if m.Seg != 0 {
i.prefix = m.Seg
}
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
if err != nil {
return err
@@ -421,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
return nil
}
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
// override with the mod=00 SIB escape's disp32 absolute form and no
// relocation.
func setSegAbs(i *instr, regField int, m SegAbs) {
i.prefix = m.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(m.Disp)
}
// memComponents computes the ModR/M byte (with the given reg field), the SIB
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
+352
View File
@@ -9,6 +9,9 @@ import (
"testing"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// decode encodes an instruction and decodes it back, returning the decoded
@@ -200,10 +203,60 @@ func TestUnary(t *testing.T) {
func TestShift(t *testing.T) {
checkSyntax(t, "shl rdx, 0x2", "SHLQ", Imm(2), DX)
checkSyntax(t, "shl rdx, cl", "SHLQ", CL, DX)
checkSyntax(t, "shl rdx, cl", "SHLQ", CX, DX)
checkSyntax(t, "shl rdx, 0x1", "SHLQ", Imm(1), DX)
checkSyntax(t, "sar rcx, 0x1f", "SARQ", Imm(31), CX)
}
// TestDoubleShift pins the three-operand SHL/SHR form, which encodes as
// SHLD/SHRD: go tool asm accepts it for SHL/SHR at W/L/Q widths and rejects
// it for SAR, SAL, the rotates and the B width. The byte pins mirror the
// oracle's objdump output (48 0f a4 fe 0d for the first case, and so on).
func TestDoubleShift(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string // hex encoding
}{
{"SHLQ imm", "SHLQ", []Operand{Imm(0x0d), DI, SI}, "480fa4fe0d"},
{"SHLQ CX high regs", "SHLQ", []Operand{CX, Reg{idx: 8, size: 8}, Reg{idx: 9, size: 8}}, "4d0fa5c1"},
{"SHRQ imm", "SHRQ", []Operand{Imm(1), AX, CX}, "480facc101"},
{"SHLW imm", "SHLW", []Operand{Imm(1), AX, CX}, "660fa4c101"},
{"SHRD CL", "SHRQ", []Operand{CL, AX, CX}, "480fadc1"},
{"SHLD imm high regs", "SHLQ", []Operand{Imm(2), Reg{idx: 10, size: 8}, Reg{idx: 11, size: 8}}, "4d0fa4d302"},
{"SHRD imm max", "SHRQ", []Operand{Imm(63), Reg{idx: 9, size: 8}, Reg{idx: 15, size: 8}}, "4d0faccf3f"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: bytes %s, want %s", c.name, got, c.want)
}
}
// Rejected forms: the oracle rejects every one of these.
rejected := []struct {
name string
mnem string
ops []Operand
}{
{"SARQ three operands", "SARQ", []Operand{Imm(1), AX, CX}},
{"SALQ three operands", "SALQ", []Operand{Imm(1), AX, CX}},
{"ROLQ three operands", "ROLQ", []Operand{Imm(1), AX, CX}},
{"SHLB three operands", "SHLB", []Operand{Imm(1), AL, CL}},
{"SHRQ memory source", "SHRQ", []Operand{Imm(1), Ptr(AX, 0, 8), CX}},
{"SHRQ ECX count", "SHRQ", []Operand{Reg{idx: 1, size: 4}, AX, CX}},
}
for _, c := range rejected {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: Encode succeeded, want rejection", c.name)
}
}
}
func TestImul(t *testing.T) {
checkSyntax(t, "imul rdx, rcx", "IMULQ", CX, DX)
checkSyntax(t, "imul edx, edx, 0x3", "IMULL", Imm(3), DX, DX)
@@ -277,6 +330,12 @@ func TestSSEMoveGroundTruth(t *testing.T) {
{"MOVSD (SI),X1", "MOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X1")}, "f20f100e", "MOVSD_XMM"},
{"MOVSD X1,X2", "MOVSD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "f20f10d1", "MOVSD_XMM"},
{"MOVSS X3,(DI)", "MOVSS", []Operand{vreg(t, "X3"), Ptr(DI, 0, 4)}, "f30f111f", "MOVSS"},
// Static-symbol (SB) references: the GOROOT crypto kernels load and
// store octa constants by name (MOVOU bswapMask<>+0(SB), X0).
{"MOVOU sym,X0", "MOVOU", []Operand{sbMem{size: 16, name: "bswapMask"}, vreg(t, "X0")}, "f30f6f0500000000", "MOVDQU"},
{"MOVOU X0,sym+8", "MOVOU", []Operand{vreg(t, "X0"), sbMem{size: 16, name: "bswapMask", addend: 8}}, "f30f7f0500000000", "MOVDQU"},
{"MOVO sym,X1", "MOVO", []Operand{sbMem{size: 16, name: "gcmPoly"}, vreg(t, "X1")}, "660f6f0d00000000", "MOVDQA"},
{"MOVO X2,sym", "MOVO", []Operand{vreg(t, "X2"), sbMem{size: 16, name: "gcmPoly"}}, "660f7f1500000000", "MOVDQA"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
@@ -869,3 +928,296 @@ func TestMOVQXMMGroundTruth(t *testing.T) {
}
}
}
// TestPrefixStatements pins LOCK, REP and REPN. go tool asm encodes each as
// a standalone one-byte instruction with a PC of its own (F0, F3, F2), not a
// prefix field merged into the following instruction, and it validates
// nothing about the pairing (LOCK before NOP assembles). The prefixed
// atomic and string shapes are the bytes the runtime's own kernels need.
func TestPrefixStatements(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"LOCK", "LOCK", nil, "f0"},
{"REP", "REP", nil, "f3"},
{"REPN", "REPN", nil, "f2"},
// LOCK; CMPXCHGQ AX, (BX)
{"LOCK CMPXCHGQ", "CMPXCHGQ", []Operand{AX, Ptr(BX, 0, 8)}, "480fb103"},
// REP; MOVSQ
{"REP MOVSQ", "MOVSQ", nil, "48a5"},
// REPN; MOVSB
{"REPN MOVSB", "MOVSB", nil, "a4"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// The prefix statements take no operands, as the toolchain reports for
// LOCK AX.
if _, err := Encode("LOCK", AX); err == nil {
t.Error("LOCK AX assembled, want an error")
}
if _, err := Encode("REP", Imm(1)); err == nil {
t.Error("REP $1 assembled, want an error")
}
}
// TestDataEmission pins BYTE, WORD, LONG and QUAD: the immediate lands in
// the text stream as 1, 2, 4 or 8 little-endian bytes with no opcode
// lookup, truncated to the width rather than range-checked (go tool asm
// emits FF for BYTE $0x1FF and 45 23 for WORD $0x12345, both silently).
func TestDataEmission(t *testing.T) {
cases := []struct {
name string
mnem string
imm Imm
want string
}{
{"BYTE", "BYTE", 0x0f, "0f"},
{"BYTE negative", "BYTE", -1, "ff"},
{"BYTE truncated", "BYTE", 0x1ff, "ff"},
{"WORD", "WORD", 0x1234, "3412"},
{"WORD negative", "WORD", -1, "ffff"},
{"WORD truncated", "WORD", 0x12345, "4523"},
{"LONG", "LONG", 0x11223344, "44332211"},
{"LONG negative", "LONG", -1, "ffffffff"},
{"QUAD", "QUAD", 0x1122334455667788, "8877665544332211"},
{"QUAD negative", "QUAD", -2, "feffffffffffffff"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.imm)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// Exactly one immediate: the toolchain rejects BYTE $1, $2, $3, and a
// register or a missing operand is no immediate at all.
if _, err := Encode("BYTE"); err == nil {
t.Error("BYTE with no operand assembled, want an error")
}
if _, err := Encode("BYTE", Imm(1), Imm(2)); err == nil {
t.Error("BYTE $1, $2 assembled, want an error")
}
if _, err := Encode("WORD", AX); err == nil {
t.Error("WORD AX assembled, want an error")
}
}
// TestEndIgnored pins END: go tool asm drops the statement entirely, so it
// encodes to zero bytes and takes any operands without complaint (the
// toolchain accepts END $0 and END AX alike).
func TestEndIgnored(t *testing.T) {
for _, ops := range [][]Operand{nil, {Imm(0)}, {AX}} {
code, err := Encode("END", ops...)
if err != nil {
t.Errorf("END: %v", err)
continue
}
if len(code) != 0 {
t.Errorf("END = %x, want no bytes", code)
}
}
}
// TestAdjsp pins ADJSP: a positive immediate is SUBQ $imm, SP, a negative
// one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude picks; $0
// has no encoding (go tool asm refuses ADJSP $0 outright).
func TestAdjsp(t *testing.T) {
cases := []struct {
name string
imm Imm
want string
}{
{"imm8", 112, "4883ec70"},
{"imm8 negative", -112, "4883c470"},
{"imm32", 200, "4881ecc8000000"},
{"imm32 negative", -200, "4881c4c8000000"},
{"small", 8, "4883ec08"},
}
for _, c := range cases {
code, err := Encode("ADJSP", c.imm)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("ADJSP %d = %s, want %s", int64(c.imm), got, c.want)
}
}
if _, err := Encode("ADJSP", Imm(0)); err == nil {
t.Error("ADJSP $0 assembled, want an error")
}
if _, err := Encode("ADJSP"); err == nil {
t.Error("ADJSP with no operand assembled, want an error")
}
if _, err := Encode("ADJSP", AX); err == nil {
t.Error("ADJSP AX assembled, want an error")
}
}
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
// byte for byte against go tool asm: the scalar moves and the scalar
// arithmetic read the constant from a synthesised read-only pool symbol
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
func TestFloatImmediateGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
// single spellings; negative zero keeps its sign bit and never takes the
// XORPS shortcut.
for _, c := range []struct {
mnem string
imm FloatImm
want string
}{
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
} {
_, name, err := floatPoolValue(c.mnem, c.imm)
if err != nil {
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
continue
}
if name != c.want {
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
}
}
// The shapes the toolchain's parser rejects: the packed and uniform
// forms, a non-vector destination, and the integer spellings.
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
// go tool asm emits no text bytes for either, on every architecture.
func TestBookkeepingGroundTruth(t *testing.T) {
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
} {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if len(code) != 0 {
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
}
}
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
// At the statement level the bookkeeping lines sit between real
// instructions and contribute nothing to the body, symbol reference
// included: the FUNCDATA operand never needs file-level resolution.
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
want := "90c3"
if got := hexCompact(img.Code); got != want {
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("FUNCDATA $1, X0 assembled, want an error")
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("PCDATA $1, X0 assembled, want an error")
}
// Encodable mirrors Encode for the names this work touched.
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
if !Encodable(mnem) {
t.Errorf("Encodable(%s) = false, want true", mnem)
}
}
}
// mustParse parses src or fails the test.
func mustParse(t *testing.T, src string) *ast.File {
t.Helper()
f, errs := parser.Parse("t_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
return f
}
+593 -24
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"slices"
"strings"
)
@@ -91,7 +92,7 @@ var evexTable = map[string]evexSpec{
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1, variable shift with an XMM count (VPSRAQ;
// the W bit distinguishes it from VPSRAD's E2 form).
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAQ": {1, 0x72, 1, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.F3.0F.W1, signed qword to packed double (reg=dst,
// rm=src, no vvvv).
@@ -138,7 +139,7 @@ var evexTable = map[string]evexSpec{
// EVEX.66.0F, the EVEX forms of the VEX two-source shuffle.
"VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F3A, lane insert ($imm, xsrc, zsrc1, zdst).
"VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
@@ -202,7 +203,7 @@ var evexTable = map[string]evexSpec{
"VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSLLVW": {2, 0x12, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x10, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSWB": {1, 0x63, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKUSWB": {1, 0x67, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -324,7 +325,7 @@ var evexTable = map[string]evexSpec{
"VCVTPD2UQQ": {1, 0x79, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 3, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F38, half-precision convert (half-width source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F3A, half-precision convert back ($imm, src, dst: reg=src,
@@ -494,6 +495,331 @@ var evexTable = map[string]evexSpec{
// destination (VPMOVDW dword→word, VPMOVQD qword→dword).
"VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
"VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
// --- the AVX-512 families the avx512enc corpus exercises, read off
// the toolchain opcodetables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VALIGNQ": {3, 0x03, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VANDNPD": {1, 0x55, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPD": {1, 0x54, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPD": {2, 0x65, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPS": {2, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBROADCASTF32X2": {2, 0x19, 0, 1, -1, vexRM, [3]int{0, 8, 8}},
"VBROADCASTF32X4": {2, 0x1A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF32X8": {2, 0x1B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTF64X2": {2, 0x1A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF64X4": {2, 0x1B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI32X2": {2, 0x59, 0, 1, -1, vexRM, [3]int{8, 8, 8}},
"VBROADCASTI32X4": {2, 0x5A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI32X8": {2, 0x5B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI64X2": {2, 0x5A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI64X4": {2, 0x5B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VCOMISD": {1, 0x2F, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VCVTSD2SS": {1, 0x5A, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VDBPSADBW": {3, 0x42, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VEXP2PD": {2, 0xC8, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VEXP2PS": {2, 0xC8, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213SD": {2, 0xA9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231SD": {2, 0xBD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8AFFINEQB": {3, 0xCE, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMOVNTDQ": {1, 0xE7, 0, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VMOVNTPD": {1, 0x2B, 1, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VORPD": {1, 0x56, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMB": {2, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMD": {2, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMQ": {2, 0x64, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMW": {2, 0x66, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBROADCASTMB2Q": {2, 0x2A, 1, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPBROADCASTMW2D": {2, 0x3A, 0, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPCLMULQDQ": {3, 0x44, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPCMPEQB": {1, 0x74, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQQ": {2, 0x29, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTQ": {2, 0x37, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCOMPRESSB": {2, 0x63, 0, 1, -1, vexRMRev, [3]int{1, 1, 1}},
"VPCOMPRESSW": {2, 0x63, 1, 1, -1, vexRMRev, [3]int{2, 2, 2}},
"VPCONFLICTD": {2, 0xC4, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPCONFLICTQ": {2, 0xC4, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPDPBUSD": {2, 0x50, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPBUSDS": {2, 0x51, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSD": {2, 0x52, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSDS": {2, 0x53, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PD": {2, 0x77, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PS": {2, 0x77, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2W": {2, 0x75, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}},
"VPERMT2B": {2, 0x7D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2PS": {2, 0x7F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2W": {2, 0x7D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPEXPANDB": {2, 0x62, 0, 1, -1, vexRM, [3]int{1, 1, 1}},
"VPEXPANDW": {2, 0x62, 1, 1, -1, vexRM, [3]int{2, 2, 2}},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm, [3]int{4, 0, 0}},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm, [3]int{8, 0, 0}},
"VPLZCNTD": {2, 0x44, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPLZCNTQ": {2, 0x44, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPMADD52HUQ": {2, 0xB5, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADD52LUQ": {2, 0xB4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULDQ": {2, 0x28, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULTISHIFTQB": {2, 0x83, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULUDQ": {1, 0xF4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPOPCNTW": {2, 0x54, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPORD": {1, 0xEB, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVD": {2, 0x15, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVQ": {2, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVD": {2, 0x14, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVQ": {2, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDD": {3, 0x71, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDQ": {3, 0x71, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDVD": {2, 0x71, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVQ": {2, 0x71, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVW": {2, 0x70, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDW": {3, 0x70, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDD": {3, 0x73, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDQ": {3, 0x73, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDVD": {2, 0x73, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVQ": {2, 0x73, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVW": {2, 0x72, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDW": {3, 0x72, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHUFBITQMB": {2, 0x8F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm, [3]int{16, 32, 64}},
"VPSRLDQ": {1, 0x73, 0, 1, 3, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.66.0F73 /7, the byte-quad shift left (the count is always an
// immediate; there is no register-count twin).
"VPSLLDQ": {1, 0x73, 0, 1, 7, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.W0, the plain-prefix (no 66) packed spellings
// whose EVEX form drops the legacy prefix entirely.
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VXORPS": {1, 0x57, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM, [3]int{16, 32, 64}},
"VCOMISS": {1, 0x2F, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev, [3]int{16, 32, 64}},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMB": {2, 0x26, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMD": {2, 0x27, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMQ": {2, 0x27, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMW": {2, 0x26, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMB": {2, 0x26, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMD": {2, 0x27, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMQ": {2, 0x27, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMW": {2, 0x26, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHQDQ": {1, 0x6D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLQDQ": {1, 0x6C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VRCP28PD": {2, 0xCA, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28PS": {2, 0xCA, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28SD": {2, 0xCB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRCP28SS": {2, 0xCB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VRSQRT28PD": {2, 0xCC, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28PS": {2, 0xCC, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28SD": {2, 0xCD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRSQRT28SS": {2, 0xCD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VSQRTPD": {1, 0x51, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VSQRTSD": {1, 0x51, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VUCOMISD": {1, 0x2E, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VXORPD": {1, 0x57, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.F3/F2.W0, word shuffles with an immediate
// ($imm, src, dst: reg = dst, rm = src, imm8). The F3/F2 prefixes
// split the high/low lane spellings.
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM, [3]int{16, 32, 64}},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.66.0F3A, lane extract to a general-purpose register or
// memory ($imm, xsrc, GPR/mem dst: reg = source, rm = destination).
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtractGPR, [3]int{1, 1, 1}},
"VPEXTRW": {3, 0x15, 0, 1, -1, vexExtractGPR, [3]int{2, 2, 2}},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtractGPR, [3]int{4, 4, 4}},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtractGPR, [3]int{8, 8, 8}},
// EVEX.66.0F3A.W1, the qword permutes with an immediate control
// ($imm, src, dst: reg = dst, rm = src, imm8); the register-count
// forms live in evexRegFormTable.
"VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMPD": {3, 0x01, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.66.0F3A, the packed permute shuffles with an immediate control.
"VPERMILPS": {3, 0x04, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMILPD": {3, 0x05, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.0F.W0, high/low half moves. VMOVHPS carries the
// three-operand insert form (rm = m64 source, vvvv = preserved,
// reg = dst) and the two-operand store (reg = source, rm = m64);
// the encoder splits on the operand count. VMOVLHPS is the
// three-operand form alone.
"VMOVHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
}
// evexQuad describes one quad-register instruction: the opcode under
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
// registers the bracketed list and the destination take (512-bit ZMM for
// the packed forms, 128-bit XMM for the scalar ones).
type evexQuad struct {
opcode byte
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
}
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
// families) to their encoding. The operand shape is fixed: a single memory
// source in r/m, the bracketed register list whose LOW register travels the
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
// destination in reg. The vector length follows the destination (512-bit
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
// multiplier stays 16 for every member, the toolchain's own tuple choice.
var evexQuadTable = map[string]evexQuad{
"V4FMADDPS": {0x9A, 64},
"V4FMADDSS": {0x9B, 16},
"V4FNMADDPS": {0xAA, 64},
"V4FNMADDSS": {0xAB, 16},
"VP4DPWSSD": {0x52, 64},
"VP4DPWSSDS": {0x53, 64},
}
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
func isEvexQuad(upper string) bool {
_, ok := evexQuadTable[upper]
return ok
}
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
// The register list is the VVVV-side source: its low register fills the
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
// suffix applies.
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
if len(ops) != 3 && len(ops) != 4 {
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
}
mem, lst := ops[0], ops[1]
dst := ops[len(ops)-1]
mask := 0
if len(ops) == 4 {
k, ok := ops[2].(Reg)
if !ok || !k.mask {
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
}
if k.idx == 0 {
return fmt.Errorf("k0 is not a usable mask register")
}
mask = k.idx
}
list, ok := lst.(RegList)
if !ok {
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
}
if list.Lo.size != q.width {
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if dstReg.size != q.width {
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
}
if !memOperand(mem) {
return fmt.Errorf("%s: the source must be a memory operand", mnem)
}
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
// its fifth bit into the same field the list's low register occupies.
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
}
if sfx.zeroing && mask == 0 {
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
}
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
// The vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones), exactly as the oracle encodes it.
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
}
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
@@ -529,32 +855,38 @@ type evexMoveSpec struct {
n [3]int
vecOK bool // the non-memory operand may be a vector register
xmmOnly bool // wider than XMM registers are rejected
nds3 bool // a three-operand register form exists (VMOVSD/VMOVSS)
}
// evexMoveTable maps an upper-case EVEX move mnemonic to its encoding.
var evexMoveTable = map[string]evexMoveSpec{
// EVEX.128/256/512.F3.0F.W0, unaligned integer move.
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F3.0F.W1, unaligned qword move.
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W0, unaligned byte move (byte/word moves use the
// F2 prefix, dword/qword moves F3; the element size only changes the tuple
// semantics).
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W1, unaligned word move (shares the qword
// encoding).
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F.W1, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false},
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512, aligned packed moves.
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false},
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F, aligned integer moves.
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128.F3.0F.W0, scalar single move, memory operands (the
// three-operand register form is not supported).
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true},
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true, true},
// EVEX.128.F2.0F.W1, scalar double move: memory operands and the
// three-operand register form (VMOVSD dst, src1, src2).
"VMOVSD": {1, 3, 0x10, 0x11, 1, [3]int{8, 8, 8}, false, true, true},
// EVEX.128/256/512.0F.W0, unaligned packed single move.
"VMOVUPS": {1, 0, 0x10, 0x11, 0, [3]int{16, 32, 64}, true, false, false},
}
// isEvex reports whether the mnemonic has an EVEX encoding we handle.
@@ -565,8 +897,10 @@ func isEvex(mnemUpper string) bool {
if _, ok := evexBcastTable[mnemUpper]; ok {
return true
}
_, ok := evexMoveTable[mnemUpper]
return ok
if _, ok := evexMoveTable[mnemUpper]; ok {
return true
}
return isEvexQuad(mnemUpper)
}
// evexRequired reports whether the operands force the EVEX encoding of a
@@ -579,6 +913,13 @@ func evexRequired(upper string, ops []Operand) bool {
if !inVex && !inVexMove {
return true // EVEX-only mnemonic
}
// The byte-quad shifts have VEX register forms but EVEX-only memory
// forms: a memory count source forces the EVEX encoding.
if upper == "VPSLLDQ" || upper == "VPSRLDQ" {
if slices.ContainsFunc(ops, memOperand) {
return true
}
}
for _, op := range ops {
if r, ok := op.(Reg); ok && (r.size == 64 || r.mask || (r.isVec() && r.idx >= 16)) {
return true
@@ -741,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexRM(spec, ops, 0, sfx)
}
spec, inTable := evexTable[mnemUpper]
if q, ok := evexQuadTable[mnemUpper]; ok {
// The quad-register family carries no rounding, SAE or broadcast;
// only masking and zeroing apply.
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
}
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
}
if inTable {
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
@@ -752,6 +1101,27 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
spec.n = [3]int{n, n, n}
}
// A mnemonic with an immediate and a register spelling (the
// variable-count shifts, the permutes) encodes the register one
// when the first operand is not an immediate.
if len(ops) > 0 {
if _, isImm := ops[0].(Imm); !isImm {
if alt, ok := evexRegFormTable[mnemUpper]; ok {
spec, inTable = alt, true
}
}
}
// The high/low half moves split by operand count: three operands
// insert, two store (VMOVHPS m64, X1).
if hs, ok := evexHptrTable[mnemUpper]; ok {
if len(ops) == 2 {
if hs.store.opcode == 0 {
return fmt.Errorf("%s has no two-operand form", mnemUpper)
}
return e.encodeEvexRMRev(hs.store, ops, 0, sfx)
}
spec = hs.insert
}
} else if sfx.evexOnly() {
return fmt.Errorf("%s: the instruction does not take rounding/SAE/broadcast suffixes", mnemUpper)
}
@@ -811,6 +1181,12 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
return e.encodeEvexMove(mnemUpper, ms, ops, mask, sfx)
}
if ps, ok := evexPrefGatherTable[mnemUpper]; ok {
if sfx.any() {
return fmt.Errorf("%s takes no EVEX suffixes", mnemUpper)
}
return e.encodeEvexPrefGather(mnemUpper, ps, ops, mask, sfx)
}
if !inTable {
return fmt.Errorf("unsupported instruction %q for ZMM/K operands", mnemUpper)
}
@@ -829,6 +1205,8 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexNDS3Imm(spec, ops, mask, sfx)
case vexExtract:
return e.encodeEvexExtract(spec, ops, mask, sfx)
case vexExtractGPR:
return e.encodeEvexExtractGPR(spec, ops, mask, sfx)
case vexRMSrcLen:
return e.encodeEvexRMSrcLen(spec, ops, mask, sfx)
}
@@ -908,6 +1286,11 @@ func (e *enc) encodeEvexImmRM(spec evexSpec, ops []Operand, mask int, sfx evexSu
if dstReg.mask {
if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
} else if l, err := soleLen(spec.n); err == nil {
// A memory source with a length-fixed mnemonic
// (VFPCLASSPDX/Y/Z): the length comes from the table's
// single valid slot, not from the operand.
ll = l
}
} else if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
@@ -934,9 +1317,11 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the length the L'L
// field and the disp8×N multiplier follow is the destination's either
// way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -946,7 +1331,7 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if err != nil {
return err
}
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, srcReg, mask, sfx); err != nil {
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, src, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
@@ -1018,10 +1403,77 @@ func (e *enc) encodeEvexExtract(spec evexSpec, ops []Operand, mask int, sfx evex
return nil
}
// encodeEvexExtractGPR encodes the lane extract to a general-purpose
// register or memory: OP $imm, xsrc, dst (reg = the XMM source, rm = the
// destination, imm8). The encoding is 128-bit regardless of register
// numbers, so L'L is fixed at 0 and the disp8×N multiplier is the extracted
// element size the table carries.
func (e *enc) encodeEvexExtractGPR(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 3 {
return fmt.Errorf("extract expects 3 operands ($imm, xsrc, dst), got %d", len(ops))
}
imm, src, dst := ops[0], ops[1], ops[2]
immVal, ok := imm.(Imm)
if !ok {
return fmt.Errorf("extract lane must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("extract source must be a vector register")
}
switch dst.(type) {
case Reg:
if dst.(Reg).isVec() {
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
case Mem, sbMem:
default:
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
immByte, err := imm8(int64(immVal))
if err != nil {
return err
}
if err := e.emitEvexFields(spec, 0, srcReg.idx, -1, dst, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
return nil
}
// encodeEvexMove encodes a two-operand EVEX move; a vector→vector move uses
// the store-form opcode (reg = source, rm = destination), matching the Go
// assembler.
// assembler. The scalar moves also carry a three-operand register form
// (VMOVSD dst, src1, src2: the load opcode with vvvv = src1), which ms.nds3
// opens.
func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) == 3 {
if !ms.nds3 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
// The masked scalar register form keeps the Go assembler's own
// layout: the store opcode with reg = op0, vvvv = op1 and the
// destination in r/m (op2) — the bytes go tool asm emits, not
// the manual's NDS reading.
src, src1, dst := ops[0], ops[1], ops[2]
reg, ok := src.(Reg)
if !ok || !reg.isVec() {
return fmt.Errorf("%s: first operand must be a vector register", mnem)
}
vvvvReg, ok := src1.(Reg)
if !ok || !vvvvReg.isVec() {
return fmt.Errorf("%s: second operand must be a vector register", mnem)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if ms.xmmOnly && (reg.size != 16 || vvvvReg.size != 16 || dstReg.size != 16) {
return fmt.Errorf("%s operates on XMM registers only", mnem)
}
spec := evexSpec{mapSel: ms.mapSel, opcode: ms.store, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n}
return e.emitEvexFields(spec, dstReg.vecLenBit(), reg.idx, vvvvReg.idx, dst, mask, sfx)
}
if len(ops) != 2 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
@@ -1141,12 +1593,20 @@ func (e *enc) encodeEvexBcast(bs evexBcastSpec, ops []Operand, mask int, sfx eve
return fmt.Errorf("broadcast destination must be a vector register")
}
spec := evexSpec{mapSel: bs.mapSel, w: bs.w, pp: 1, opdigit: -1}
switch src.(type) {
switch r := src.(type) {
case Mem, sbMem:
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
case Reg:
spec.opcode = bs.opReg
// A GPR source uses the register broadcast opcode; a vector
// source shares the xmm/mem one (the low byte is copied from
// the lane or from the memory operand).
if r.isVec() {
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
} else {
spec.opcode = bs.opReg
}
default:
return fmt.Errorf("broadcast source must be a register or memory")
}
@@ -1349,6 +1809,102 @@ func isScatter(upper string) bool {
return ok
}
// isEvexPrefGather reports whether the mnemonic is a gather/scatter
// prefetch hint.
func isEvexPrefGather(upper string) bool {
_, ok := evexPrefGatherTable[upper]
return ok
}
// evexRegFormTable holds the register-count twin of the immediate-form
// entries in evexTable. Several mnemonics name two encodings: an immediate
// count or control ($imm, src, dst …) and a register-count one whose second
// operand is a vector register or memory (count, src2, src1, dst). The
// immediate spelling lives in evexTable, this table carries the register
// spelling, and encodeEvex picks by whether the first operand is an
// immediate, the way vexVarShift does on the VEX side.
var evexRegFormTable = map[string]evexSpec{
"VPSLLD": {1, 0xF2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLQ": {1, 0xF3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLW": {1, 0xF1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAD": {1, 0xE2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAW": {1, 0xE1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLD": {1, 0xD2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLQ": {1, 0xD3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLW": {1, 0xD1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
// EVEX.NDS.0F38.W1, the register-count permutes (the immediate
// controls live in evexTable under 0F3A).
"VPERMQ": {2, 0x36, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPD": {2, 0x16, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.NDS.0F38, the register-count permil shuffles.
"VPERMILPS": {2, 0x0C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMILPD": {2, 0x0D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
}
// evexPrefGatherSpec describes a gather/scatter prefetch hint: one memory
// operand with a VSIB index and an opmask register, no destination. The
// ModRM.reg field carries a fixed /digit, the L'L field is fixed at 512, and
// the mask register is the instruction's only register operand.
type evexPrefGatherSpec struct {
mapSel int
opcode byte
w int
pp int
opdigit int
n int
}
var evexPrefGatherTable = map[string]evexPrefGatherSpec{
"VGATHERPF0DPD": {2, 0xC6, 1, 1, 1, 8},
"VGATHERPF0DPS": {2, 0xC6, 0, 1, 1, 4},
"VGATHERPF0QPD": {2, 0xC7, 1, 1, 1, 8},
"VGATHERPF0QPS": {2, 0xC7, 0, 1, 1, 4},
"VGATHERPF1DPD": {2, 0xC6, 1, 1, 2, 8},
"VGATHERPF1DPS": {2, 0xC6, 0, 1, 2, 4},
"VGATHERPF1QPD": {2, 0xC7, 1, 1, 2, 8},
"VGATHERPF1QPS": {2, 0xC7, 0, 1, 2, 4},
"VSCATTERPF0DPD": {2, 0xC6, 1, 1, 5, 8},
"VSCATTERPF0DPS": {2, 0xC6, 0, 1, 5, 4},
"VSCATTERPF0QPD": {2, 0xC7, 1, 1, 5, 8},
"VSCATTERPF0QPS": {2, 0xC7, 0, 1, 5, 4},
"VSCATTERPF1DPD": {2, 0xC6, 1, 1, 6, 8},
"VSCATTERPF1DPS": {2, 0xC6, 0, 1, 6, 4},
"VSCATTERPF1QPD": {2, 0xC7, 1, 1, 6, 8},
"VSCATTERPF1QPS": {2, 0xC7, 0, 1, 6, 4},
}
// evexHptrSpec describes the high/low half moves (VMOVHPS family): the
// three-operand insert shares an opcode with a two-operand store whose
// source is the vector register and whose destination is m64.
type evexHptrSpec struct {
insert evexSpec
store evexSpec // store.opcode == 0 when the mnemonic has no store form
}
var evexHptrTable = map[string]evexHptrSpec{
"VMOVHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
store: evexSpec{mapSel: 1, opcode: 0x17, w: 0, pp: 0, opdigit: -1, form: vexRMRev, n: [3]int{8, 0, 0}},
},
"VMOVLHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
},
}
// encodeEvexPrefGather encodes a gather/scatter prefetch hint: OP K, vsib.
func (e *enc) encodeEvexPrefGather(upper string, ps evexPrefGatherSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 2 operands (K, vsib memory), got %d", upper, len(ops)+1)
}
m, ok := ops[0].(Mem)
if !ok || !m.HasIndex || !m.Index.isVec() {
return fmt.Errorf("%s: operand must be a VSIB memory reference with a vector index", upper)
}
spec := evexSpec{mapSel: ps.mapSel, opcode: ps.opcode, w: ps.w, pp: ps.pp, opdigit: ps.opdigit, n: [3]int{ps.n, ps.n, ps.n}}
return e.emitEvexFields(spec, 2, ps.opdigit, -1, m, mask, sfx)
}
// vsibLen validates a VSIB memory operand (the index must be a vector
// register) and returns it with the vector length the index selects, the
// EVEX L'L field follows the index register, not the data register.
@@ -1370,7 +1926,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
return err
}
if mask != 0 || sfx.any() {
// EVEX form: OP vsib, K, dst.
// EVEX form: OP vsib, K, dst. The L'L field is the wider of the
// index and the data register lengths (the Go assembler's
// layout); the disp8×N multiplier stays the index element size.
if len(rest) != 2 {
return fmt.Errorf("%s expects 3 operands (vsib, K, dst), got %d", upper, len(ops))
}
@@ -1382,6 +1940,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", upper)
}
if d := dst.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: gs.opcode, w: gs.w, pp: 1, opdigit: -1, n: [3]int{gs.n, gs.n, gs.n}}
return e.emitEvexFields(evex, ll, dst.idx, -1, vsib, mask, sfx)
}
@@ -1431,6 +1992,11 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
if err != nil {
return err
}
// The L'L field is the wider of the data register and the VSIB index
// lengths, the bytes go tool asm emits.
if d := src.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: ss.opcode, w: ss.w, pp: 1, opdigit: -1, n: [3]int{ss.n, ss.n, ss.n}}
return e.emitEvexFields(evex, ll, src.idx, -1, vsib, mask, sfx)
}
@@ -1441,6 +2007,8 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
var evexKOperand = map[string]bool{
"VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true,
"VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true,
// The K-to-vector broadcast reads its opmask source from r/m.
"VPBROADCASTMB2Q": true, "VPBROADCASTMW2D": true,
}
// kmovSpec describes a KMOV width: the opcode depends on the operand
@@ -1541,6 +2109,7 @@ var kOpsTable = map[string]kOpSpec{
"KXORD": {1, 0x47, 1, 1, 1, vexNDS3},
"KXORQ": {1, 0x47, 1, 0, 1, vexNDS3},
"KUNPCKBW": {1, 0x4B, 0, 1, 1, vexNDS3},
"KUNPCKWD": {1, 0x4B, 0, 0, 1, vexNDS3},
"KUNPCKDQ": {1, 0x4B, 1, 0, 1, vexNDS3},
"KADDB": {1, 0x4A, 0, 1, 1, vexNDS3},
"KADDW": {1, 0x4A, 0, 0, 1, vexNDS3},
+213
View File
@@ -721,3 +721,216 @@ func hexCompact(b []byte) string {
}
return string(out)
}
// TestAvx512CorpusFamilies pins representative encodings of the AVX-512
// families the toolchain's avx512enc corpus exercises: the bytes are the
// go tool asm output for exactly these operands, and the same families are
// covered end to end by the avx512_amd64.s differential kernel.
func TestAvx512CorpusFamilies(t *testing.T) {
vsib := func(base, idx string, scale int) Operand {
return Idx(vreg(t, base), vreg(t, idx), scale, 0, 0)
}
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
// AES rounds (EVEX NDS, VEX twin routed by operand width).
{"VAESDEC Z", "VAESDEC", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d48ded9"},
// Integer VNNI and the bit algorithm group.
{"VPDPBUSD", "VPDPBUSD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K2"), vreg(t, "Z3")}, "62f26d4a50d9"},
{"VPOPCNTW", "VPOPCNTW", []Operand{vreg(t, "Z1"), vreg(t, "K3"), vreg(t, "Z2")}, "62f2fd4b54d1"},
{"VPCONFLICTD", "VPCONFLICTD", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d49c4d1"},
{"VPLZCNTQ masked", "VPLZCNTQ", []Operand{vreg(t, "Z7"), vreg(t, "K1"), vreg(t, "Z8")}, "6272fd4944c7"},
{"VPERMT2B", "VPERMT2B", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f26d497dd9"},
{"VPMULTISHIFTQB", "VPMULTISHIFTQB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z4")}, "62f2ed4b83e1"},
{"VDBPSADBW", "VDBPSADBW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z3")}, "62f36d4b42d903"},
{"VPSHUFBITQMB", "VPSHUFBITQMB", []Operand{vreg(t, "Z9"), vreg(t, "Z10"), vreg(t, "K3")}, "62d22d488fd9"},
{"VPTESTNMQ", "VPTESTNMQ", []Operand{vreg(t, "Z13"), vreg(t, "Z14"), vreg(t, "K5")}, "62d28e4827ed"},
// Permutations: immediate and register counts.
{"VALIGNQ", "VALIGNQ", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f3ed4903d903"},
{"VPERMQ imm", "VPERMQ", []Operand{Imm(1), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f3fd4a00d101"},
{"VPERMQ reg", "VPERMQ", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K2"), vreg(t, "Z5")}, "62f2dd4a36eb"},
{"VPERMPD reg", "VPERMPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f2ed4816d9"},
{"VPERMILPS imm", "VPERMILPS", []Operand{Imm(5), vreg(t, "Z9"), vreg(t, "K2"), vreg(t, "Z10")}, "62537d4a04d105"},
{"VPERMILPS reg", "VPERMILPS", []Operand{vreg(t, "Z11"), vreg(t, "Z12"), vreg(t, "K2"), vreg(t, "Z13")}, "62521d4a0ceb"},
// Shifts: immediate, register-count and memory-count forms; the
// count source carries its own XMM tuple width.
{"VPSLLW imm mask", "VPSLLW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f16d4a71f103"},
{"VPSLLD reg count", "VPSLLD", []Operand{vreg(t, "X1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f16d49f2d9"},
{"VPSLLDQ", "VPSLLDQ", []Operand{Imm(9), vreg(t, "Z7"), vreg(t, "Z8")}, "62f13d4873ff09"},
{"VPSRLDQ mem", "VPSRLDQ", []Operand{Imm(11), Ptr(SI, 16, 16), vreg(t, "Z4")}, "62f15d48739e100000000b"},
{"VPSRLVW", "VPSRLVW", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K1"), vreg(t, "Z5")}, "62f2dd4910eb"},
// Conversions and shuffles with the F2 prefix and no prefix.
{"VCVTUDQ2PS", "VCVTUDQ2PS", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f17f497ad1"},
{"VSHUFPS", "VSHUFPS", []Operand{Imm(2), vreg(t, "Z4"), vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f15449c6f402"},
// Gather and scatter prefetch hints (memory-only, /digit in reg).
{"VGATHERPF0DPD", "VGATHERPF0DPD", []Operand{vreg(t, "K5"), vsib("R10", "Y29", 8)}, "6292fd45c60cea"},
{"VSCATTERPF1DPS", "VSCATTERPF1DPS", []Operand{vreg(t, "K2"), vsib("R10", "Z28", 4)}, "62927d42c634a2"},
// Opmask broadcasts and the K logic.
{"VPBROADCASTMB2Q", "VPBROADCASTMB2Q", []Operand{vreg(t, "K1"), vreg(t, "Z2")}, "62f2fe482ad1"},
{"VPBROADCASTMW2D", "VPBROADCASTMW2D", []Operand{vreg(t, "K3"), vreg(t, "Z4")}, "62f27e483ae3"},
{"KUNPCKWD", "KUNPCKWD", []Operand{vreg(t, "K6"), vreg(t, "K4"), vreg(t, "K1")}, "c5dc4bce"},
{"KADDB", "KADDB", []Operand{vreg(t, "K2"), vreg(t, "K3"), vreg(t, "K5")}, "c5e54aea"},
// Lane extracts to general registers (EVEX and VEX routes).
{"VPEXTRB", "VPEXTRB", []Operand{Imm(3), vreg(t, "X26"), AX}, "62637d0814d003"},
{"VPEXTRD", "VPEXTRD", []Operand{Imm(1), vreg(t, "X26"), vreg(t, "R9")}, "62437d0816d101"},
{"VPEXTRD vex", "VPEXTRD", []Operand{Imm(1), vreg(t, "X2"), DI}, "c4e37916d701"},
{"VPINSRQ", "VPINSRQ", []Operand{Imm(1), DI, vreg(t, "X3"), vreg(t, "X4")}, "c4e3e122e701"},
// Moves: masked unaligned, masked scalar register form, half moves
// and non-temporal stores.
{"VMOVUPS mask", "VMOVUPS", []Operand{vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z3")}, "62f17c4a11cb"},
{"VMOVSD 3op", "VMOVSD", []Operand{vreg(t, "X14"), vreg(t, "X5"), vreg(t, "K3"), vreg(t, "X22")}, "6231d70b11f6"},
{"VMOVSS 3op", "VMOVSS", []Operand{vreg(t, "X18"), vreg(t, "X3"), vreg(t, "K2"), vreg(t, "X25")}, "6281660a11d1"},
{"VMOVHPS insert", "VMOVHPS", []Operand{Ptr(SI, 0, 8), vreg(t, "X18"), vreg(t, "X19")}, "62e16c00161e"},
{"VMOVHPS store", "VMOVHPS", []Operand{vreg(t, "X20"), Ptr(SI, 8, 8)}, "62e17c08176601"},
{"VMOVLHPS", "VMOVLHPS", []Operand{vreg(t, "X16"), vreg(t, "X5"), vreg(t, "X17")}, "62a1540816c8"},
{"VMOVNTDQ", "VMOVNTDQ", []Operand{vreg(t, "Z7"), Ptr(SI, 0, 64)}, "62f17d48e73e"},
{"VMOVNTDQA", "VMOVNTDQA", []Operand{Ptr(SI, 64, 64), vreg(t, "Z8")}, "62727d482a4601"},
{"VMOVNTPS", "VMOVNTPS", []Operand{vreg(t, "Z9"), Ptr(SI, 0, 64)}, "62717c482b0e"},
// Scalar compares with and without the 66 prefix.
{"VCOMISD", "VCOMISD", []Operand{vreg(t, "X5"), vreg(t, "X6")}, "c5f92ff5"},
{"VUCOMISS", "VUCOMISS", []Operand{vreg(t, "X7"), vreg(t, "X8")}, "c5782ec7"},
// Floating point helpers.
{"VSQRTSD", "VSQRTSD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "K1"), vreg(t, "X3")}, "62f1ef0951d9"},
{"VEXP2PD", "VEXP2PD", []Operand{vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f2fd49c8f5"},
{"VRCP28SD", "VRCP28SD", []Operand{vreg(t, "X9"), vreg(t, "X8"), vreg(t, "K1"), vreg(t, "X10")}, "6252bd09cbd1"},
{"VBROADCASTF32X2", "VBROADCASTF32X2", []Operand{vreg(t, "X1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d4919d1"},
{"VPCOMPRESSB", "VPCOMPRESSB", []Operand{vreg(t, "Z1"), vreg(t, "K1"), Ptr(SI, 0, 64)}, "62f27d49630e"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
// source keeps r/m, the bracketed list's LOW register travels the inverted
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
// the vector length follows the destination (L'L=512 for the ZMM forms,
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
// member. The x86 decoder has no view of these forms, so no decode check
// runs.
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
sp := vreg(t, "RSP")
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a842411000000"},
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f22f4a9a842411000000"},
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f25f429a842411000000"},
{"V4FMADDPS Z8 dst", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
"62727f4a9a842411000000"},
{"V4FMADDPS disp8x16", "V4FMADDPS",
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a442404"},
{"V4FMADDPS unmasked", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f489a842411000000"},
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9bb007000000"},
{"V4FMADDSS (DI)", "V4FMADDSS",
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9b37"},
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
"62e22f0d9bb007000000"},
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
"62e25f059bb007000000"},
{"V4FMADDSS X30 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
"62627f0d9bb007000000"},
{"V4FMADDSS X3 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
"62f27f0d9b9807000000"},
{"V4FMADDSS disp8x16", "V4FMADDSS",
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
"62625f059b7001"},
{"V4FNMADDPS", "V4FNMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4aaa842411000000"},
{"V4FNMADDSS", "V4FNMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0dabb007000000"},
{"VP4DPWSSD", "VP4DPWSSD",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a52842411000000"},
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f4853842411000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
// the register class the list and the destination take is fixed per
// instruction, the source is memory only, the opmask slot is positional and
// the list's low register owns V'VVVV.
func TestEvexQuadRegisterErrors(t *testing.T) {
sp := vreg(t, "RSP")
list := func(lo, hi string) RegList {
return RegList{vreg(t, lo), vreg(t, hi)}
}
cases := []struct {
name string
mnem string
ops []Operand
}{
{"X list on the PS form", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"Z list on the SS form", "V4FMADDSS",
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
{"Y destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
{"register source", "V4FMADDPS",
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"non-mask third operand", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
{"k0 mask", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
{"K after the destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
{"zeroing without a mask", "V4FMADDPS.Z",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
{"SAE suffix", "V4FMADDPS.SAE",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"high index source", "VP4DPWSSD",
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"short operand list", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
}
for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
+55
View File
@@ -463,6 +463,61 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
symRelocs[si] = append(symRelocs[si], rec[:]...)
}
}
// The data symbols' own relocations: the symbol-valued DATA fields
// ("DATA s+0(SB)/8, $other(SB)"). The toolchain patches each field
// with the target's absolute address through an R_ADDR of the DATA
// line's width, on every architecture (the code relocations are
// per-architecture PC-relative shapes; a data pointer word is not), so
// this mapping bypasses relocField. The definitions were appended in
// DataSyms order, so data symbol i is definition index i.
for i, d := range img.DataSyms {
for _, r := range d.Relocs {
if r.Kind != RelAddr {
return nil, fmt.Errorf("GOOBJ emission: data symbol %q carries a non-data relocation", d.Name)
}
var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = r.Siz
binary.LittleEndian.PutUint16(rec[5:], relocAddr)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
switch {
case r.External && r.Name == goobjBuiltinMorestack:
binary.LittleEndian.PutUint32(rec[15:], pkgIdxBuiltin)
binary.LittleEndian.PutUint32(rec[19:], goobjBuiltinMorestackNoctxt)
case r.External:
pkg, name := splitQualified(r.Name)
if pkg == "" {
return nil, fmt.Errorf("GOOBJ emission: external symbol %q has no package prefix", r.Name)
}
pIdx, ok := extPkgIdx[pkg]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: package %q not resolved", pkg)
}
sIdx, ok := extSymIdx[pkg+"·"+name]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: symbol %s·%s not resolved", pkg, name)
}
binary.LittleEndian.PutUint32(rec[15:], uint32(pIdx))
binary.LittleEndian.PutUint32(rec[19:], uint32(sIdx))
default:
if di, ok := defIdx[r.Name]; ok {
binary.LittleEndian.PutUint32(rec[15:], pkgIdxSelf)
binary.LittleEndian.PutUint32(rec[19:], uint32(di))
break
}
// A DATA field may hold the address of a TEXT function of
// the same file (the rt0 lib entry spelling), which is a
// non-package definition.
ni, isText := textNpIdx[r.Name]
if !isText {
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name)
}
binary.LittleEndian.PutUint32(rec[15:], pkgIdxNone)
binary.LittleEndian.PutUint32(rec[19:], uint32(ni))
}
symRelocs[i] = append(symRelocs[i], rec[:]...)
}
}
// The DWARF symbols' own relocations (the function address references).
for _, ds := range dwarfRelocs {
for _, r := range ds.relocs {
+4 -4
View File
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
func TestStackGuardGOObjInternalCall(t *testing.T) {
for _, tt := range []struct {
src string
assemble func(*ast.File) (*Image, error)
assemble func(*ast.File, ...AssembleOption) (*Image, error)
}{
{"g_amd64.s", AssembleFile},
{"g_arm64.s", AssembleFileARM64},
{"g_riscv64.s", AssembleFileRISCV},
{"g_loong64.s", AssembleFileLOONG64},
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
} {
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
+150 -26
View File
@@ -65,6 +65,15 @@ var bitTestOp = map[string]int{
// noOperandTable maps a fixed no-operand mnemonic to its opcode bytes. The
// fence names carry their opcode inside the 0F AE /digit group spelled out in
// full (E8/F0/F8), and PAUSE is F3 90.
//
// LOCK, REP and REPN are the prefix statements. go tool asm encodes each as
// a standalone one-byte instruction with a PC of its own (F0, F3 and F2
// respectively), not as a prefix field merged into the next instruction: the
// statement that follows is encoded unaware of it, and nothing validates
// that the pairing is a legal one (LOCK before NOP assembles without
// complaint, each byte pinned against the toolchain). Because the bytes
// land in the stream before the following statement anyway, a LOCKed
// CMPXCHGQ encodes identically to a prefixed form.
var noOperandTable = map[string][]byte{
"CPUID": {0x0F, 0xA2},
"RDTSC": {0x0F, 0x31},
@@ -78,6 +87,9 @@ var noOperandTable = map[string][]byte{
"MFENCE": {0x0F, 0xAE, 0xF0},
"SFENCE": {0x0F, 0xAE, 0xF8},
"UNDEF": {0x0F, 0x0B},
"LOCK": {0xF0},
"REP": {0xF3},
"REPN": {0xF2},
}
// --- MOV --------------------------------------------------------------------
@@ -180,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
}
return e.emit(i)
case TLSMem:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
// offset with its R_TLSLE patch site.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case SegAbs:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, 0x30(GS): the segment-absolute load.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case Imm:
if dstIsReg {
v := int64(src)
@@ -220,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
i.imm = imm
return e.emit(i)
}
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
// destination slot is the absolute-address crash-store spelling,
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
// as an immediate, and the store's disp32 carries the address.
op := byte(0xC7)
if size == 1 {
op = 0xC6
}
if d, ok := dst.(Imm); ok {
i := newInstr(size, []byte{op})
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
immBytes, err := immediate(int64(src), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
i := newInstr(size, []byte{op})
if err := setRMDigit(i, 0, dst, size); err != nil {
return err
@@ -498,13 +547,34 @@ func (e *enc) encodeUnary(op struct {
// --- SHL/SHR/SAR ------------------------------------------------------------
func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
// doubleShiftOp maps the two mnemonics whose three-operand form go tool asm
// accepts to the SHLD/SHRD opcode pair (imm8 form, CL form). SAR, SAL and
// the rotates have no such form: the oracle rejects SARQ/ROLQ with three
// operands, and so do we.
var doubleShiftOp = map[string][2]byte{
"SHL": {0xA4, 0xA5}, // SHLD
"SHR": {0xAC, 0xAD}, // SHRD
}
// isShiftCountCL reports whether a count operand is the CL register or its
// CX spelling: go tool asm accepts both (CX names the same low byte) and
// rejects ECX/RCX.
func isShiftCountCL(o Operand) bool {
reg, ok := o.(Reg)
return ok && reg.idx == 1 && (reg.size == 1 || reg.size == 2)
}
func (e *enc) encodeShift(base string, ops []Operand, size int) error {
digit := shiftOp[base]
if len(ops) == 3 {
return e.encodeDoubleShift(base, ops, size)
}
if len(ops) != 2 {
return fmt.Errorf("shift expects 2 operands, got %d", len(ops))
}
count, dst := ops[0], ops[1]
// Count is $1, %CL, or an imm8.
if reg, ok := count.(Reg); ok && reg.idx == 1 && reg.size <= 1 {
// Count is $1, CL (or its CX spelling), or an imm8.
if isShiftCountCL(count) {
// CL: 0xD2 (8-bit) / 0xD3.
op := byte(0xD3)
if size == 1 {
@@ -551,12 +621,59 @@ func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
return e.emit(i)
}
// encodeDoubleShift emits the three-operand SHL/SHR form, which the Go
// assembler spells as a shift but encodes as SHLD/SHRD (0F A4/A5, 0F AC/AD):
// the first operand is the count ($imm or CL), the second feeds the vacated
// bits (the reg field) and the third is the shifted value (the r/m field),
// matching go tool asm byte for byte. The W/L/Q widths exist; the oracle
// rejects the three-operand B form and every SAR/rotate one.
func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
opc, ok := doubleShiftOp[base]
if !ok || size == 1 {
return fmt.Errorf("%s: shift expects 2 operands, got %d", base, len(ops))
}
count, src, dst := ops[0], ops[1], ops[2]
srcReg, ok := src.(Reg)
if !ok {
return fmt.Errorf("%s: middle operand must be a register, like go tool asm", base)
}
i := newInstr(size, []byte{0x0F, opc[0]})
if isShiftCountCL(count) {
// CL (or CX) form: 0F A5/AD.
i.opcode[1] = opc[1]
} else {
imm, ok := count.(Imm)
if !ok {
return fmt.Errorf("shift count must be $1, CL or an immediate")
}
// The count is an unsigned imm8: the same range convention as the
// two-operand shift above.
if imm < 0 || imm > 255 {
return fmt.Errorf("shift count $%d is out of the 0..255 range", int64(imm))
}
i.imm = []byte{byte(imm)}
}
if err := setRMReg(i, srcReg.idx, srcReg.idx >= 8, false, dst, size); err != nil {
return err
}
return e.emit(i)
}
// --- IMUL -------------------------------------------------------------------
func (e *enc) encodeImul(ops []Operand, size int) error {
switch len(ops) {
case 2:
// IMUL r, r/m: 0x0F 0xAF.
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
// Otherwise IMUL r, r/m: 0x0F 0xAF.
if imm, ok := ops[0].(Imm); ok {
dstReg, isReg := ops[1].(Reg)
if !isReg {
return fmt.Errorf("IMUL: destination must be a register")
}
return e.encodeImulImm(imm, dstReg, dstReg, size)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("IMUL: destination must be a register")
@@ -576,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
if !ok {
return fmt.Errorf("IMUL: immediate operand expected first")
}
// Plan 9 order: IMUL $imm, src, dst.
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
// r/m operand (setRM takes registers and memory alike).
return e.encodeImulImm(imm, ops[1], dstReg, size)
}
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
}
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
// when the value fits, 0x69 with a 32-bit immediate otherwise.
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
// --- PUSH / POP -------------------------------------------------------------
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
@@ -986,12 +1110,12 @@ func (e *enc) encodeSSEMove(m sseMove, ops []Operand) error {
op = m.load
reg, rm = dstReg, src
case srcVec:
if _, ok := dst.(Mem); !ok {
if !isX86Mem(dst) {
return fmt.Errorf("SSE move: invalid destination operand")
}
reg, rm = srcReg, dst
case dstVec:
if _, ok := src.(Mem); !ok {
if !isX86Mem(src) {
return fmt.Errorf("SSE move: invalid source operand")
}
op = m.load
+219
View File
@@ -0,0 +1,219 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"encoding/binary"
"os"
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// The differential kernels for the DATA-path and front-end gaps are kept in
// testdata/verify beside the campaign's other kernels; the verify package's
// suites are not open to the asm package, so this test is their runner: each
// kernel assembles through gasm and through go tool asm, and the functions'
// bytes must agree with the relocation sites masked on both sides.
// toolAsmObject assembles path with the installed toolchain's assembler for
// goarch ("" = the host) and returns the object bytes.
func toolAsmObject(t *testing.T, path, goarch string) []byte {
t.Helper()
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
out, err := exec.Command(goBin, "env", "GOROOT").Output()
if err != nil {
t.Fatalf("go env GOROOT: %v", err)
}
includeDir := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
pkg := strings.TrimSuffix(filepath.Base(path), ".s")
pkg = strings.TrimSuffix(pkg, "_amd64")
pkg = strings.TrimSuffix(pkg, "_arm64")
objPath := filepath.Join(t.TempDir(), "oracle.o")
cmd := exec.Command(goBin, "tool", "asm", "-I", includeDir, "-p", pkg, "-o", objPath, path)
if goarch != "" {
environ := os.Environ()
env := make([]string, 0, len(environ)+1)
for _, e := range environ {
if !strings.HasPrefix(e, "GOARCH=") {
env = append(env, e)
}
}
cmd.Env = append(env, "GOARCH="+goarch)
}
if out, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("go tool asm %s: %v\n%s", filepath.Base(path), err, out)
}
obj, err := os.ReadFile(objPath)
if err != nil {
t.Fatal(err)
}
return obj
}
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
// toolchain object, keyed by the name the object records (pkg.name). Each
// function's span is its own symbol size: a toolchain object that follows
// the text with data symbols (the synthesised float-constant pool) would
// otherwise fold them into the last function's bytes.
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
t.Helper()
v := openGoobj(t, obj)
le := binary.LittleEndian
const symSize = 21
nps := v.syms(blkNonpkgdef)
data := v.blk(blkData)
didx := v.blk(blkDataIdx)
preceding := 0
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
preceding += len(v.blk(bi)) / symSize
}
out := make(map[string][]byte, len(nps))
for i, s := range nps {
if s.typ != kindSTEXT {
continue
}
start := le.Uint32(didx[4*(preceding+i):])
out[s.name] = data[start : start+s.size]
}
return out
}
// maskCode zeroes every relocation field, the way the toolchain's object
// leaves them for the linker.
func maskCode(code []byte, relocs []Reloc) []byte {
for _, r := range relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
return code
}
// code assembles src for amd64 and returns the image's code bytes.
func code(path, src string) []byte {
f, errs := parser.Parse(path, src)
if len(errs) > 0 {
return nil
}
img, err := AssembleFile(f)
if err != nil {
return nil
}
return img.Code
}
// TestDifferentialKernels pins the new kernels against the oracle.
func TestDifferentialKernels(t *testing.T) {
if runtime.GOARCH != "amd64" {
t.Skip("the amd64 kernels assume an amd64 host assembler default")
}
for _, k := range []struct {
path string
goarch string
arm64 bool
}{
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
} {
t.Run(filepath.Base(k.path), func(t *testing.T) {
src, err := os.ReadFile(k.path)
if err != nil {
t.Fatalf("read: %v", err)
}
f, errs := parser.Parse(k.path, string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
var img *Image
if k.arm64 {
img, err = AssembleFileARM64(f)
} else {
img, err = AssembleFile(f)
}
if err != nil {
t.Fatalf("assemble: %v", err)
}
gt := oracleFuncCode(t, toolAsmObject(t, k.path, k.goarch))
// The oracle keys its functions by the qualified object name
// (pkg.name); match on the local part.
byLocal := make(map[string][]byte, len(gt))
for name, code := range gt {
if _, after, ok := strings.Cut(name, "."); ok {
name = after
}
byLocal[name] = code
}
matched := 0
for _, fn := range img.Funcs {
gasmCode := maskCode(append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...), fn.Relocs)
goCode, ok := byLocal[fn.Name]
if !ok {
t.Errorf("%s: not in ground truth (%d functions: %v)", fn.Name, len(gt), keysOf(byLocal))
continue
}
goCode = maskCode(append([]byte(nil), goCode...), fn.Relocs)
cmpLen := min(len(goCode), len(gasmCode))
if !bytes.Equal(gasmCode[:cmpLen], goCode[:cmpLen]) {
t.Errorf("%s: MISMATCH gasm=%d go=%d bytes\ngasm %x\ngo %x", fn.Name, len(gasmCode), len(goCode), gasmCode, goCode)
continue
}
for _, b := range goCode[len(gasmCode):] {
if b != 0 {
t.Errorf("%s: non-zero trailing bytes in go tool asm output", fn.Name)
break
}
}
matched++
t.Logf("%s: MATCH (%d bytes)", fn.Name, len(gasmCode))
}
if matched == 0 {
t.Fatal("no functions matched")
}
})
}
}
func keysOf(m map[string][]byte) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
return out
}
// TestSemicolonSpellingParity pins that the ';' statement separator changes
// nothing about the encoding: the one-line spelling assembles to exactly the
// bytes of the same statements written one per line.
func TestSemicolonSpellingParity(t *testing.T) {
for _, tt := range []struct{ one, two string }{
{"\tROLQ $3, DI; ROLQ $13, DI\n", "\tROLQ $3, DI\n\tROLQ $13, DI\n"},
{"\tREP; MOVSQ\n", "\tREP\n\tMOVSQ\n"},
{"\tXORQ AX, AX; XORQ CX, CX\n", "\tXORQ AX, AX\n\tXORQ CX, CX\n"},
} {
one := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.one+"\tRET\n")
two := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.two+"\tRET\n")
if !bytes.Equal(one, two) {
t.Errorf("semicolon spelling %q: %x, want the two-line bytes %x", tt.one, one, two)
}
}
}
+176 -11
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"math"
"sort"
"strconv"
@@ -103,6 +104,7 @@ const (
RelArm64Branch // R_CALLARM64 (BL instruction)
RelArm64LDST64 // R_ARM64_PCREL_LDST64 (ADRP + 64-bit LDR/STR pair)
RelLoong64Branch // R_CALLLOONG64 (BL instruction)
RelAddr // R_ADDR: the absolute address of a symbol held in a DATA field
)
type Reloc struct {
@@ -112,13 +114,17 @@ type Reloc struct {
// Addend select the target: the symbol plus the byte offset. An
// External relocation names a symbol no GLOBL in the file defines;
// the object-file emitters carry it into the output's relocation
// table.
// table. Siz is the width of the patched field and is set only for
// data-field relocations (RelAddr, Off relative to the data symbol),
// whose width is the DATA line's; code relocations take their width
// from the architecture's instruction encoding.
Off int
After int
Name string
Addend int64
External bool
Kind RelocKind
Siz uint8
}
// DataSymbol describes one GLOBL symbol laid out in the data section.
@@ -130,6 +136,11 @@ type DataSymbol struct {
Static bool // the <> marker: file-local, not exported
Rodata bool // the RODATA flag: read-only data
Dupok bool // the DUPOK flag: duplicate-OK
// Relocs carries the symbol-valued DATA initialisers ("DATA s+0(SB)/8,
// $other(SB)"): fields of this symbol's data that hold another symbol's
// address, resolved by the linker. Off is relative to the symbol's
// data start.
Relocs []Reloc
}
// Bytes returns the whole image: code, then data.
@@ -139,6 +150,18 @@ func (img *Image) Bytes() []byte {
return append(out, img.Data...)
}
// AssembleOption adjusts the file-level assembly context.
type AssembleOption func(*linkInfo)
// WithGOOS selects the target operating system for the forms that depend on
// it, the TLS access shape above all: linux and freebsd take the
// one-instruction form, windows and plan9 keep the two-instruction load.
func WithGOOS(goos string) AssembleOption {
return func(l *linkInfo) {
l.goos = goos
}
}
// AssembleFile assembles every TEXT function of a parsed file and lays out
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
// reference to a file-local static symbol becomes a RIP-relative load whose
@@ -146,7 +169,7 @@ func (img *Image) Bytes() []byte {
// GLOBL defines is recorded as an external relocation (Externals) with its
// displacement left zero, the object-file emitters resolve it at link
// time, while the raw image (Bytes) cannot represent it.
func AssembleFile(f *ast.File) (*Image, error) {
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
dataSyms, err := collectData(f)
if err != nil {
return nil, err
@@ -155,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
for _, d := range dataSyms {
known[d.name] = true
}
// TEXT symbols are file-level definitions too: a symbol immediate
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
for _, d := range f.Decls {
if t, ok := d.(*ast.Text); ok {
known[t.Name.Name] = true
}
}
link := &linkInfo{symbols: known, allowExternal: true}
for _, o := range opts {
o(link)
}
poolSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
textOff := map[string]int{}
@@ -169,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, patches, labels, steps, lines, err := assemble(t, link)
code, patches, labels, steps, lines, pool, err := assemble(t, link)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
// The pooled floating-point constants join the declared data as
// read-only symbols, deduplicated across the file (the toolchain
// synthesises the same symbols into its rodata).
for _, entry := range pool {
if poolSeen[entry.name] {
continue
}
poolSeen[entry.name] = true
dataSyms = append(dataSyms, dataSym{
name: entry.name,
buf: entry.data,
size: len(entry.data),
rodata: true,
dupok: true,
})
}
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
@@ -257,6 +310,22 @@ func AssembleFile(f *ast.File) (*Image, error) {
img.Funcs[i].Relocs = append(img.Funcs[i].Relocs, reloc)
}
}
// The data symbols' symbol-valued DATA fields resolve the same way the
// code references do: a name the file defines (GLOBL or TEXT) stays an
// internal reference the emitters resolve, anything else is external.
// img.DataSyms was laid out in dataSyms order, so the indexes line up.
for i := range img.DataSyms {
for _, r := range dataSyms[i].relocs {
reloc := r
if _, ok := img.Symbols[reloc.Name]; !ok {
if _, ok := textOff[reloc.Name]; !ok {
reloc.External = true
externals[reloc.Name] = true
}
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -273,6 +342,10 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if err != nil {
return nil, err
}
// The pooled $i64 constants the wide MOV immediate loads refer to join
// the declared data as read-only symbols, deduplicated across the file
// (the toolchain synthesises the same symbols into its rodata).
litSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls {
@@ -280,10 +353,23 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, labels, relocs, lines, spadj, err := assembleRISCV(t)
code, labels, relocs, lines, spadj, lits, err := assembleRISCV(t)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
for _, lit := range lits {
if litSeen[lit.Name] {
continue
}
litSeen[lit.Name] = true
dataSyms = append(dataSyms, dataSym{
name: lit.Name,
buf: lit.Data,
size: len(lit.Data),
rodata: true,
dupok: true,
})
}
fl := FuncLayout{
Name: t.Name.Name,
Pkg: t.Name.Pkg,
@@ -429,6 +515,23 @@ func markExternals(img *Image, dataSyms []dataSym) {
}
}
}
// The declared data symbols carry the file's own relocations (the
// symbol-valued DATA fields); the layouts appended img.DataSyms in
// dataSyms order, so the indexes line up. The trailing entries (the
// pooled arm64 literals) have no source relocations.
for i := range img.DataSyms {
if i >= len(dataSyms) {
break
}
for _, r := range dataSyms[i].relocs {
reloc := r
if !known[reloc.Name] {
reloc.External = true
externals[reloc.Name] = true
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -444,6 +547,9 @@ type dataSym struct {
static bool
rodata bool
dupok bool
// relocs are the symbol-valued DATA fields, in declaration order; Off
// is relative to the symbol's data start.
relocs []Reloc
}
// collectData gathers the file's static symbols (GLOBL) and their initial
@@ -511,20 +617,79 @@ func collectData(f *ast.File) ([]dataSym, error) {
if !ok {
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
}
if dd.Value == nil || !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
if dd.Value == nil {
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
}
w := dd.Width
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
off := dd.Name.Offset
buf := syms[i].buf
if off < 0 || off+int64(w) > int64(len(buf)) {
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
}
// A symbol value ("DATA s+0(SB)/8, $other(SB)", the rt0 spelling)
// leaves the field zero and records a relocation against the named
// symbol: the linker patches the absolute address at this data
// offset. The toolchain emits the same shape, an R_ADDR of the
// DATA width with the value's offset as the addend, on every
// architecture.
if sym := dd.Value.Imm.Sym; !dd.Value.Imm.HasVal && sym != nil {
syms[i].relocs = append(syms[i].relocs, Reloc{
Off: int(off),
Name: sym.Name,
Addend: sym.Offset,
Kind: RelAddr,
Siz: uint8(w),
})
continue
}
// A string or rune value ("DATA s+0(SB)/20, $"text"") writes its
// bytes into the field and leaves the rest zero, the toolchain's
// WriteString: the declared width must hold every byte, and any
// width is legal.
if s := dd.Value.Imm.Str; s != "" && !dd.Value.Imm.HasVal {
text, err := strconv.Unquote(s)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid string value %s", dd.Name.Name, s)
}
if len(text) > w {
return nil, fmt.Errorf("DATA %q: string of %d bytes does not fit width %d", dd.Name.Name, len(text), w)
}
copy(buf[off:], text)
continue
}
// A floating-point value stores its IEEE-754 bits: /4 the float32
// rounding of the parsed double, /8 the full 64 bits, the
// toolchain's WriteFloat32 and WriteFloat64.
if f := dd.Value.Imm.Float; f != "" && !dd.Value.Imm.HasVal {
num, err := strconv.ParseFloat(f, 64)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid floating-point value %q", dd.Name.Name, f)
}
if dd.Value.Imm.Neg {
num = -num
}
var v uint64
switch w {
case 4:
v = uint64(math.Float32bits(float32(num)))
case 8:
v = math.Float64bits(num)
default:
return nil, fmt.Errorf("DATA %q: invalid width %d for a float (want 4 or 8)", dd.Name.Name, w)
}
for j := range w {
buf[off+int64(j)] = byte(v >> (8 * j))
}
continue
}
if !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
}
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
v := dd.Value.Imm.Val
if dd.Value.Imm.Neg {
v = -v
+325
View File
@@ -4,6 +4,11 @@
package asm
import (
"encoding/binary"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
@@ -166,3 +171,323 @@ func TestCollectDataNumericFlags(t *testing.T) {
}
}
}
// TestCollectDataSymbolValue covers the symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)", the rt0 spelling): the field stays zero in the
// image and the relocation is recorded against the named symbol, whatever
// the file defines (a TEXT function, a GLOBL) or leaves external.
func TestCollectDataSymbolValue(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
MOVQ target+0(FP), AX
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep(SB)
DATA holder+8(SB)/8, $·Keep+5(SB)
DATA holder+16(SB)/8, $holder(SB)
GLOBL spare(SB), NOPTR, $8
DATA spare+0(SB)/8, $extvar(SB)
`
f, errs := parser.Parse("f_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
want := []struct {
sym string
off int
name string
addend int64
ext bool
}{
{"holder", 0, "Keep", 0, false},
{"holder", 8, "Keep", 5, false},
{"holder", 16, "holder", 0, false},
{"spare", 0, "extvar", 0, true},
}
var flat []struct {
sym string
r Reloc
}
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
flat = append(flat, struct {
sym string
r Reloc
}{d.Name, r})
}
}
if len(flat) != len(want) {
t.Fatalf("data relocations = %d, want %d", len(flat), len(want))
}
for i, w := range want {
g := flat[i]
r := g.r
if g.sym != w.sym {
t.Errorf("relocation %d sits on %q, want %q", i, g.sym, w.sym)
continue
}
if r.Off != w.off || r.Name != w.name || r.Addend != w.addend || r.External != w.ext {
t.Errorf("relocation %d = {+%d %q addend %d ext %v}, want {+%d %q addend %d ext %v}",
i, r.Off, r.Name, r.Addend, r.External, w.off, w.name, w.addend, w.ext)
}
if r.Kind != RelAddr {
t.Errorf("relocation %d kind = %v, want RelAddr", i, r.Kind)
}
if r.Siz != 8 {
t.Errorf("relocation %d siz = %d, want 8", i, r.Siz)
}
}
// The fields themselves stay zero: only the linker fills them.
for _, b := range img.Data {
if b != 0 {
t.Fatal("data section is not all zero before relocation")
}
}
if len(img.Externals) != 1 || img.Externals[0] != "extvar" {
t.Errorf("Externals = %v, want [extvar]", img.Externals)
}
}
// TestGOObjectDataSymbolReloc pins the GOOBJ record a symbol-valued DATA
// field produces, against the shape the toolchain emits for the same
// source: an R_ADDR of the DATA width at the field offset, pkgIdxNone plus
// the non-package definition index when the target is the file's own TEXT
// function (the rt0 lib entry spelling).
func TestGOObjectDataSymbolReloc(t *testing.T) {
f, errs := parser.Parse("f_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $16
DATA holder+0(SB)/8, $·Keep+5(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.GOObject("main", "f_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
v := openGoobj(t, obj)
// Walk every relocation record; the data record is the one of Siz 8
// and type R_ADDR.
var off, add int64
var pkg, sym uint32
found := false
for data := v.blk(blkReloc); len(data) >= 23; data = data[23:] {
if data[4] != 8 || binary.LittleEndian.Uint16(data[5:]) != relocAddr {
continue
}
found = true
off = int64(int32(binary.LittleEndian.Uint32(data[0:])))
add = int64(binary.LittleEndian.Uint64(data[7:]))
pkg = binary.LittleEndian.Uint32(data[15:])
sym = binary.LittleEndian.Uint32(data[19:])
break
}
if !found {
t.Fatal("no data relocation record in the object")
}
if off != 0 || add != 5 {
t.Errorf("data reloc = {off %d addend %d}, want {off 0 addend 5}", off, add)
}
if pkg != pkgIdxNone {
t.Errorf("data reloc pkg = %#x, want pkgIdxNone (the TEXT function)", pkg)
}
// The function's non-package definition index: the four pc tables
// precede it, so index 4.
if sym != 4 {
t.Errorf("data reloc sym = %d, want 4", sym)
}
}
// TestGOObjectDataSymbolLink is the end-to-end proof for symbol-valued DATA
// fields: the gasm object is substituted for the toolchain's and re-linked,
// then executed, and the linked data word must hold the real address of the
// function the DATA line named (runtime.FuncForPC identifies it).
func TestGOObjectDataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVQ entry+0(SB), AX
MOVQ AX, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_amd64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
import "runtime"
func keepme()
func entryptr() uintptr
func main() {
pc := entryptr()
fn := runtime.FuncForPC(pc)
if fn == nil {
panic("the entry word does not point at a function")
}
if fn.Name() != "main.keepme" {
panic("the entry word points at " + fn.Name())
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
// Capture the build: the package archive's asm object and the link line.
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_amd64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
// Assemble the same source with gasm and substitute the object.
src, err := os.ReadFile(filepath.Join(dir, "main_amd64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_amd64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
gasmObj, err := img.GOObject("dlink", "main_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
// The linked program must run and find the right function behind the
// data word.
out, err := exec.Command(filepath.Join(dir, "prog")).CombinedOutput()
if err != nil {
t.Fatalf("linked program failed: %v\n%s", err, out)
}
}
// TestCollectDataFloatAndStringValues covers the non-integer DATA values the
// runtime's math and asm files use: floating-point initialisers store their
// IEEE-754 bits (/4 the float32 rounding, /8 the full double) and string
// initialisers write their bytes zero-padded within the declared width.
func TestCollectDataFloatAndStringValues(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL vals<>(SB), RODATA, $44
DATA vals<>+0(SB)/8, $0.5
DATA vals<>+8(SB)/8, $-1.0
DATA vals<>+16(SB)/4, $1.5
DATA vals<>+20(SB)/16, $"call frame too "
DATA vals<>+36(SB)/4, $"hi"
`
f, errs := parser.Parse("fvals_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
d := byName["vals"]
if d.Size != 44 {
t.Fatalf("vals size = %d, want 44", d.Size)
}
buf := img.Data[d.Offset : d.Offset+44]
// 0.5 = 0x3FE0000000000000, -1.0 = 0xBFF0000000000000 (float64);
// 1.5 = 0x3FC00000 (float32).
for _, c := range []struct {
off int
want []byte
}{
{0, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xE0, 0x3F}},
{8, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xF0, 0xBF}},
{16, []byte{0x00, 0x00, 0xC0, 0x3F}},
{20, []byte("call frame too ")},
{36, []byte{'h', 'i', 0x00, 0x00}},
} {
if string(buf[c.off:c.off+len(c.want)]) != string(c.want) {
t.Errorf("vals+%d: got % x, want % x", c.off, buf[c.off:c.off+len(c.want)], c.want)
}
}
}
// TestCollectDataValueErrors pins the value-kind width rules: a float needs
// width 4 or 8, a string must fit its declared width, and a bad float
// literal is diagnosed rather than stored.
func TestCollectDataValueErrors(t *testing.T) {
cases := []string{
`GLOBL v<>(SB), RODATA, $4
DATA v<>+0(SB)/1, $0.5`,
`GLOBL v<>(SB), RODATA, $2
DATA v<>+0(SB)/2, $"toolarge"`,
}
for i, src := range cases {
full := "#include \"textflag.h\"\nTEXT ·Keep(SB), NOSPLIT, $0-8\n\tRET\n" + src
f, errs := parser.Parse(fmt.Sprintf("verr%d_amd64.s", i), full)
if len(errs) > 0 {
t.Fatalf("case %d parse: %v", i, errs)
}
if _, err := AssembleFile(f); err == nil {
t.Errorf("case %d: expected an error, got none", i)
}
}
}
+450 -45
View File
@@ -46,30 +46,70 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
spadj = append(spadj, SpadjStep{PC: guardLen + (loong64StoreWords(fi.autosize)+loong64AdjustWords(-int64(fi.autosize)))*4, Value: fi.autosize})
}
// Pass 1: label offsets from the instruction sizes.
offsets := map[string]int{}
pos := guardLen + len(prologue)
// The toolchain's parser counts N(PC) displacements over the source
// instructions at a uniform 4 bytes each, so a PC-relative branch
// resolves to the instruction N slots away in body order; the resolved
// target then participates in layout and loop-head padding like any
// branch target.
instrs := make([]*ast.Instr, 0, len(t.Body))
for _, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
offsets[s.Name.Text] = pos
case *ast.Instr:
pos += loong64InstrSize(s, fi)
if in, ok := stmt.(*ast.Instr); ok && strings.ToUpper(in.Mnemonic.Text) != "PCALIGN" {
instrs = append(instrs, in)
}
}
// Pass 2: encode. The guard prefix precedes the prologue; its branches
// target the morestack block at the end of the function, which the first
// pass has sized.
bodyLen := 0
{
p := guardLen + len(prologue)
for _, stmt := range t.Body {
if in, ok := stmt.(*ast.Instr); ok {
p += loong64InstrSize(in, fi)
}
parseIndex := make(map[*ast.Instr]int, len(instrs))
for i, in := range instrs {
parseIndex[in] = i
}
pcRelTarget := make(map[*ast.Instr]*ast.Instr)
for _, in := range instrs {
off, ok := loong64PCRelOffset(in)
if !ok {
continue
}
bodyLen = p - (guardLen + len(prologue))
tgt := parseIndex[in] + off
if tgt < 0 || tgt >= len(instrs) {
continue
}
pcRelTarget[in] = instrs[tgt]
}
// Pass 1: label offsets from the instruction sizes. PCALIGN contributes
// only its padding. On top of the explicit PCALIGNs, the toolchain pads
// every backward-branch target (loop head) to a 16-byte boundary, so the
// layout runs to a fixpoint over the alignment set.
loopAligns := map[string]bool{}
alignInstrs := map[*ast.Instr]bool{}
for {
offsets, _, pcs, _ := loong64Layout(t, guardLen+len(prologue), fi, loopAligns, alignInstrs)
changed := false
for _, in := range instrs {
// A backward PC-relative target is the resolved instruction.
if tgt, ok := pcRelTarget[in]; ok && pcs[tgt] < pcs[in] && !alignInstrs[tgt] {
alignInstrs[tgt] = true
changed = true
}
target, ok := loong64BranchTarget(in)
if !ok {
continue
}
tOff, ok := offsets[target]
if !ok || tOff >= pcs[in] || loopAligns[target] {
continue
}
loopAligns[target] = true
changed = true
}
if !changed {
break
}
}
// Final layout with the complete alignment set.
offsets, alignPad, pcs, bodyEnd := loong64Layout(t, guardLen+len(prologue), fi, loopAligns, alignInstrs)
bodyLen := bodyEnd - (guardLen + len(prologue))
pcRelPcs := make(map[*ast.Instr]int, len(pcRelTarget))
for in, tgt := range pcRelTarget {
pcRelPcs[in] = pcs[tgt]
}
var out []byte
if fi.needSplit {
@@ -84,7 +124,20 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
if !ok {
continue
}
code, err := encodeLOONG64Instr(in, pc, offsets, fi, &relocs, resolve)
// PCALIGN pads to the requested boundary with andi $0, $0, 0, the
// architecture's NOP, and encodes to nothing itself.
if strings.ToUpper(in.Mnemonic.Text) == "PCALIGN" {
pad := loong64PCAlignPad(pc, in)
out = append(out, loong64PadBytes(pad)...)
pc += pad
continue
}
// Loop-head alignment padding precedes the instruction.
if pad := alignPad[in]; pad > 0 {
out = append(out, loong64PadBytes(pad)...)
pc += pad
}
code, err := encodeLOONG64Instr(in, pc, offsets, fi, &relocs, resolve, pcRelPcs)
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err)
}
@@ -115,6 +168,115 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
return out, offsets, relocs, lines, spadj, nil
}
// loong64PCRelOffset reports the N of a branch operand spelled N(PC): the
// displacement counted in source instructions from the branch itself.
func loong64PCRelOffset(instr *ast.Instr) (int, bool) {
mnem := strings.ToUpper(instr.Mnemonic.Text)
branch := false
switch mnem {
case "JMP":
branch = len(instr.Operands) == 1
case "JAL", "CALL", "BL":
branch = len(instr.Operands) == 1 || len(instr.Operands) == 2
case "BFPT", "BFPF":
branch = len(instr.Operands) == 1
case "BEQ", "BNE", "BLT", "BGE", "BLTU", "BGEU",
"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ":
branch = len(instr.Operands) >= 2
}
if !branch {
return 0, false
}
op := instr.Operands[len(instr.Operands)-1]
if op.Kind == ast.OpAddr && op.Addr.Sym == nil && op.Addr.Base == "PC" {
return int(op.Addr.Offset), true
}
return 0, false
}
// loong64Layout walks the function body once and returns the label offsets,
// the loop-alignment padding due before each instruction (a pad of 0 needs
// nothing), the pc each instruction starts at (its padding included) and the
// first pc past the body. Explicit PCALIGN pads, the alignment pads for the
// labels in aligns and those for the instructions in alignInstrs (backward
// PC-relative targets) all contribute, mirroring the toolchain's layout
// pass.
func loong64Layout(t *ast.Text, start int, fi loong64FrameInfo, aligns map[string]bool, alignInstrs map[*ast.Instr]bool) (map[string]int, map[*ast.Instr]int, map[*ast.Instr]int, int) {
offsets := map[string]int{}
alignPad := map[*ast.Instr]int{}
pcs := map[*ast.Instr]int{}
pos := start
pendingAlign := false
var pendingNames []string
explicit := false
for _, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
if aligns[s.Name.Text] {
pendingAlign = true
}
pendingNames = append(pendingNames, s.Name.Text)
// Provisional: a branch to the label lands here unless a loop
// alignment pad follows, in which case the label resolves to the
// padded instruction (the toolchain's labels bind to the branch
// target instruction, which the padding pass precedes).
offsets[s.Name.Text] = pos
case *ast.Instr:
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
pos += loong64PCAlignPad(pos, s)
explicit = true
continue
}
if pendingAlign {
pendingAlign = false
if pos&15 != 0 {
alignPad[s] = 16 - pos&15
}
}
if alignInstrs[s] && pos&15 != 0 {
alignPad[s] = 16 - pos&15
}
if !explicit {
for _, n := range pendingNames {
offsets[n] = pos + alignPad[s]
}
}
pendingNames = nil
explicit = false
pcs[s] = pos + alignPad[s]
pos += alignPad[s] + loong64InstrSize(s, fi)
}
}
return offsets, alignPad, pcs, pos
}
// loong64BranchTarget reports the local label a branch-like instruction
// transfers to, the loop-head signal the toolchain derives from backward
// branch targets.
func loong64BranchTarget(instr *ast.Instr) (string, bool) {
mnem := strings.ToUpper(instr.Mnemonic.Text)
ops := instr.Operands
var op *ast.Operand
switch {
case mnem == "JMP" || mnem == "JAL" || mnem == "BFPT" || mnem == "BFPF":
if len(ops) != 1 {
return "", false
}
op = ops[0]
case mnem == "TEQ" || mnem == "TNE":
return "", false
case len(ops) >= 2:
op = ops[len(ops)-1]
default:
return "", false
}
if op.Kind == ast.OpAddr && op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" &&
op.Addr.Base == "" && op.Addr.Sym.Name != "" {
return op.Addr.Sym.Name, true
}
return "", false
}
// loong64JumpChain precomputes jump-to-jump folding, mirroring the linker's
// branch-chasing pass: a label whose first instruction is an unconditional
// local jump redirects its own jumpers to the ultimate target. The Go
@@ -177,21 +339,55 @@ func l64LabelOK(op *ast.Operand) (string, bool) {
return "", false
}
// l64SubToAdd rewrites the SUB family with an immediate first operand onto
// its ADD counterpart with the negated immediate: LoongArch has no
// subtract-immediate instructions, and the toolchain folds SUB $v into the
// ADD immediate form through the same optab matching (the $0 fold into 3R
// and the large-constant materialisations included). The negation is the
// second result; the operand is left untouched because the size pass
// normalises the same instruction.
func l64SubToAdd(mnem string, ops []*ast.Operand) (string, bool) {
if len(ops) >= 2 && isImmOperand(ops[0]) {
switch mnem {
case "SUB":
return "ADD", true
case "SUBW":
return "ADDW", true
case "SUBV", "SUBVU":
return "ADDV", true
}
}
return mnem, false
}
// loong64InstrSize returns the encoded size of an instruction: 4 bytes for
// most, more for the multi-instruction expansions.
func loong64InstrSize(instr *ast.Instr, fi loong64FrameInfo) int {
mnem := strings.ToUpper(instr.Mnemonic.Text)
ops := instr.Operands
var neg bool
mnem, neg = l64SubToAdd(mnem, ops)
if mnem == "RET" {
return len(loong64Return(fi))
}
switch mnem {
case "END", "FUNCDATA", "PCDATA":
return 0 // bookkeeping statements contribute no bytes
case "GETCALLERPC":
return 4 // or rd, r1, r0
case "TEQ", "TNE":
return 8 // bne/beq over the BREAK, then BREAK
case "PRELDX":
return 20 // the four-instruction constant materialisation + preldx
case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD":
return loong64MovSize(mnem, ops, fi)
case "ADD", "ADDW", "ADDV", "ADDVU", "AND", "OR", "XOR", "SGT", "SGTU":
if len(ops) >= 2 && isImmOperand(ops[0]) {
v := l64Imm64(ops[0])
if neg {
v = -v
}
if v == 0 {
return 4 // folds into the 3R form (rk = R0)
}
@@ -229,11 +425,50 @@ func loong64InstrSize(instr *ast.Instr, fi loong64FrameInfo) int {
return 4
}
// loong64PCAlignPad returns the padding PCALIGN inserts before the next
// instruction so that it starts at the requested boundary relative to the
// function start. The boundary must be a power of two between 8 and 2048, as
// the toolchain requires; anything else pads nothing.
func loong64PCAlignPad(pos int, instr *ast.Instr) int {
if len(instr.Operands) != 1 || !isImmOperand(instr.Operands[0]) {
return 0
}
align := int(immFromOperand(instr.Operands[0]))
if align < 8 || align > 2048 || align&(align-1) != 0 {
return 0
}
return (align - pos%align) % align
}
// loong64PadBytes renders PCALIGN padding: the toolchain emits andi $0, $0, 0
// (the architecture's NOP) for every full 4 bytes of pad.
func loong64PadBytes(pad int) []byte {
nop := l64wordLE(l64irr(l64DualTable["AND"].imm, 0, 0, 0))
out := make([]byte, 0, pad/4*len(nop))
for i := 0; i < pad/4; i++ {
out = append(out, nop...)
}
return out
}
// encodeLOONG64Instr encodes a single LoongArch instruction.
func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loong64FrameInfo, relocs *[]Reloc, resolve func(string) string) ([]byte, error) {
func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loong64FrameInfo, relocs *[]Reloc, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
mnem := strings.ToUpper(instr.Mnemonic.Text)
ops := instr.Operands
// The SUB family with an immediate first operand folds onto the ADD
// immediate form with the negated immediate; the negation happens on a
// copy of the operand, never on the shared syntax tree.
mnem, neg := l64SubToAdd(mnem, ops)
if neg {
c := *ops[0]
c.Imm.Val = -c.Imm.Val
ops2 := make([]*ast.Operand, len(ops))
ops2[0] = &c
copy(ops2[1:], ops[1:])
ops = ops2
}
// Pseudo-instructions and the branches first.
switch mnem {
case "RET":
@@ -249,10 +484,110 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
return nil, fmt.Errorf("WORD expects 1 operand, got %d", len(ops))
}
return l64wordLE(uint32(immFromOperand(ops[0]))), nil
case "END", "FUNCDATA", "PCDATA", "GETCALLERPC":
// The assembler's bookkeeping statements. END, FUNCDATA and PCDATA
// contribute no bytes, the same shapes GOARCH=loong64 go tool asm
// accepts and emits nothing for; GETCALLERPC reads the caller's
// address out of R1 (RA) as or rd, r1, r0.
switch mnem {
case "END":
if len(ops) != 0 {
return nil, fmt.Errorf("END expects no operands, got %d", len(ops))
}
return nil, nil
case "FUNCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) {
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
}
return nil, nil
case "PCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
return nil, fmt.Errorf("PCDATA expects $n, $n")
}
return nil, nil
}
if len(ops) != 1 || isMemOperand(ops[0]) || loong64RegClass(operandRegName(ops[0])) != l64ClsGR {
return nil, fmt.Errorf("GETCALLERPC expects a general register")
}
rd := loong64RegNum(operandRegName(ops[0]))
if rd < 0 {
return nil, fmt.Errorf("GETCALLERPC: invalid register operand")
}
return l64wordLE(l64rrr(l64movRegTable["MOVV"].op, 0, 1, rd)), nil
case "NEGW", "NEGV":
// The integer negation pseudo is a subtract from zero:
// NEGW src, dst → sub.w r0, src, dst.
if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
}
src, dst := l64Reg(ops[0]), l64Reg(ops[1])
if src < 0 || dst < 0 {
return nil, fmt.Errorf("invalid register operand")
}
sub := l64InstrTable["SUBW"].op
if mnem == "NEGV" {
sub = l64InstrTable["SUBV"].op
}
return l64wordLE(l64rrr(sub, src, 0, dst)), nil
case "TEQ", "TNE":
// The trap pseudo expands to two instructions: bne/beq rj, rd over
// the BREAK (offset 2 instruction units), then BREAK $code.
if len(ops) != 2 && len(ops) != 3 {
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
}
code := int(immFromOperand(ops[0]))
rj, rd := 0, l64Reg(ops[len(ops)-1])
if len(ops) == 3 {
rj = l64Reg(ops[1])
}
if rj < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand")
}
bop := l64branchTable["BNE"]
if mnem == "TNE" {
bop = l64branchTable["BEQ"]
}
return l64WordsLE(
l64irr16(bop, 2, rj, rd),
l64i15(l64InstrTable["BREAK"].op, code),
), nil
case "PRELDX":
// preldx offset(Rbase), $n, $hint: the 64-bit descriptor n packs
// (addrSeq, blockSize, blockNums, stride); the constant v built from
// it materialises in R30 across four instructions, then the preldx.
if len(ops) != 3 || !isMemOperand(ops[0]) || !isImmOperand(ops[1]) || !isImmOperand(ops[2]) {
return nil, fmt.Errorf("PRELDX expects offset(reg), $n, $hint")
}
rj := loong64RegNum(ops[0].Addr.Base)
if rj < 0 {
return nil, fmt.Errorf("invalid register operand")
}
n := uint64(l64Imm64(ops[1]))
hint := int(l64Imm64(ops[2]))
addrSeq := (n >> 0) & 0x1
blkSize := (n >> 1) & 0x7ff
blkNums := (n >> 12) & 0x1ff
stride := (n >> 21) & 0xffff
v := uint64(ops[0].Addr.Offset)&0xffff + addrSeq<<16 +
((blkSize/16)-1)<<20 + (blkNums-1)<<32 + stride<<44
const (
lu12iw = 0x0a << 25
lu32id = 0x0b << 25
lu52id = 0x00c << 22
ori = 0x00e << 22
preldx = 0x7058 << 15
)
return l64WordsLE(
l64ir(lu12iw, int(uint32(v>>12)), 30),
l64irr(ori, int(uint32(v)), 30, 30),
l64ir(lu32id, int(uint32(v>>32)), 30),
l64irr(lu52id, int(uint32(v>>52)), 30, 30),
l64rrr(preldx, 30, rj, hint),
), nil
case "JMP", "B":
return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve, relocs)
return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve, relocs, pcRelPcs)
case "JAL", "CALL", "BL":
return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve, relocs)
return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve, relocs, pcRelPcs)
case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD":
return encodeLOONG64Mov(instr, mnem, fi, relocs)
}
@@ -262,12 +597,12 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
if mnem == "JIRL" {
return encodeLOONG64Jirl(op, ops)
}
return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve)
return encodeLOONG64Branch16(instr, mnem, op, ops, pc, offsets, resolve, pcRelPcs)
}
// Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ,
// BFPT/BFPF; BEQZ/BNEZ are reached through BEQ/BNE with R0).
if op, ok := l64branch21Table[mnem]; ok {
return encodeLOONG64Branch21(mnem, op, ops, pc, offsets, resolve)
return encodeLOONG64Branch21(instr, mnem, op, ops, pc, offsets, resolve, pcRelPcs)
}
// B/BL aliases reached only via JMP/JAL above.
@@ -533,11 +868,23 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
//
// JMP/B label → b label JMP/B (rj) → jirl r0, rj, 0
// JAL/CALL/BL label → bl label JAL/CALL/BL (rj) → jirl r1, rj, 0
func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string, relocs *[]Reloc) ([]byte, error) {
func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
if len(instr.Operands) != 1 {
return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(instr.Operands))
}
op := instr.Operands[0]
// PC-relative displacement: N(PC) resolves to the instruction N slots
// away in source order (the toolchain's parse-time count), and the field
// carries the final pc distance in instruction units.
if op.Addr.Sym == nil && op.Addr.Base == "PC" {
targetPc, ok := pcRelPcs[instr]
if !ok {
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, op.Addr.Offset)
}
v := (targetPc - pc) >> 2
opc := l64jumpTable[mnem]
return l64wordLE(l64bbl(opc, v)), nil
}
if isMemOperand(op) && op.Addr.Base != "" && op.Addr.Index == "" && op.Addr.Sym == nil {
// Indirect: (rj) → jirl.
rj := loong64RegNum(op.Addr.Base)
@@ -620,16 +967,28 @@ func l64offsetOperand(op *ast.Operand) (int32, bool) {
// encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU):
// INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain
// turns into the 21-bit BEQZ/BNEZ form when the register is the only operand.
func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
func encodeLOONG64Branch16(instr *ast.Instr, mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
if len(ops) != 2 && len(ops) != 3 {
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
}
target := resolve(l64Label(ops[len(ops)-1]))
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
var target string
var v int
lastOp := ops[len(ops)-1]
if lastOp.Kind == ast.OpAddr && lastOp.Addr.Sym == nil && lastOp.Addr.Base == "PC" {
// N(PC) resolves to the instruction N slots away in source order.
targetPc, ok := pcRelPcs[instr]
if !ok {
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, lastOp.Addr.Offset)
}
v = (targetPc - pc) >> 2
} else {
target = resolve(l64Label(lastOp))
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
}
v = (targetOff - pc) >> 2
}
v := (targetOff - pc) >> 2
if len(ops) == 2 {
// Single register: BEQ rj, label → beqz (21-bit), and the BLTZ/
// BGEZ-family aliases encoded with rj in the rj field.
@@ -690,33 +1049,55 @@ func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, o
// BFPT/BFPF use the 21-bit offset form (register in the rj field), while
// BGTZ/BLEZ, which the toolchain encodes with the register in the rd field
// and a 16-bit offset, are handled separately.
func encodeLOONG64Branch21(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
if len(ops) != 2 {
func encodeLOONG64Branch21(instr *ast.Instr, mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
isBF := mnem == "BFPT" || mnem == "BFPF"
if len(ops) != 2 && !(isBF && (len(ops) == 1 || len(ops) == 2)) {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
}
target := resolve(l64Label(ops[1]))
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
}
v := (targetOff - pc) >> 2
rj := 0 // BFPT/BFPF default to FCC0
if mnem != "BFPT" && mnem != "BFPF" {
var rj int
tgtOp := ops[len(ops)-1]
if isBF {
// BFPT/BFPF test an FCC condition register, defaulting to FCC0 when
// spelled without one.
rj = 0
if len(ops) == 2 {
rj = l64Reg(ops[0])
if rj < 0 {
return nil, fmt.Errorf("invalid register operand")
}
}
} else {
rj = l64Reg(ops[0])
if rj < 0 {
return nil, fmt.Errorf("invalid register operand")
}
}
var v int
if tgtOp.Kind == ast.OpAddr && tgtOp.Addr.Sym == nil && tgtOp.Addr.Base == "PC" {
// N(PC) resolves to the instruction N slots away in source order.
targetPc, ok := pcRelPcs[instr]
if !ok {
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, tgtOp.Addr.Offset)
}
v = (targetPc - pc) >> 2
} else {
target := resolve(l64Label(tgtOp))
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
}
v = (targetOff - pc) >> 2
}
if mnem == "BGTZ" || mnem == "BLEZ" {
// The toolchain swaps the register into the rd field and keeps the
// 16-bit offset form.
if (v<<16)>>16 != v {
return nil, fmt.Errorf("branch to %q too far (16-bit range)", target)
return nil, fmt.Errorf("branch %d too far (16-bit range)", v)
}
return l64wordLE(l64irr16(op, v, 0, rj)), nil
}
if (v<<11)>>11 != v {
return nil, fmt.Errorf("branch to %q too far (21-bit range)", target)
return nil, fmt.Errorf("branch %d too far (21-bit range)", v)
}
return l64wordLE(l64ir21(op, v, rj)), nil
}
@@ -1737,6 +2118,30 @@ func encodeLOONG64Vector(instr *ast.Instr, mnem string, fi loong64FrameInfo) ([]
return l64wordLE(l64rr(l64InstrTable[mnem].op, vj, fcc)), true, nil
}
// Four-register forms (vshuf.b): INSTR va, vk, vj, vd.
if l64Vec4R[mnem] {
if len(ops) != 4 {
return nil, true, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
}
va, err := vec(ops[0])
if err != nil {
return nil, true, err
}
vk, err := vec(ops[1])
if err != nil {
return nil, true, err
}
vj, err := vec(ops[2])
if err != nil {
return nil, true, err
}
vd, err := vec(ops[3])
if err != nil {
return nil, true, err
}
return l64wordLE(l64InstrTable[mnem].op | uint32(va&0x1f)<<15 | uint32(vk&0x1f)<<10 | uint32(vj&0x1f)<<5 | uint32(vd&0x1f)), true, nil
}
// Three-register forms: INSTR vk, vj, vd or INSTR vk, vd (vj = vd).
if len(ops) != 2 && len(ops) != 3 {
return nil, true, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
+449 -12
View File
@@ -278,6 +278,7 @@ const (
l64Fpreld // preld (2RI12 + 5-bit hint)
l64Fvvv // 3R vector (LSX/LASX): op | vk<<10 | vj<<5 | vd
l64Fvcf // vector-to-condition: op | subop<<10 | vj<<5 | fcc
l64Fvvvv // 4R vector shuffle: op | va<<15 | vk<<10 | vj<<5 | vd
)
// l64Enc is one instruction's encoding: its bit layout (format) and the
@@ -336,6 +337,10 @@ var l64VecImmInfo = map[string]l64VecImmEnc{}
// vpcnt.v).
var l64Vec2R = map[string]bool{}
// l64Vec4R marks the four-operand vector mnemonics (INSTR va, vk, vj, vd,
// such as vshuf.b).
var l64Vec4R = map[string]bool{}
// l64VmovqOps holds the VMOVQ/XVMOVQ opcode constants (pre-shifted to bit
// 15), read off `go tool objdump` of GOARCH=loong64 `go tool asm` kernels.
type l64VmovqEnc struct {
@@ -445,6 +450,14 @@ func init() {
// bank is the FP registers (the toolchain spells it `FFINTDV F0, F1`),
// so the entry stays on the 2R integer/FP format.
"FFINTDV": 0x474a << 10,
// The rest of the scalar conversions (all F-bank, 2R).
"FFINTFW": 0x4744 << 10, // ffint.s.w
"FFINTFV": 0x4746 << 10, // ffint.s.l
"FFINTDW": 0x4748 << 10, // ffint.d.w
"FTINTWF": 0x46c1 << 10, // ftint.w.s
"FTINTWD": 0x46c2 << 10, // ftint.w.d
"FTINTVF": 0x46c9 << 10, // ftint.l.s
"FTINTVD": 0x46ca << 10, // ftint.l.d
}
for m, op := range rr {
l64InstrTable[m] = l64Enc{format: l64Frr, op: op}
@@ -568,6 +581,12 @@ func init() {
"AMADDDBW": 0x070D4 << 15, "AMADDDBV": 0x070D5 << 15,
"AMANDDBW": 0x070D6 << 15, "AMANDDBV": 0x070D7 << 15,
"AMORDBW": 0x070D8 << 15, "AMORDBV": 0x070D9 << 15,
// The remaining _dbar exchange variants (loong64enc1.s).
"AMXORDBW": 0x070DA << 15, "AMXORDBV": 0x070DB << 15,
"AMMAXDBW": 0x070DC << 15, "AMMAXDBV": 0x070DD << 15,
"AMMINDBW": 0x070DE << 15, "AMMINDBV": 0x070DF << 15,
"AMMAXDBWU": 0x070E0 << 15, "AMMAXDBVU": 0x070E1 << 15,
"AMMINDBWU": 0x070E2 << 15, "AMMINDBVU": 0x070E3 << 15,
}
for m, op := range am {
l64InstrTable[m] = l64Enc{format: l64Fam, op: op}
@@ -590,6 +609,240 @@ func init() {
"XVANDV": {0xEA4C << 15, true}, "XVXORV": {0xEA4E << 15, true},
"XVSEQB": {0xE800 << 15, true}, "XVSEQV": {0xE803 << 15, true},
}
// The integer and FP add/subtract families: [X]VADD and [X]VSUB by lane
// width, plus the [X]VSADD/[X]VSSUB saturating pairs.
// Opcodes transcribed from the toolchain's loong64enc1.s.
addsub := map[string]l64Vec3Enc{
"VADDB": {0xE014 << 15, false}, "VADDH": {0xE015 << 15, false},
"VADDD": {0xE262 << 15, false}, "VADDF": {0xE261 << 15, false},
"VADDQ": {0xE25A << 15, false},
"VSUBB": {0xE018 << 15, false}, "VSUBH": {0xE019 << 15, false},
"VSUBW": {0xE01A << 15, false}, "VSUBV": {0xE01B << 15, false},
"VSUBQ": {0xE25B << 15, false},
"VSUBF": {0xE265 << 15, false}, "VSUBD": {0xE266 << 15, false},
"VSADDB": {0xE08C << 15, false}, "VSADDH": {0xE08D << 15, false},
"VSADDW": {0xE08E << 15, false}, "VSADDV": {0xE08F << 15, false},
"VSADDBU": {0xE094 << 15, false}, "VSADDHU": {0xE095 << 15, false},
"VSADDWU": {0xE096 << 15, false}, "VSADDVU": {0xE097 << 15, false},
"VSSUBB": {0xE090 << 15, false}, "VSSUBH": {0xE091 << 15, false},
"VSSUBW": {0xE092 << 15, false}, "VSSUBV": {0xE093 << 15, false},
"VSSUBBU": {0xE098 << 15, false}, "VSSUBHU": {0xE099 << 15, false},
"VSSUBWU": {0xE09A << 15, false}, "VSSUBVU": {0xE09B << 15, false},
"XVADDB": {0xE814 << 15, true}, "XVADDH": {0xE815 << 15, true},
"XVADDW": {0xE816 << 15, true},
"XVADDD": {0xEA62 << 15, true}, "XVADDF": {0xEA61 << 15, true},
"XVADDQ": {0xEA5A << 15, true},
"XVSUBB": {0xE818 << 15, true}, "XVSUBH": {0xE819 << 15, true},
"XVSUBW": {0xE81A << 15, true}, "XVSUBV": {0xE81B << 15, true},
"XVSUBQ": {0xEA5B << 15, true},
"XVSUBF": {0xEA65 << 15, true}, "XVSUBD": {0xEA66 << 15, true},
"XVSADDB": {0xE88C << 15, true}, "XVSADDH": {0xE88D << 15, true},
"XVSADDW": {0xE88E << 15, true}, "XVSADDV": {0xE88F << 15, true},
"XVSADDBU": {0xE894 << 15, true}, "XVSADDHU": {0xE895 << 15, true},
"XVSADDWU": {0xE896 << 15, true}, "XVSADDVU": {0xE897 << 15, true},
"XVSSUBB": {0xE890 << 15, true}, "XVSSUBH": {0xE891 << 15, true},
"XVSSUBW": {0xE892 << 15, true}, "XVSSUBV": {0xE893 << 15, true},
"XVSSUBBU": {0xE898 << 15, true}, "XVSSUBHU": {0xE899 << 15, true},
"XVSSUBWU": {0xE89A << 15, true}, "XVSSUBVU": {0xE89B << 15, true},
}
// The multiply families: plain and high-half [X]VMUL/[X]VMUH, the
// widening [X]VMULW{EV,OD} ladder and its accumulating [X]VMADDW twins,
// plus the [X]VMADD/[X]VMSUB fused multiply-add and the [X]VDIV/[X]VMOD
// divide and modulo pairs.
muldiv := map[string]l64Vec3Enc{
"VMULB": {0xE108 << 15, false}, "VMULH": {0xE109 << 15, false},
"VMULW": {0xE10A << 15, false}, "VMULV": {0xE10B << 15, false},
"VMUHB": {0xE10C << 15, false}, "VMUHH": {0xE10D << 15, false},
"VMUHW": {0xE10E << 15, false}, "VMUHV": {0xE10F << 15, false},
"VMUHBU": {0xE110 << 15, false}, "VMUHHU": {0xE111 << 15, false},
"VMUHWU": {0xE112 << 15, false}, "VMUHVU": {0xE113 << 15, false},
"VMULWEVHB": {0xE120 << 15, false}, "VMULWEVWH": {0xE121 << 15, false},
"VMULWEVVW": {0xE122 << 15, false}, "VMULWEVQV": {0xE123 << 15, false},
"VMULWODHB": {0xE124 << 15, false}, "VMULWODWH": {0xE125 << 15, false},
"VMULWODVW": {0xE126 << 15, false}, "VMULWODQV": {0xE127 << 15, false},
"VMULWEVHBU": {0xE130 << 15, false}, "VMULWEVWHU": {0xE131 << 15, false},
"VMULWEVVWU": {0xE132 << 15, false}, "VMULWEVQVU": {0xE133 << 15, false},
"VMULWODHBU": {0xE134 << 15, false}, "VMULWODWHU": {0xE135 << 15, false},
"VMULWODVWU": {0xE136 << 15, false}, "VMULWODQVU": {0xE137 << 15, false},
"VMULWEVHBUB": {0xE140 << 15, false}, "VMULWEVWHUH": {0xE141 << 15, false},
"VMULWEVVWUW": {0xE142 << 15, false}, "VMULWEVQVUV": {0xE143 << 15, false},
"VMULWODHBUB": {0xE144 << 15, false}, "VMULWODWHUH": {0xE145 << 15, false},
"VMULWODVWUW": {0xE146 << 15, false}, "VMULWODQVUV": {0xE147 << 15, false},
"VMADDB": {0xE150 << 15, false}, "VMADDH": {0xE151 << 15, false},
"VMADDW": {0xE152 << 15, false}, "VMADDV": {0xE153 << 15, false},
"VMSUBB": {0xE154 << 15, false}, "VMSUBH": {0xE155 << 15, false},
"VMSUBW": {0xE156 << 15, false}, "VMSUBV": {0xE157 << 15, false},
"VMADDWEVHB": {0xE158 << 15, false}, "VMADDWEVWH": {0xE159 << 15, false},
"VMADDWEVVW": {0xE15A << 15, false}, "VMADDWEVQV": {0xE15B << 15, false},
"VMADDWODHB": {0xE15C << 15, false}, "VMADDWODWH": {0xE15D << 15, false},
"VMADDWODVW": {0xE15E << 15, false}, "VMADDWODQV": {0xE15F << 15, false},
"VMADDWEVHBU": {0xE168 << 15, false}, "VMADDWEVWHU": {0xE169 << 15, false},
"VMADDWEVVWU": {0xE16A << 15, false}, "VMADDWEVQVU": {0xE16B << 15, false},
"VMADDWODHBU": {0xE16C << 15, false}, "VMADDWODWHU": {0xE16D << 15, false},
"VMADDWODVWU": {0xE16E << 15, false}, "VMADDWODQVU": {0xE16F << 15, false},
"VMADDWEVHBUB": {0xE178 << 15, false}, "VMADDWEVWHUH": {0xE179 << 15, false},
"VMADDWEVVWUW": {0xE17A << 15, false}, "VMADDWEVQVUV": {0xE17B << 15, false},
"VMADDWODHBUB": {0xE17C << 15, false}, "VMADDWODWHUH": {0xE17D << 15, false},
"VMADDWODVWUW": {0xE17E << 15, false}, "VMADDWODQVUV": {0xE17F << 15, false},
"VDIVB": {0xE1C0 << 15, false}, "VDIVH": {0xE1C1 << 15, false},
"VDIVW": {0xE1C2 << 15, false}, "VDIVV": {0xE1C3 << 15, false},
"VMODB": {0xE1C4 << 15, false}, "VMODH": {0xE1C5 << 15, false},
"VMODW": {0xE1C6 << 15, false}, "VMODV": {0xE1C7 << 15, false},
"VDIVBU": {0xE1C8 << 15, false}, "VDIVHU": {0xE1C9 << 15, false},
"VDIVWU": {0xE1CA << 15, false}, "VDIVVU": {0xE1CB << 15, false},
"VMODBU": {0xE1CC << 15, false}, "VMODHU": {0xE1CD << 15, false},
"VMODWU": {0xE1CE << 15, false}, "VMODVU": {0xE1CF << 15, false},
"VMULF": {0xE271 << 15, false}, "VMULD": {0xE272 << 15, false},
"VDIVF": {0xE275 << 15, false}, "VDIVD": {0xE276 << 15, false},
"XVMULB": {0xE908 << 15, true}, "XVMULH": {0xE909 << 15, true},
"XVMULW": {0xE90A << 15, true}, "XVMULV": {0xE90B << 15, true},
"XVMUHB": {0xE90C << 15, true}, "XVMUHH": {0xE90D << 15, true},
"XVMUHW": {0xE90E << 15, true}, "XVMUHV": {0xE90F << 15, true},
"XVMUHBU": {0xE910 << 15, true}, "XVMUHHU": {0xE911 << 15, true},
"XVMUHWU": {0xE912 << 15, true}, "XVMUHVU": {0xE913 << 15, true},
"XVMULWEVHB": {0xE920 << 15, true}, "XVMULWEVWH": {0xE921 << 15, true},
"XVMULWEVVW": {0xE922 << 15, true}, "XVMULWEVQV": {0xE923 << 15, true},
"XVMULWODHB": {0xE924 << 15, true}, "XVMULWODWH": {0xE925 << 15, true},
"XVMULWODVW": {0xE926 << 15, true}, "XVMULWODQV": {0xE927 << 15, true},
"XVMULWEVHBU": {0xE930 << 15, true}, "XVMULWEVWHU": {0xE931 << 15, true},
"XVMULWEVVWU": {0xE932 << 15, true}, "XVMULWEVQVU": {0xE933 << 15, true},
"XVMULWODHBU": {0xE934 << 15, true}, "XVMULWODWHU": {0xE935 << 15, true},
"XVMULWODVWU": {0xE936 << 15, true}, "XVMULWODQVU": {0xE937 << 15, true},
"XVMULWEVHBUB": {0xE940 << 15, true}, "XVMULWEVWHUH": {0xE941 << 15, true},
"XVMULWEVVWUW": {0xE942 << 15, true}, "XVMULWEVQVUV": {0xE943 << 15, true},
"XVMULWODHBUB": {0xE944 << 15, true}, "XVMULWODWHUH": {0xE945 << 15, true},
"XVMULWODVWUW": {0xE946 << 15, true}, "XVMULWODQVUV": {0xE947 << 15, true},
"XVMADDB": {0xE950 << 15, true}, "XVMADDH": {0xE951 << 15, true},
"XVMADDW": {0xE952 << 15, true}, "XVMADDV": {0xE953 << 15, true},
"XVMSUBB": {0xE954 << 15, true}, "XVMSUBH": {0xE955 << 15, true},
"XVMSUBW": {0xE956 << 15, true}, "XVMSUBV": {0xE957 << 15, true},
"XVMADDWEVHB": {0xE958 << 15, true}, "XVMADDWEVWH": {0xE959 << 15, true},
"XVMADDWEVVW": {0xE95A << 15, true}, "XVMADDWEVQV": {0xE95B << 15, true},
"XVMADDWODHB": {0xE95C << 15, true}, "XVMADDWODWH": {0xE95D << 15, true},
"XVMADDWODVW": {0xE95E << 15, true}, "XVMADDWODQV": {0xE95F << 15, true},
"XVMADDWEVHBU": {0xE968 << 15, true}, "XVMADDWEVWHU": {0xE969 << 15, true},
"XVMADDWEVVWU": {0xE96A << 15, true}, "XVMADDWEVQVU": {0xE96B << 15, true},
"XVMADDWODHBU": {0xE96C << 15, true}, "XVMADDWODWHU": {0xE96D << 15, true},
"XVMADDWODVWU": {0xE96E << 15, true}, "XVMADDWODQVU": {0xE96F << 15, true},
"XVMADDWEVHBUB": {0xE978 << 15, true}, "XVMADDWEVWHUH": {0xE979 << 15, true},
"XVMADDWEVVWUW": {0xE97A << 15, true}, "XVMADDWEVQVUV": {0xE97B << 15, true},
"XVMADDWODHBUB": {0xE97C << 15, true}, "XVMADDWODWHUH": {0xE97D << 15, true},
"XVMADDWODVWUW": {0xE97E << 15, true}, "XVMADDWODQVUV": {0xE97F << 15, true},
"XVDIVB": {0xE9C0 << 15, true}, "XVDIVH": {0xE9C1 << 15, true},
"XVDIVW": {0xE9C2 << 15, true}, "XVDIVV": {0xE9C3 << 15, true},
"XVMODB": {0xE9C4 << 15, true}, "XVMODH": {0xE9C5 << 15, true},
"XVMODW": {0xE9C6 << 15, true}, "XVMODV": {0xE9C7 << 15, true},
"XVDIVBU": {0xE9C8 << 15, true}, "XVDIVHU": {0xE9C9 << 15, true},
"XVDIVWU": {0xE9CA << 15, true}, "XVDIVVU": {0xE9CB << 15, true},
"XVMODBU": {0xE9CC << 15, true}, "XVMODHU": {0xE9CD << 15, true},
"XVMODWU": {0xE9CE << 15, true}, "XVMODVU": {0xE9CF << 15, true},
"XVMULF": {0xEA71 << 15, true}, "XVMULD": {0xEA72 << 15, true},
"XVDIVF": {0xEA75 << 15, true}, "XVDIVD": {0xEA76 << 15, true},
}
// The lane-wise shifts and rotates (three-register forms; the immediate
// forms live in l64VecImmInfo), the interleave families, the bit
// clear/set/rev register forms, the remaining logic and compare
// spellings, the widening add/subtract ladder and the vector FP
// arithmetic.
vecmisc := map[string]l64Vec3Enc{
"VSLLB": {0xE1D0 << 15, false}, "VSLLH": {0xE1D1 << 15, false},
"VSLLW": {0xE1D2 << 15, false}, "VSLLV": {0xE1D3 << 15, false},
"VSRLB": {0xE1D4 << 15, false}, "VSRLH": {0xE1D5 << 15, false},
"VSRLW": {0xE1D6 << 15, false}, "VSRLV": {0xE1D7 << 15, false},
"VSRAH": {0xE1D9 << 15, false}, "VSRAW": {0xE1DA << 15, false},
"VSRAV": {0xE1DB << 15, false},
"VROTRB": {0xE1DC << 15, false}, "VROTRH": {0xE1DD << 15, false},
"VROTRV": {0xE1DF << 15, false},
"VILVLB": {0xE234 << 15, false}, "VILVLH": {0xE235 << 15, false},
"VILVLW": {0xE236 << 15, false}, "VILVLV": {0xE237 << 15, false},
"VILVHB": {0xE238 << 15, false}, "VILVHH": {0xE239 << 15, false},
"VILVHW": {0xE23A << 15, false}, "VILVHV": {0xE23B << 15, false},
"VBITCLRB": {0xE218 << 15, false}, "VBITCLRH": {0xE219 << 15, false},
"VBITCLRW": {0xE21A << 15, false}, "VBITCLRV": {0xE21B << 15, false},
"VBITSETB": {0xE21C << 15, false}, "VBITSETH": {0xE21D << 15, false},
"VBITSETW": {0xE21E << 15, false}, "VBITSETV": {0xE21F << 15, false},
"VBITREVB": {0xE220 << 15, false}, "VBITREVH": {0xE221 << 15, false},
"VBITREVW": {0xE222 << 15, false}, "VBITREVV": {0xE223 << 15, false},
"VORV": {0xE24D << 15, false}, "VNORV": {0xE24F << 15, false},
"VANDNV": {0xE250 << 15, false}, "VORNV": {0xE251 << 15, false},
"VSEQH": {0xE001 << 15, false}, "VSEQW": {0xE002 << 15, false},
"VSLTB": {0xE00C << 15, false}, "VSLTH": {0xE00D << 15, false},
"VSLTW": {0xE00E << 15, false}, "VSLTV": {0xE00F << 15, false},
"VSLTBU": {0xE010 << 15, false}, "VSLTHU": {0xE011 << 15, false},
"VSLTWU": {0xE012 << 15, false}, "VSLTVU": {0xE013 << 15, false},
"VADDWEVHB": {0xE03C << 15, false}, "VADDWEVWH": {0xE03D << 15, false},
"VADDWEVVW": {0xE03E << 15, false}, "VADDWEVQV": {0xE03F << 15, false},
"VSUBWEVHB": {0xE040 << 15, false}, "VSUBWEVWH": {0xE041 << 15, false},
"VSUBWEVVW": {0xE042 << 15, false}, "VSUBWEVQV": {0xE043 << 15, false},
"VADDWODHB": {0xE044 << 15, false}, "VADDWODWH": {0xE045 << 15, false},
"VADDWODVW": {0xE046 << 15, false}, "VADDWODQV": {0xE047 << 15, false},
"VSUBWODHB": {0xE048 << 15, false}, "VSUBWODWH": {0xE049 << 15, false},
"VSUBWODVW": {0xE04A << 15, false}, "VSUBWODQV": {0xE04B << 15, false},
"VSUBWEVHBU": {0xE060 << 15, false}, "VSUBWEVWHU": {0xE061 << 15, false},
"VSUBWEVVWU": {0xE062 << 15, false}, "VSUBWEVQVU": {0xE063 << 15, false},
"VADDWEVHBU": {0xE05C << 15, false}, "VADDWEVWHU": {0xE05D << 15, false},
"VADDWEVVWU": {0xE05E << 15, false}, "VADDWEVQVU": {0xE05F << 15, false},
"VADDWODHBU": {0xE064 << 15, false}, "VADDWODWHU": {0xE065 << 15, false},
"VADDWODVWU": {0xE066 << 15, false}, "VADDWODQVU": {0xE067 << 15, false},
"VSUBWODHBU": {0xE068 << 15, false}, "VSUBWODWHU": {0xE069 << 15, false},
"VSUBWODVWU": {0xE06A << 15, false}, "VSUBWODQVU": {0xE06B << 15, false},
"VSHUFH": {0xE2F5 << 15, false}, "VSHUFW": {0xE2F6 << 15, false},
"VSHUFV": {0xE2F7 << 15, false},
"XVSLLB": {0xE9D0 << 15, true}, "XVSLLH": {0xE9D1 << 15, true},
"XVSLLW": {0xE9D2 << 15, true}, "XVSLLV": {0xE9D3 << 15, true},
"XVSRLB": {0xE9D4 << 15, true}, "XVSRLH": {0xE9D5 << 15, true},
"XVSRLW": {0xE9D6 << 15, true}, "XVSRLV": {0xE9D7 << 15, true},
"XVSRAB": {0xE9D8 << 15, true}, "XVSRAH": {0xE9D9 << 15, true},
"XVSRAW": {0xE9DA << 15, true}, "XVSRAV": {0xE9DB << 15, true},
"XVROTRB": {0xE9DC << 15, true}, "XVROTRH": {0xE9DD << 15, true},
"XVROTRW": {0xE9DE << 15, true}, "XVROTRV": {0xE9DF << 15, true},
"XVILVLB": {0xEA34 << 15, true}, "XVILVLH": {0xEA35 << 15, true},
"XVILVLW": {0xEA36 << 15, true}, "XVILVLV": {0xEA37 << 15, true},
"XVILVHB": {0xEA38 << 15, true}, "XVILVHH": {0xEA39 << 15, true},
"XVILVHW": {0xEA3A << 15, true}, "XVILVHV": {0xEA3B << 15, true},
"XVBITCLRB": {0xEA18 << 15, true}, "XVBITCLRH": {0xEA19 << 15, true},
"XVBITCLRW": {0xEA1A << 15, true}, "XVBITCLRV": {0xEA1B << 15, true},
"XVBITSETB": {0xEA1C << 15, true}, "XVBITSETH": {0xEA1D << 15, true},
"XVBITSETW": {0xEA1E << 15, true}, "XVBITSETV": {0xEA1F << 15, true},
"XVBITREVB": {0xEA20 << 15, true}, "XVBITREVH": {0xEA21 << 15, true},
"XVBITREVW": {0xEA22 << 15, true}, "XVBITREVV": {0xEA23 << 15, true},
"XVORV": {0xEA4D << 15, true}, "XVNORV": {0xEA4F << 15, true},
"XVANDNV": {0xEA50 << 15, true}, "XVORNV": {0xEA51 << 15, true},
"XVSEQH": {0xE801 << 15, true}, "XVSEQW": {0xE802 << 15, true},
"XVSLTB": {0xE80C << 15, true}, "XVSLTH": {0xE80D << 15, true},
"XVSLTW": {0xE80E << 15, true}, "XVSLTV": {0xE80F << 15, true},
"XVSLTBU": {0xE810 << 15, true}, "XVSLTHU": {0xE811 << 15, true},
"XVSLTWU": {0xE812 << 15, true}, "XVSLTVU": {0xE813 << 15, true},
"XVADDWEVHB": {0xE83C << 15, true}, "XVADDWEVWH": {0xE83D << 15, true},
"XVADDWEVVW": {0xE83E << 15, true}, "XVADDWEVQV": {0xE83F << 15, true},
"XVSUBWEVHB": {0xE840 << 15, true}, "XVSUBWEVWH": {0xE841 << 15, true},
"XVSUBWEVVW": {0xE842 << 15, true}, "XVSUBWEVQV": {0xE843 << 15, true},
"XVADDWODHB": {0xE844 << 15, true}, "XVADDWODWH": {0xE845 << 15, true},
"XVADDWODVW": {0xE846 << 15, true}, "XVADDWODQV": {0xE847 << 15, true},
"XVSUBWODHB": {0xE848 << 15, true}, "XVSUBWODWH": {0xE849 << 15, true},
"XVSUBWODVW": {0xE84A << 15, true}, "XVSUBWODQV": {0xE84B << 15, true},
"XVADDWEVHBU": {0xE85C << 15, true}, "XVADDWEVWHU": {0xE85D << 15, true},
"XVADDWEVVWU": {0xE85E << 15, true}, "XVADDWEVQVU": {0xE85F << 15, true},
"XVSUBWEVHBU": {0xE860 << 15, true}, "XVSUBWEVWHU": {0xE861 << 15, true},
"XVSUBWEVVWU": {0xE862 << 15, true}, "XVSUBWEVQVU": {0xE863 << 15, true},
"XVADDWODHBU": {0xE864 << 15, true}, "XVADDWODWHU": {0xE865 << 15, true},
"XVADDWODVWU": {0xE866 << 15, true}, "XVADDWODQVU": {0xE867 << 15, true},
"XVSUBWODHBU": {0xE868 << 15, true}, "XVSUBWODWHU": {0xE869 << 15, true},
"XVSUBWODVWU": {0xE86A << 15, true}, "XVSUBWODQVU": {0xE86B << 15, true},
"XVSHUFH": {0xEAF5 << 15, true}, "XVSHUFW": {0xEAF6 << 15, true},
"XVSHUFV": {0xEAF7 << 15, true},
}
for _, tab := range []map[string]l64Vec3Enc{addsub, muldiv, vecmisc} {
for m, e := range tab {
if _, dup := vec3[m]; dup {
panic("loong64: duplicate vector mnemonic " + m)
}
vec3[m] = e
}
}
for m, e := range vec3 {
l64InstrTable[m] = l64Enc{format: l64Fvvv, op: e.op}
l64VecBank[m] = e.lasx
@@ -597,21 +850,156 @@ func init() {
// Immediate forms: INSTR $imm, vj, vd (or INSTR $imm, vd). The immediate
// range, bias and field mask are the ones the toolchain encodes: vandi.b
// stores the raw 8-bit constant, vsrai.b stores imm+8 (byte-lane bias),
// vseqi.b and vseqi.d store 5-bit and 7-bit two's-complement values.
// The mnemonics that also have a register form (VSEQB, VSEQV, VSRAB,
// VROTRW) keep their three-register entry in l64InstrTable; the
// dispatcher picks the immediate opcode from l64VecImmInfo by operand
// kind, so the immediate entries must not overwrite the table.
// stores the raw 8-bit constant, vsrari.b stores imm+8 (lane-width
// bias), the si5 compares store 5-bit two's-complement values and vseqi.d
// a 7-bit field the toolchain range-checks down to si5.
// The mnemonics that also have a register form (the shifts, the bit
// clear/set/rev families, VSEQ and the logic immediates) keep their
// three-register entry in l64InstrTable; the dispatcher picks the
// immediate opcode from l64VecImmInfo by operand kind, so the immediate
// entries must not overwrite the table.
vecImm := map[string]l64VecImmEnc{
"VANDB": {0xE7A0 << 15, false, 0, 255, 0, 0xFF},
"XVANDB": {0xEFA0 << 15, true, 0, 255, 0, 0xFF},
"VORB": {0xE7A8 << 15, false, 0, 255, 0, 0xFF},
"XVORB": {0xEFA8 << 15, true, 0, 255, 0, 0xFF},
"VXORB": {0xE7B0 << 15, false, 0, 255, 0, 0xFF},
"XVXORB": {0xEFB0 << 15, true, 0, 255, 0, 0xFF},
"VNORB": {0xE7B8 << 15, false, 0, 255, 0, 0xFF},
"XVNORB": {0xEFB8 << 15, true, 0, 255, 0, 0xFF},
"VSEQB": {0xE500 << 15, false, -16, 15, 0, 0x1F},
"XVSEQB": {0xE900 << 15, true, -16, 15, 0, 0x1F},
"VSEQV": {0xE503 << 15, false, -64, 63, 0, 0x7F},
"XVSEQV": {0xE903 << 15, true, -64, 63, 0, 0x7F},
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
// vseqi.h/w accept the same si5 window as vseqi.b; vseqi.d carries a
// 7-bit field, but the toolchain range-checks it down to si5 as well
// (GOARCH=loong64 go tool asm rejects VSEQV $32 and VSEQV $-64).
"VSEQH": {0xE501 << 15, false, -16, 15, 0, 0x1F},
"XVSEQH": {0xED01 << 15, true, -16, 15, 0, 0x1F},
"VSEQW": {0xE502 << 15, false, -16, 15, 0, 0x1F},
"XVSEQW": {0xED02 << 15, true, -16, 15, 0, 0x1F},
"VSEQV": {0xE503 << 15, false, -16, 15, 0, 0x7F},
"XVSEQV": {0xE903 << 15, true, -16, 15, 0, 0x7F},
// vslti compares against a signed (or, in the U spellings, unsigned)
// si5/ui5 constant.
"VSLTB": {0xE50C << 15, false, -16, 15, 0, 0x1F},
"XVSLTB": {0xED0C << 15, true, -16, 15, 0, 0x1F},
"VSLTH": {0xE50D << 15, false, -16, 15, 0, 0x1F},
"XVSLTH": {0xED0D << 15, true, -16, 15, 0, 0x1F},
"VSLTW": {0xE50E << 15, false, -16, 15, 0, 0x1F},
"XVSLTW": {0xED0E << 15, true, -16, 15, 0, 0x1F},
"VSLTV": {0xE50F << 15, false, -16, 15, 0, 0x1F},
"XVSLTV": {0xED0F << 15, true, -16, 15, 0, 0x1F},
"VSLTBU": {0xE510 << 15, false, 0, 31, 0, 0x1F},
"XVSLTBU": {0xED10 << 15, true, 0, 31, 0, 0x1F},
"VSLTHU": {0xE511 << 15, false, 0, 31, 0, 0x1F},
"XVSLTHU": {0xED11 << 15, true, 0, 31, 0, 0x1F},
"VSLTWU": {0xE512 << 15, false, 0, 31, 0, 0x1F},
"XVSLTWU": {0xED12 << 15, true, 0, 31, 0, 0x1F},
"VSLTVU": {0xE513 << 15, false, 0, 31, 0, 0x1F},
"XVSLTVU": {0xED13 << 15, true, 0, 31, 0, 0x1F},
// vaddi/vsubi take ui5 constants for every width on this toolchain
// (VADDVU $32 is rejected by the oracle although the field is ui8).
"VADDBU": {0xE514 << 15, false, 0, 31, 0, 0x1F},
"XVADDBU": {0xED14 << 15, true, 0, 31, 0, 0x1F},
"VADDHU": {0xE515 << 15, false, 0, 31, 0, 0x1F},
"XVADDHU": {0xED15 << 15, true, 0, 31, 0, 0x1F},
"VADDWU": {0xE516 << 15, false, 0, 31, 0, 0x1F},
"XVADDWU": {0xED16 << 15, true, 0, 31, 0, 0x1F},
"VADDVU": {0xE517 << 15, false, 0, 31, 0, 0x1F},
"XVADDVU": {0xED17 << 15, true, 0, 31, 0, 0x1F},
"VSUBBU": {0xE518 << 15, false, 0, 31, 0, 0x1F},
"XVSUBBU": {0xED18 << 15, true, 0, 31, 0, 0x1F},
"VSUBHU": {0xE519 << 15, false, 0, 31, 0, 0x1F},
"XVSUBHU": {0xED19 << 15, true, 0, 31, 0, 0x1F},
"VSUBWU": {0xE51A << 15, false, 0, 31, 0, 0x1F},
"XVSUBWU": {0xED1A << 15, true, 0, 31, 0, 0x1F},
"VSUBVU": {0xE51B << 15, false, 0, 31, 0, 0x1F},
"XVSUBVU": {0xED1B << 15, true, 0, 31, 0, 0x1F},
// The shift/rotate immediates ride in a width-sized field whose upper
// bits carry the lane-width code: vslli.b stores ui3 at [12:0] with
// bits [14:13] inside the opcode, vslli.h ui4 under a 4 bit mask, and
// the .w/.d spellings a raw ui5/ui6.
"VSLLB": {0x732C2000, false, 0, 7, 0, 0x7},
"XVSLLB": {0x772C2000, true, 0, 7, 0, 0x7},
"VSLLH": {0x732C4000, false, 0, 15, 0, 0xF},
"XVSLLH": {0x772C4000, true, 0, 15, 0, 0xF},
"VSLLW": {0xE659 << 15, false, 0, 31, 0, 0x1F},
"XVSLLW": {0xEE59 << 15, true, 0, 31, 0, 0x1F},
"VSLLV": {0xE65A << 15, false, 0, 63, 0, 0x3F},
"XVSLLV": {0xEE5A << 15, true, 0, 63, 0, 0x3F},
"VSRLB": {0x73302000, false, 0, 7, 0, 0x7},
"XVSRLB": {0x77302000, true, 0, 7, 0, 0x7},
"VSRLH": {0x73304000, false, 0, 15, 0, 0xF},
"XVSRLH": {0x77304000, true, 0, 15, 0, 0xF},
"VSRLW": {0xE661 << 15, false, 0, 31, 0, 0x1F},
"XVSRLW": {0xEE61 << 15, true, 0, 31, 0, 0x1F},
"VSRLV": {0xE662 << 15, false, 0, 63, 0, 0x3F},
"XVSRLV": {0xEE62 << 15, true, 0, 63, 0, 0x3F},
// vsrari/vrotri bias the field so the lane-width code rides above the
// shift amount (.b adds 8, .h 16, .w 32; .d is a raw ui6).
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
"XVSRAB": {0xEE68 << 15, true, 0, 7, 8, 0x1F},
"VSRAH": {0x73344000, false, 0, 15, 0, 0xF},
"XVSRAH": {0x77344000, true, 0, 15, 0, 0xF},
"VSRAW": {0xE669 << 15, false, 0, 31, 0, 0x1F},
"XVSRAW": {0xEE69 << 15, true, 0, 31, 0, 0x1F},
"VSRAV": {0xE66A << 15, false, 0, 63, 0, 0x3F},
"XVSRAV": {0xEE6A << 15, true, 0, 63, 0, 0x3F},
"VROTRB": {0x72A02000, false, 0, 7, 0, 0x7},
"XVROTRB": {0x76A02000, true, 0, 7, 0, 0x7},
"VROTRH": {0x72A04000, false, 0, 15, 0, 0xF},
"XVROTRH": {0x76A04000, true, 0, 15, 0, 0xF},
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
"XVROTRW": {0xED41 << 15, true, 0, 31, 0, 0x1F},
"VROTRV": {0xE542 << 15, false, 0, 63, 0, 0x3F},
"XVROTRV": {0xED42 << 15, true, 0, 63, 0, 0x3F},
// vbitclri/vbitseti/vbitrevi follow the same width-coded layout.
"VBITCLRB": {0x73102000, false, 0, 7, 0, 0x7},
"XVBITCLRB": {0x77102000, true, 0, 7, 0, 0x7},
"VBITCLRH": {0x73104000, false, 0, 15, 0, 0xF},
"XVBITCLRH": {0x77104000, true, 0, 15, 0, 0xF},
"VBITCLRW": {0xE621 << 15, false, 0, 31, 0, 0x1F},
"XVBITCLRW": {0xEE21 << 15, true, 0, 31, 0, 0x1F},
"VBITCLRV": {0xE622 << 15, false, 0, 63, 0, 0x3F},
"XVBITCLRV": {0xEE22 << 15, true, 0, 63, 0, 0x3F},
"VBITSETB": {0x73142000, false, 0, 7, 0, 0x7},
"XVBITSETB": {0x77142000, true, 0, 7, 0, 0x7},
"VBITSETH": {0x73144000, false, 0, 15, 0, 0xF},
"XVBITSETH": {0x77144000, true, 0, 15, 0, 0xF},
"VBITSETW": {0xE629 << 15, false, 0, 31, 0, 0x1F},
"XVBITSETW": {0xEE29 << 15, true, 0, 31, 0, 0x1F},
"VBITSETV": {0xE62A << 15, false, 0, 63, 0, 0x3F},
"XVBITSETV": {0xEE2A << 15, true, 0, 63, 0, 0x3F},
"VBITREVB": {0x73182000, false, 0, 7, 0, 0x7},
"XVBITREVB": {0x77182000, true, 0, 7, 0, 0x7},
"VBITREVH": {0x73184000, false, 0, 15, 0, 0xF},
"XVBITREVH": {0x77184000, true, 0, 15, 0, 0xF},
"VBITREVW": {0xE631 << 15, false, 0, 31, 0, 0x1F},
"XVBITREVW": {0xEE31 << 15, true, 0, 31, 0, 0x1F},
"VBITREVV": {0xE632 << 15, false, 0, 63, 0, 0x3F},
"XVBITREVV": {0xEE32 << 15, true, 0, 63, 0, 0x3F},
// The 4-bit-select shuffles and the byte-extract/insert permutations
// take ui8 (the .d shuffle ui4 range-checked to 0..15 by the
// toolchain) packing both position nibbles.
"VSHUF4IB": {0xE720 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IB": {0xEF20 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IH": {0xE728 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IH": {0xEF28 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IW": {0xE730 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IW": {0xEF30 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IV": {0xE738 << 15, false, 0, 15, 0, 0xFF},
"XVSHUF4IV": {0xEF38 << 15, true, 0, 15, 0, 0xFF},
"VPERMIW": {0xE7C8 << 15, false, 0, 255, 0, 0xFF},
"XVPERMIW": {0xEFC8 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIV": {0xEFD0 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIQ": {0xEFD8 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSB": {0xE718 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSB": {0xEF18 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSH": {0xE710 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSH": {0xEF10 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSW": {0xE708 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSW": {0xEF08 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSV": {0xE700 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSV": {0xEF00 << 15, true, 0, 255, 0, 0xFF},
}
for m, e := range vecImm {
l64VecImmInfo[m] = e
@@ -625,22 +1013,71 @@ func init() {
"VSETANYEQB": 0xE539<<15 | 8<<10, "XVSETANYEQB": 0xED39<<15 | 8<<10,
"VSETANYEQV": 0xE539<<15 | 11<<10, "XVSETANYEQV": 0xED39<<15 | 11<<10,
"VSETALLNEV": 0xE539<<15 | 15<<10, "XVSETALLNEV": 0xED39<<15 | 15<<10,
"VSETEQV": 0xE539<<15 | 6<<10, "XVSETEQV": 0xED39<<15 | 6<<10,
"VSETANYEQH": 0xE539<<15 | 9<<10, "XVSETANYEQH": 0xED39<<15 | 9<<10,
"VSETANYEQW": 0xE539<<15 | 10<<10, "XVSETANYEQW": 0xED39<<15 | 10<<10,
"VSETALLNEB": 0xE539<<15 | 12<<10, "XVSETALLNEB": 0xED39<<15 | 12<<10,
"VSETALLNEH": 0xE539<<15 | 13<<10, "XVSETALLNEH": 0xED39<<15 | 13<<10,
"VSETALLNEW": 0xE539<<15 | 14<<10, "XVSETALLNEW": 0xED39<<15 | 14<<10,
}
for m, op := range vecCf {
l64InstrTable[m] = l64Enc{format: l64Fvcf, op: op}
l64VecBank[m] = strings.HasPrefix(m, "XV")
}
// Lane popcount: INSTR vj, vd (the 2R layout with the opcode extending
// over the unused vk field).
// Lane popcount and the two-operand vector FP/unary spellings: INSTR vj,
// vd (the 2R layout with the opcode extending over the unused vk field;
// the low byte of each constant is the instruction's own sub-op).
vec2r := map[string]l64Vec3Enc{
"VPCNTV": {0x1CA70B << 10, false}, "XVPCNTV": {0x1DA70B << 10, true},
}
// The rest of the lane popcounts, the vector negations and the vector FP
// unary conversions (loong64enc1.s).
vec2rMore := map[string]l64Vec3Enc{
"VPCNTB": {0x1CA708 << 10, false}, "VPCNTH": {0x1CA709 << 10, false},
"VPCNTW": {0x1CA70A << 10, false},
"VNEGB": {0x1CA70C << 10, false}, "VNEGH": {0x1CA70D << 10, false},
"VNEGW": {0x1CA70E << 10, false}, "VNEGV": {0x1CA70F << 10, false},
"VFCLASSF": {0x1CA735 << 10, false}, "VFCLASSD": {0x1CA736 << 10, false},
"VFSQRTF": {0x1CA739 << 10, false}, "VFSQRTD": {0x1CA73A << 10, false},
"VFRECIPF": {0x1CA73D << 10, false}, "VFRECIPD": {0x1CA73E << 10, false},
"VFRSQRTF": {0x1CA741 << 10, false}, "VFRSQRTD": {0x1CA742 << 10, false},
"VFRINTF": {0x1CA74D << 10, false}, "VFRINTD": {0x1CA74E << 10, false},
"VFRINTRMF": {0x1CA751 << 10, false}, "VFRINTRMD": {0x1CA752 << 10, false},
"VFRINTRPF": {0x1CA755 << 10, false}, "VFRINTRPD": {0x1CA756 << 10, false},
"VFRINTRZF": {0x1CA759 << 10, false}, "VFRINTRZD": {0x1CA75A << 10, false},
"VFRINTRNEF": {0x1CA75D << 10, false}, "VFRINTRNED": {0x1CA75E << 10, false},
"XVPCNTB": {0x1DA708 << 10, true}, "XVPCNTH": {0x1DA709 << 10, true},
"XVPCNTW": {0x1DA70A << 10, true},
"XVNEGB": {0x1DA70C << 10, true}, "XVNEGH": {0x1DA70D << 10, true},
"XVNEGW": {0x1DA70E << 10, true}, "XVNEGV": {0x1DA70F << 10, true},
"XVFCLASSF": {0x1DA735 << 10, true}, "XVFCLASSD": {0x1DA736 << 10, true},
"XVFSQRTF": {0x1DA739 << 10, true}, "XVFSQRTD": {0x1DA73A << 10, true},
"XVFRECIPF": {0x1DA73D << 10, true}, "XVFRECIPD": {0x1DA73E << 10, true},
"XVFRSQRTF": {0x1DA741 << 10, true}, "XVFRSQRTD": {0x1DA742 << 10, true},
"XVFRINTF": {0x1DA74D << 10, true}, "XVFRINTD": {0x1DA74E << 10, true},
"XVFRINTRMF": {0x1DA751 << 10, true}, "XVFRINTRMD": {0x1DA752 << 10, true},
"XVFRINTRPF": {0x1DA755 << 10, true}, "XVFRINTRPD": {0x1DA756 << 10, true},
"XVFRINTRZF": {0x1DA759 << 10, true}, "XVFRINTRZD": {0x1DA75A << 10, true},
"XVFRINTRNEF": {0x1DA75D << 10, true}, "XVFRINTRNED": {0x1DA75E << 10, true},
}
maps.Copy(vec2r, vec2rMore)
for m, e := range vec2r {
l64InstrTable[m] = l64Enc{format: l64Frr, op: e.op}
l64VecBank[m] = e.lasx
l64Vec2R[m] = true
}
// The four-register byte shuffle: INSTR va, vk, vj, vd (the operand the
// table reads in each field position, va at bits [19:15]).
vec4r := map[string]l64Vec3Enc{
"VSHUFB": {0x0D50 << 16, false}, "XVSHUFB": {0x0D60 << 16, true},
}
for m, e := range vec4r {
l64InstrTable[m] = l64Enc{format: l64Fvvvv, op: e.op}
l64VecBank[m] = e.lasx
l64Vec4R[m] = true
}
}
// l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the
+242
View File
@@ -507,6 +507,218 @@ TEXT ·v(SB), NOSPLIT, $0
0x4C000020,
)
})
// The integer and FP add/subtract families with their saturating pairs
// and immediate spellings (loong64enc1.s words).
t.Run("add and subtract families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDB V1, V2, V3
VADDF V1, V2, V3
VADDD V1, V2, V3
VSUBD V1, V2, V3
VSADDV V1, V2, V3
VSSUBVU V1, V2, V3
VADDBU $1, V2, V1
VADDBU $1, V2
VSUBVU $31, V2
XVSADDV X3, X2, X1
XVSUBD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x700A0443, // vadd.b
0x71308443, // vadd.f
0x71310443, // vadd.d
0x71330443, // vsub.d
0x70478443, // vsadd.v
0x704D8443, // vssub.u.d
0x728A0441, // vaddi.bu v1, v2, 1
0x728A0442, // vaddi.bu v2, v2, 1 (two-operand form)
0x728DFC42, // vsubi.du v2, v2, 31 (two-operand form)
0x74478C41, // xvsadd.d x1, x2, x3
0x75330443, // xvsub.d x3, x2, x1
0x4C000020,
)
})
// The multiply, divide and accumulate families.
t.Run("multiply and divide families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VMULV V1, V2, V3
VMUHHU V1, V2, V3
VDIVBU V1, V2, V3
VMODV V1, V2, V3
VMADDB V1, V2, V3
VMSUBV V1, V2, V3
VMULWEVHB V1, V2, V3
VMULWODQV V1, V2, V3
VMADDWEVHBUB V1, V2, V3
XVDIVD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70858443, // vmul.v
0x70888443, // vmuh.u.d
0x70E40443, // vdiv.u.b
0x70E38443, // vmod.d
0x70A80443, // vmadd.b
0x70AB8443, // vmsub.d
0x70900443, // vmulwev.h.b
0x70938443, // vmulwod.q.d
0x70BC0443, // vmaddwev.h.bu.b
0x753B0443, // xvdiv.d
0x4C000020,
)
})
// The shift, bit and interleave families in register and immediate
// spellings, with the width-coded shift immediates.
t.Run("shift, bit and interleave families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSLLV V1, V2, V3
VROTRB V1, V2, V3
VBITCLRV V1, V2, V3
VBITSETW V1, V2, V3
VBITREVV V1, V2, V3
VILVLB V1, V2, V3
VILVHV V1, V2, V3
VSLLB $7, V1, V2
VSLLB $5, V1
VSRLH $15, V1, V2
VSRAW $31, V1, V2
VSRAV $63, V1, V2
VROTRV $63, V1, V2
VBITCLRB $7, V2, V3
VBITREVV $63, V2, V3
VSEQH $-16, V2, V3
VSLTB $1, V2, V3
VSLTHU $31, V2, V3
XVILVLV X3, X2, X1
XVSLLB $7, X2, X1
XVSRAV $63, X2, X1
XVBITREVV $63, X2, X1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70E98443, // vsll.d
0x70EE0443, // vrotr.b
0x710D8443, // vbitclr.d
0x710F0443, // vbitset.w
0x71118443, // vbitrev.d
0x711A0443, // vilvl.b
0x711D8443, // vilvh.d
0x732C3C22, // vslli.b v2, v1, 7
0x732C3421, // vslli.b v1, v1, 5 (two-operand form)
0x73307C22, // vsrli.h v2, v1, 15
0x7334FC22, // vsrai.w v2, v1, 31
0x7335FC22, // vsrai.d v2, v1, 63
0x72A1FC22, // vrotri.d v2, v1, 63
0x73103C43, // vbitclri.b v3, v2, 7
0x7319FC43, // vbitrevi.d v3, v2, 63
0x7280C043, // vseqi.h v3, v2, -16
0x72860443, // vslti.b v3, v2, 1
0x7288FC43, // vslti.hu v3, v2, 31
0x751B8C41, // xvilvl.d x1, x2, x3
0x772C3C41, // xvslli.b x1, x2, 7
0x7735FC41, // xvsrai.d x1, x2, 63
0x7719FC41, // xvbitrevi.d x1, x2, 63
0x4C000020,
)
})
// The shuffle, select and permutation families, including the
// four-register byte shuffle.
t.Run("shuffle and permutation families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSHUFH V1, V2, V3
VSHUFW V1, V2, V3
VSHUFV V1, V2, V3
VSHUFB V1, V2, V3, V4
XVSHUFB X1, X2, X3, X4
VSHUF4IB $255, V2, V1
VSHUF4IV $15, V2, V1
XVSHUF4IV $15, X1, X2
VEXTRINSB $0x18, V1, V2
XVEXTRINSV $0x81, X1, X2
VPERMIW $0x1B, V1, V2
XVPERMIQ $0x4B, X1, X2
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x717A8443, // vshuf.h
0x717B0443, // vshuf.w
0x717B8443, // vshuf.d
0x0D508864, // vshuf.b v4, v3, v2, v1
0x0D608864, // xvshuf.b
0x7393FC41, // vshuf4i.b v1, v2, 255
0x739C3C41, // vshuf4i.d v1, v2, 15
0x779C3C22, // xvshuf4i.d x2, x1, 15
0x738C6022, // vextrins.b v2, v1, 0x18
0x77820422, // xvextrins.d x2, x1, 0x81
0x73E46C22, // vpermi.w v2, v1, 0x1b
0x77ED2C22, // xvpermi.q x2, x1, 0x4b
0x4C000020,
)
})
// The vector FP families, the unary spellings, the compare-to-flag
// additions and the scalar int/float conversions.
t.Run("FP and conversion families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDF V1, V2, V3
VMULF V1, V2, V3
VFCLASSD V1, V2
VFSQRTF V1, V2
VFRECIPD V1, V2
VFRSQRTF V1, V2
VFRINTF V1, V2
VFRINTRNED V1, V2
VNEGB V1, V2
VPCNTB V1, V2
XVNEGV X2, X1
XVPCNTW X3, X2
XVFRINTRNEF X1, X2
VSETEQV V1, FCC0
VSETANYEQH V1, FCC0
VSETALLNEB V1, FCC0
XVSETALLNEW X1, FCC0
FFINTFW F0, F1
FTINTVD F0, F1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x71308443, // vfadd.s
0x71388443, // vfmul.s
0x729CD822, // vfclass.d
0x729CE422, // vfsqrt.s
0x729CF822, // vfrecip.d
0x729D0422, // vfrsqrt.s
0x729D3422, // vfrint.s
0x729D7822, // vfrintne.s
0x729C3022, // vneg.b
0x729C2022, // vpcnt.b
0x769C3C41, // xvneg.d x1, x2
0x769C2862, // xvpcnt.w x2, x3
0x769D7422, // xvfrintne.s x2, x1
0x729C9820, // vseteqz.d fcc0, v1
0x729CA420, // vsetanyeqz.h
0x729CB020, // vsetallnez.b
0x769CB820, // xvsetallnez.w
0x011D1001, // ffint.s.w f1, f0
0x011B2801, // ftint.l.d f1, f0
0x4C000020,
)
})
}
// TestLOONG64_vectorErrors pins the register-class and range diagnostics of
@@ -545,6 +757,36 @@ func TestLOONG64_vectorErrors(t *testing.T) {
`TEXT ·e(SB), NOSPLIT, $0
VROTRW $32, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VADDVU $32, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSEQV $32, V2, V3
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSHUF4IV $16, V2, V1
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VEXTRINSB $256, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSLTV $-17, V2, V3
RET
`,
// VSHUFB wants four vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSHUFB V1, V2, V3
RET
`,
// The FCC forms still refuse vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSETEQV V1, V2
RET
`,
// VSET* wants an FCC flag, not a vector register.
`TEXT ·e(SB), NOSPLIT, $0
+59
View File
@@ -14,6 +14,51 @@ type Imm int64
func (Imm) isOperand() {}
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
// low register through the inverted 5-bit V'VVVV field; the three higher
// registers are implied by the instruction, so only the pair travels here.
type RegList struct {
Lo Reg
Hi Reg // implied by the encoding; Lo.idx+3 by construction
}
func (RegList) isOperand() {}
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
// encoding takes an XMM/memory source at that position rewrite it as a read
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
// own behaviour; every other instruction rejects it.
type FloatImm struct {
Text string // the numeric text as written, sign excluded
Neg bool // a leading minus
}
func (FloatImm) isOperand() {}
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
// (the linker fills the TLS slot offset).
type TLSMem struct {
Disp int64
Size int
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
}
func (TLSMem) isOperand() {}
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
// prefixes a disp32 absolute reference with no relocation. The base
// register spellings GS and FS produce it.
type SegAbs struct {
Disp int64
Size int
Seg byte // 0x64 FS, 0x65 GS
}
func (SegAbs) isOperand() {}
// Mem is a memory operand of the form disp(base)(index*scale).
type Mem struct {
Base Reg
@@ -23,6 +68,7 @@ type Mem struct {
Size int // operand width in bytes
HasBase bool
HasIndex bool
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
}
func (Mem) isOperand() {}
@@ -48,3 +94,16 @@ type sbMem struct {
}
func (sbMem) isOperand() {}
// isX86Mem reports whether the operand is an amd64 memory reference: a base
// or indexed Mem, or an SB-relative sbMem. Encoders that gate on "memory in
// this position" must accept both; the r/m emitters distinguish the two
// themselves.
func isX86Mem(o Operand) bool {
switch o.(type) {
case Mem, sbMem:
return true
default:
return false
}
}
+828 -58
View File
File diff suppressed because it is too large Load Diff
+27 -5
View File
@@ -63,7 +63,7 @@ func riscvRegNum(name string) int {
return 24
case "X25", "S9":
return 25
case "X26", "S10":
case "X26", "S10", "CTXT":
return 26
case "X27", "S11", "g":
return 27
@@ -219,6 +219,9 @@ var riscvInstrTable = map[string]riscvEnc{
"DIVUW": {0x3B, 0x5, 0x01},
"REMW": {0x3B, 0x6, 0x01},
"REMUW": {0x3B, 0x7, 0x01},
// Zicond conditional zeroing.
"CZEROEQZ": {0x33, 0x5, 0x07},
"CZERONEZ": {0x33, 0x7, 0x07},
// RV64I, I-type arithmetic.
"ADDI": {0x13, 0x0, 0x00},
"ADDIW": {0x1B, 0x0, 0x00},
@@ -247,13 +250,21 @@ var riscvInstrTable = map[string]riscvEnc{
"BGE": {0x63, 0x5, 0x00},
"BLTU": {0x63, 0x6, 0x00},
"BGEU": {0x63, 0x7, 0x00},
// The swapped-spelling comparison forms: encoded as BLT/BGE/BLTU/BGEU
// with the register operands swapped.
"BGT": {0x63, 0x4, 0x00},
"BLE": {0x63, 0x5, 0x00},
"BGTU": {0x63, 0x6, 0x00},
"BLEU": {0x63, 0x7, 0x00},
// U-type.
"LUI": {0x37, 0x0, 0x00},
"AUIPC": {0x17, 0x0, 0x00},
// System.
"ECALL": {0x73, 0x0, 0x00},
"EBREAK": {0x73, 0x0, 0x00},
"FENCE": {0x0F, 0x0, 0x00},
"ECALL": {0x73, 0x0, 0x00},
"EBREAK": {0x73, 0x0, 0x00},
"FENCE": {0x0F, 0x0, 0x00},
"FENCE.TSO": {0x0F, 0x0, 0x00},
"PAUSE": {0x0F, 0x0, 0x00},
// JALR, indirect jump/call (I-type).
"JALR": {0x67, 0x0, 0x00},
@@ -303,7 +314,14 @@ var riscvInstrTable = map[string]riscvEnc{
"FMIND": {0x53, 0x0, 0x15},
"FMAXD": {0x53, 0x1, 0x15},
// FP sign injection (double): rs2 carries the sign source.
"FSGNJD": {0x53, 0x0, 0x11},
"FSGNJD": {0x53, 0x0, 0x11},
"FSGNJS": {0x53, 0x0, 0x10},
"FSGNJX": {0x53, 0x0, 0x14},
"FSGNJXD": {0x53, 0x0, 0x15},
"FSGNJXS": {0x53, 0x0, 0x14},
"FSGNJND": {0x53, 0x1, 0x11},
"FSGNJNS": {0x53, 0x1, 0x10},
"FSGNJNX": {0x53, 0x1, 0x14},
// RV64A, load-reserved / store-conditional (funct5 0x02 / 0x03).
// The toolchain gives LR acquire ordering (aq = 1) and SC release
@@ -375,6 +393,10 @@ var riscvCvtTable = map[string]riscvCvtEnc{
"FMVDX": {0x79, 0x0, 0x53}, // int64 → float64 (bit move)
"FMVXW": {0x70, 0x0, 0x53}, // float32 → int32 (bit move)
"FMVWX": {0x78, 0x0, 0x53}, // int32 → float32 (bit move)
// The toolchain's W/D suffix spellings of the same moves.
"FMVXS": {0x70, 0x0, 0x53},
"FMVFS": {0x78, 0x0, 0x53},
"FMVSX": {0x79, 0x0, 0x53},
}
// riscvCvtType encodes an FP conversion instruction.
+169 -23
View File
@@ -33,7 +33,7 @@ func firstTextRISCV(t *testing.T, src string) *ast.Text {
// assembleRISCVHelper assembles one TEXT function and returns its code bytes.
func assembleRISCVHelper(t *testing.T, fn *ast.Text) []byte {
t.Helper()
code, _, _, _, _, err := assembleRISCV(fn)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -785,16 +785,140 @@ TEXT ·sys(SB), NOSPLIT, $0
}
}
func TestRISCV_MOV_sym_FP_error(t *testing.T) {
// MOV $sym(FP), rd should return an error (unsupported).
func TestRISCV_MOV_sym_FP(t *testing.T) {
// MOV $sym(FP), rd lowers to the frame-adjusted ADDI against SP: the
// toolchain's argframe spelling. A zero frame leaves the offset at the
// 8-byte link slot, compressed to C.ADDI4SPN.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·badfp(SB), NOSPLIT, $0
TEXT ·argfp(SB), NOSPLIT, $0
MOV $arg(FP), X10
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for MOV $arg(FP), got nil")
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// prologue (0: leaf, zero frame) + C.ADDI4SPN (2) + RET (4) = 6
want := []byte{0x28, 0x00, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_Bookkeeping(t *testing.T) {
// FUNCDATA and PCDATA contribute no bytes; UNDEF is the toolchain's
// ebreak, compressed to C.EBREAK under RVC.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·book(SB), NOSPLIT, $0-8
FUNCDATA $0, marks<>(SB)
PCDATA $1, $1
UNDEF
MOV $1, X10
MOV X10, ret+0(FP)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.EBREAK (2) + C.LI X10, 1 (2) + C.SWSP (2) + RET (4) = 10: the
// FUNCDATA and PCDATA statements contribute nothing.
want := []byte{0x02, 0x90, 0x05, 0x45, 0x2a, 0xe4, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_JMPPCRel(t *testing.T) {
// JMP N(PC): the displacement tracks the instruction N source slots
// away in the final layout (0 the jump itself, negative backwards).
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·slots(SB), NOSPLIT, $0-0
JMP 2(PC)
MOV $1, X11
MOV $2, X12
MOV X12, X11
JMP -3(PC)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// JMP 2(PC) lands on the C.MV six bytes ahead; JMP -3(PC) lands back on
// the first C.LI, six bytes behind.
want := []byte{
0x6f, 0x00, 0x60, 0x00, // JAL X0, 6
0x85, 0x45, // C.LI X11, 1
0x09, 0x46, // C.LI X12, 2
0xb2, 0x85, // C.MV X11, X12
0x6f, 0xf0, 0xbf, 0xff, // JAL X0, -6
0x67, 0x80, 0x00, 0x00, // RET
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVWideImm(t *testing.T) {
// Shift-sequence constants compress like the toolchain's expansion.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·wide(SB), NOSPLIT, $0-0
MOV $0x8000000000000000, X5
MOV $0x100000000, X5
MOV $0x000fffffffffffda, X5
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.LI -1, C.SLLI 63; C.LI 1, C.SLLI 32; C.LI -19, C.SLLI 13, SRLI 12.
want := []byte{
0xfd, 0x52, 0xfe, 0x12,
0x85, 0x42, 0x82, 0x12,
0xb5, 0x52, 0xb6, 0x02, 0x93, 0xd2, 0xc2, 0x00,
0x67, 0x80, 0x00, 0x00,
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVImmPool(t *testing.T) {
// A constant outside the shift shapes loads from the pooled $i64 data
// symbol via AUIPC+LD, named like the toolchain's pool.
src := `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0-8
MOV $0x0101010101010101, X16
MOV X16, ret+0(FP)
RET
`
f, errs := parser.Parse("pool_riscv64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
// AUIPC X16, 0 + LD X16, 0(X16): the relocation pair carries the symbol.
wantCode := []byte{0x17, 0x08, 0x00, 0x00, 0x03, 0x38, 0x08, 0x00}
if string(img.Code[0:8]) != string(wantCode) {
t.Errorf("pool load: got % x", img.Code[0:8])
}
var lit *DataSymbol
for i := range img.DataSyms {
if img.DataSyms[i].Name == "$i64.0101010101010101" {
lit = &img.DataSyms[i]
}
}
if lit == nil {
t.Fatalf("pool symbol missing: %v", img.DataSyms)
}
wantData := []byte{0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01}
if string(img.Data[lit.Offset:lit.Offset+8]) != string(wantData) {
t.Errorf("pool bytes: got % x", img.Data[lit.Offset:lit.Offset+8])
}
}
@@ -805,7 +929,7 @@ TEXT ·calltest(SB), NOSPLIT, $0
CALL ext(SB)
RET
`)
code, _, relocs, _, _, err := assembleRISCV(fn)
code, _, relocs, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -834,7 +958,7 @@ TEXT ·calllocal(SB), NOSPLIT, $0
sub:
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
_, _, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for CALL to local label, got nil")
}
@@ -868,7 +992,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
t.Helper()
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
instr := fn.Body[0].(*ast.Instr)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil, nil)
}
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
@@ -905,9 +1029,10 @@ func TestRISCVBranchJumpRange(t *testing.T) {
}
}
// TestRISCVBranchFarBody drives the range check through the full two-pass
// assembler: a forward branch over a body larger than the B-type span must
// error rather than wrap.
// TestRISCVBranchFarBody drives the relaxation pass through the full
// assembler: a forward branch over a body larger than the B-type span is
// rewritten as an inverted branch over an inserted JMP, the same layout the
// toolchain produces, instead of wrapping to a wrong target.
func TestRISCVBranchFarBody(t *testing.T) {
var sb strings.Builder
sb.WriteString("#include \"textflag.h\"\nTEXT ·far(SB), NOSPLIT, $0\n\tBEQ X10, X11, done\n")
@@ -916,8 +1041,20 @@ func TestRISCVBranchFarBody(t *testing.T) {
}
sb.WriteString("done:\n\tRET\n")
fn := firstTextRISCV(t, sb.String())
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected a branch-out-of-range error, got none")
out, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
// The relaxed branch at offset 0 targets the inserted JMP at 4 (bne
// x10, x11, +4); the JMP at 4 carries the far forward displacement.
wantBranch := wordLE(riscvBType(riscvEnc{0x63, 0x1, 0x00}, 10, 11, 4))
if !bytes.Equal(out[0:4], wantBranch) {
t.Errorf("relaxed branch = %x, want %x", out[0:4], wantBranch)
}
// done sits after 1100 ADDs: 4 + 4400, i.e. offset 4404 from the JMP at 4.
wantJmp := wordLE(riscvJType(0, 4404))
if !bytes.Equal(out[4:8], wantJmp) {
t.Errorf("inserted JMP = %x, want %x", out[4:8], wantJmp)
}
}
@@ -930,7 +1067,7 @@ TEXT ·csrhi(SB), NOSPLIT, $0
CSRRW $4096, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected an out-of-range error for CSR $4096, got none")
}
fn = firstTextRISCV(t, `#include "textflag.h"
@@ -938,25 +1075,24 @@ TEXT ·csrmax(SB), NOSPLIT, $0
CSRRW $4095, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("CSR $4095 must assemble: %v", err)
}
}
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
// span are diagnosed instead of silently truncated to their low 32 bits (the
// toolchain materialises such constants via SLLI expansion, which this
// assembler does not implement).
// span are diagnosed instead of silently truncated to their low 32 bits for
// the I-type arithmetic; the MOV forms materialise the wide constant instead
// (shift sequence or pooled load), like the toolchain.
func TestRISCV_Imm64Rejected(t *testing.T) {
cases := []string{
"MOV $0x123456789, X10",
"ADDI $0x100000000, X10, X11",
"ANDI $-0x800000001, X10, X11",
"SUB $0x100000000, X10, X11",
}
for _, src := range cases {
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Errorf("%s: expected an out-of-range error, got none", src)
}
}
@@ -969,9 +1105,19 @@ TEXT ·edge(SB), NOSPLIT, $0
SUB $0x80000000, X12, X13
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("int32-span immediates must assemble: %v", err)
}
// Beyond the span the MOV forms materialise the constant like the
// toolchain instead of diagnosing it.
fn = firstTextRISCV(t, `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0
MOV $0x123456789, X10
RET
`)
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("MOV with a 64-bit immediate must assemble: %v", err)
}
}
// riscvWants decodes code as little-endian words and pins each one; the
+221 -7
View File
@@ -61,6 +61,21 @@ const (
// vexImmRMGPR is the immediate form over general-purpose registers
// (RORX): reg = dst, rm = src, imm8 = op0, L = 0.
vexImmRMGPR
// vexRMOpGPR is the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): ModRM.reg = /digit, ModRM.rm = src
// (op0), VEX.vvvv = dst (op1), L = 0.
vexRMOpGPR
// vexCountGPR is the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): the first operand rides
// VEX.vvvv and the second is r/m, the opposite pairing of the ANDN
// family, with reg = dst (op2), L = 0.
vexCountGPR
// vexExtractGPR is the lane-extract-to-GPR form `OP $imm, xsrc, GPR/mem
// dst`: ModRM.reg = xsrc (op1), ModRM.rm = destination (op2), imm8 =
// op0, the VPEXTRB/W/D/Q layout. EVEX only; the destination never
// carries a vector length, so the register the L'L field follows is the
// XMM source.
vexExtractGPR
)
// vexSpec describes one VEX instruction's encoding parameters.
@@ -230,8 +245,34 @@ var vexTable = map[string]vexSpec{
"ANDNQ": {2, 0xF2, 1, 0, -1, vexNDS3GPR},
"MULXL": {2, 0xF6, 0, 3, -1, vexNDS3GPR},
"MULXQ": {2, 0xF6, 1, 3, -1, vexNDS3GPR},
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
// VEX.NDS.LZ.0F38, the BMI2 three-operand bit ops: BEXTR and BZHI
// share the F7/F5 opcodes across W, the variable shifts carry their
// direction in the prefix (SHLX 66, SHRX F2, SARX F3) and PDEP/PEXT
// in F2/F3.
"BEXTRL": {2, 0xF7, 0, 0, -1, vexCountGPR},
"BEXTRQ": {2, 0xF7, 1, 0, -1, vexCountGPR},
"BZHIL": {2, 0xF5, 0, 0, -1, vexCountGPR},
"BZHIQ": {2, 0xF5, 1, 0, -1, vexCountGPR},
"SARXL": {2, 0xF7, 0, 2, -1, vexCountGPR},
"SARXQ": {2, 0xF7, 1, 2, -1, vexCountGPR},
"SHLXL": {2, 0xF7, 0, 1, -1, vexCountGPR},
"SHLXQ": {2, 0xF7, 1, 1, -1, vexCountGPR},
"SHRXL": {2, 0xF7, 0, 3, -1, vexCountGPR},
"SHRXQ": {2, 0xF7, 1, 3, -1, vexCountGPR},
"PDEPL": {2, 0xF5, 0, 3, -1, vexNDS3GPR},
"PDEPQ": {2, 0xF5, 1, 3, -1, vexNDS3GPR},
"PEXTL": {2, 0xF5, 0, 2, -1, vexNDS3GPR},
"PEXTQ": {2, 0xF5, 1, 2, -1, vexNDS3GPR},
// VEX.LZ.0F38.W, the BMI1 unary bit ops (src, dst: ModRM.reg = /digit,
// rm = src, vvvv = dst).
"BLSIL": {2, 0xF3, 0, 0, 3, vexRMOpGPR},
"BLSIQ": {2, 0xF3, 1, 0, 3, vexRMOpGPR},
"BLSMSKL": {2, 0xF3, 0, 0, 2, vexRMOpGPR},
"BLSMSKQ": {2, 0xF3, 1, 0, 2, vexRMOpGPR},
"BLSRL": {2, 0xF3, 0, 0, 1, vexRMOpGPR},
"BLSRQ": {2, 0xF3, 1, 0, 1, vexRMOpGPR},
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
// VEX.128.0F.W0, mask-register test (KTESTW k1, k2: reg = dst, rm = src).
"KTESTW": {1, 0x99, 0, 0, -1, vexRM},
@@ -290,6 +331,128 @@ var vexTable = map[string]vexSpec{
"VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
"VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
"VCVTTPD2DQY": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
// --- the VEX forms the avx512enc corpus exercises alongside the EVEX
// spellings, read off the toolchain opcode tables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3},
"VANDNPD": {1, 0x55, 0, 1, -1, vexNDS3},
"VANDPD": {1, 0x54, 0, 1, -1, vexNDS3},
"VCOMISD": {1, 0x2F, 0, 1, -1, vexRM},
"VCVTSD2SS": {1, 0x5A, 0, 3, -1, vexNDS3},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM},
"VMOVNTPD": {1, 0x2B, 0, 1, -1, vexRMRev},
"VORPD": {1, 0x56, 0, 1, -1, vexNDS3},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3},
"VPCMPEQQ": {2, 0x29, 0, 1, -1, vexNDS3},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3},
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtract},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtract},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtract},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3},
"VPMULUDQ": {1, 0xF4, 0, 1, -1, vexNDS3},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3},
"VPUNPCKHQDQ": {1, 0x6D, 0, 1, -1, vexNDS3},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3},
"VSQRTPD": {1, 0x51, 0, 1, -1, vexRM},
"VSQRTSD": {1, 0x51, 0, 3, -1, vexNDS3},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3},
"VUCOMISD": {1, 0x2E, 0, 1, -1, vexRM},
// VEX.0F.WIG, the plain-prefix single/double arithmetic and unpack
// spellings (no 66 prefix; WIG, so W = 0).
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev},
// VEX.128.66.0F, the scalar and packed compare forms.
"VCOMISS": {1, 0x2F, 0, 1, -1, vexRM},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM},
// VEX.128.0F.F3/F2.W0, the high/low word shuffles ($imm, src, dst).
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM},
}
// vexSrcLen maps a source-length conversion mnemonic (the X/Y spellings of
@@ -420,6 +583,10 @@ func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
return e.encodeVexNDS3GPR(spec, ops)
case vexImmRMGPR:
return e.encodeVexImmRMGPR(spec, ops)
case vexRMOpGPR:
return e.encodeVexRMOpGPR(spec, ops)
case vexCountGPR:
return e.encodeVexCountGPR(spec, ops)
case vexRMRev:
return e.encodeVexRMRev(spec, ops)
}
@@ -523,9 +690,10 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the VEX length
// follows the destination register either way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -533,7 +701,7 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
}
vvvvBar := 15 - (dstReg.idx & 15)
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, srcReg); err != nil {
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, src); err != nil {
return err
}
immByte, err := imm8(int64(immVal))
@@ -700,7 +868,11 @@ func (e *enc) encodeVexNDS3GPR(spec vexSpec, ops []Operand) error {
if !ok || vvvvReg.isVec() {
return fmt.Errorf("VEX vvvv operand must be a general-purpose register")
}
return e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15-(vvvvReg.idx&15), src2)
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(vvvvReg.idx&15), src2)
}
// encodeVexImmRMGPR encodes the immediate form over general-purpose
@@ -729,6 +901,48 @@ func (e *enc) encodeVexImmRMGPR(spec vexSpec, ops []Operand) error {
return nil
}
// encodeVexRMOpGPR encodes the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): OP src, dst with ModRM.reg = /digit,
// ModRM.rm = src and VEX.vvvv = dst.
func (e *enc) encodeVexRMOpGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("instruction expects 2 operands (src, dst), got %d", len(ops))
}
src, dst := ops[0], ops[1]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
return e.emitVexFields(spec, 0, spec.opdigit, 0, 15-(dstReg.idx&15), src)
}
// encodeVexCountGPR encodes the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): OP src, count, dst with
// VEX.vvvv = src (op0), ModRM.rm = count (op1), ModRM.reg = dst (op2).
func (e *enc) encodeVexCountGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("VEX count instruction expects 3 operands, got %d", len(ops))
}
src, count, dst := ops[0], ops[1], ops[2]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
countReg, ok := count.(Reg)
if !ok || countReg.isVec() {
return fmt.Errorf("VEX count operand must be a general-purpose register")
}
srcReg, ok := src.(Reg)
if !ok || srcReg.isVec() {
return fmt.Errorf("VEX count source must be a general-purpose register")
}
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(srcReg.idx&15), count)
}
// encodeVexRMRev encodes the reversed two-operand form: OP src, dst with the
// vector source in ModRM.reg and the memory destination in r/m (VMOVNTDQ,
// a store with no register-destination form).
+57
View File
@@ -31,6 +31,51 @@ var x86asmUnrecognised = map[string]bool{
"RORXQ": true,
"VFMADD213SD": true,
"VFNMADD231SD": true,
// The scalar FMA spellings the decoder's tables lack entirely.
"VFMADD132SD": true,
"VFMADD132SS": true,
"VFMADD213SS": true,
"VFMADD231SD": true,
"VFMADD231SS": true,
"VFMSUB132SD": true,
"VFMSUB132SS": true,
"VFMSUB213SD": true,
"VFMSUB213SS": true,
"VFMSUB231SD": true,
"VFMSUB231SS": true,
"VFNMADD132SD": true,
"VFNMADD132SS": true,
"VFNMADD213SD": true,
"VFNMADD213SS": true,
"VFNMADD231SS": true,
"VFNMSUB132SD": true,
"VFNMSUB132SS": true,
"VFNMSUB213SD": true,
"VFNMSUB213SS": true,
"VFNMSUB231SD": true,
"VFNMSUB231SS": true,
// The BMI1 unary bit ops the decoder's AVX tables lack.
"BLSIL": true,
"BLSIQ": true,
"BLSMSKL": true,
"BLSMSKQ": true,
"BLSRL": true,
"BLSRQ": true,
// The BMI2 bit ops whose W1/LZ rows the decoder misses.
"BEXTRL": true,
"BEXTRQ": true,
"BZHIL": true,
"BZHIQ": true,
"PDEPL": true,
"PDEPQ": true,
"PEXTL": true,
"PEXTQ": true,
"SARXL": true,
"SARXQ": true,
"SHLXL": true,
"SHLXQ": true,
"SHRXL": true,
"SHRXQ": true,
}
// TestVexNDS3 encodes `mnem Y0, Y1, Y2` for every three-operand NDS
@@ -237,6 +282,18 @@ func TestVexGroundTruth(t *testing.T) {
{"MULXQ AX,BX,CX", "MULXQ", []Operand{AX, BX, CX}, "c4e2e3f6c8", ""},
{"RORXL $3,AX,CX", "RORXL", []Operand{Imm(3), AX, CX}, "c4e37bf0c803", ""},
{"RORXQ $3,AX,CX", "RORXQ", []Operand{Imm(3), AX, CX}, "c4e3fbf0c803", ""},
// BMI2 variable shifts and bit ops (three general registers).
{"SHLXL AX,CX,R15", "SHLXL", []Operand{AX, CX, vreg(t, "R15")}, "c46279f7f9", ""},
{"SHRXQ R8,DX,AX", "SHRXQ", []Operand{vreg(t, "R8"), DX, AX}, "c4e2bbf7c2", ""},
{"SARXQ AX,DX,R9", "SARXQ", []Operand{AX, DX, vreg(t, "R9")}, "c462faf7ca", ""},
{"BEXTRL AX,CX,R15", "BEXTRL", []Operand{AX, CX, vreg(t, "R15")}, "c46278f7f9", ""},
{"BZHIQ AX,CX,R15", "BZHIQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f8f5f9", ""},
{"PDEPQ AX,CX,R15", "PDEPQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f3f5f8", ""},
{"PEXTQ AX,CX,R15", "PEXTQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f2f5f8", ""},
// BMI1 unary bit ops (src, dst: /digit in ModRM.reg, dst in vvvv).
{"BLSIL AX,CX", "BLSIL", []Operand{AX, CX}, "c4e270f3d8", ""},
{"BLSRQ AX,CX", "BLSRQ", []Operand{AX, CX}, "c4e2f0f3c8", ""},
{"BLSMSKQ AX,CX", "BLSMSKQ", []Operand{AX, CX}, "c4e2f0f3d0", ""},
// Two-operand reg/rm form (v̄vvv must be 1111).
{"VPMOVSXDQ X0,Y4", "VPMOVSXDQ", []Operand{vreg(t, "X0"), vreg(t, "Y4")}, "c4e27d25e0", ""},
{"VPMOVSXWD (SI),Y0", "VPMOVSXWD", []Operand{Ptr(SI, 0, 8), vreg(t, "Y0")}, "c4e27d2306", ""},
+17 -7
View File
@@ -151,11 +151,21 @@ type Immediate struct {
// Address is a non-immediate operand: a register, a memory reference, a symbol
// reference or a label. Fields are populated best-effort from the syntax.
type Address struct {
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Range *RegRange // bracketed register range; nil for every other form
}
// RegRange is a bracketed register range, [Z0-Z3]: the amd64 spelling of
// the four-register source of the 4FMAPS/4VNNIW families. Lo and Hi carry
// the verbatim register spellings; the range is inclusive at both ends.
type RegRange struct {
Lo string
Hi string
Pos token.Position
}
+367
View File
@@ -0,0 +1,367 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"errors"
"fmt"
"go/ast"
"go/build"
"go/constant"
"go/parser"
"go/token"
"go/types"
"os"
"path/filepath"
"regexp"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
)
// go_asm.h is the header the Go compiler writes for every package that
// carries assembly (the compiler's -asmhdr output): "#define const_NAME
// value" for each package constant, and for each named struct type
// "#define TYPE__size size" plus one "#define TYPE_field offset" per field.
// GOROOT assembly includes it, and a standalone assembler has no compiler
// to have produced it, so gasm generates the equivalent itself: the package
// the .s file lives in is parsed and type-checked here, with the target
// architecture's own sizes, and the same defines are written out. The
// type-checking GOOS is selected by the caller: a GOOS-specific file
// (sys_darwin_arm64.s) needs its platform's defines, which a header from
// the ambient GOOS silently omits.
//
// The emitter mirrors cmd/compile's dumpasmhdr exactly: constants come out
// as "const_NAME", struct entries as "NAME__size" followed by the fields in
// declaration order, blank names are skipped, and float and complex
// constants are omitted (the assembler carries integers, bools and strings
// only). Aliases to structs are emitted, generic types are not: they have
// no fixed size. A define the assembly references but this header does not
// carry surfaces later as the assembler's own "undefined" diagnostic naming
// the define, which is the honest failure.
// goAsmInclude matches the #include "go_asm.h" directive, tolerant of
// whitespace, so the wiring knows which files need a generated header
// before the preprocessor runs and would report the header as missing.
var goAsmInclude = regexp.MustCompile(`(?m)^\s*#\s*include\s+"go_asm\.h"`)
// needsGoAsmHeader reports whether src includes go_asm.h.
func needsGoAsmHeader(src string) bool {
return goAsmInclude.MatchString(src)
}
// goAsmHeaderResolved reports whether the include of go_asm.h from a file in
// asmDir already resolves: to a header in the package directory itself, or
// in one of the -I directories, the way the preprocessor searches. Only an
// unresolved include is generated for; a header someone placed by hand is
// the tool the author chose, and it also wins the preprocessor's own search
// order, so generating a second copy would be dead weight at best.
func goAsmHeaderResolved(asmDir string, dirs []string) bool {
candidates := []string{filepath.Join(asmDir, "go_asm.h")}
for _, d := range dirs {
candidates = append(candidates, filepath.Join(d, "go_asm.h"))
}
for _, candidate := range candidates {
if st, err := os.Stat(candidate); err == nil && !st.IsDir() {
return true
}
}
return false
}
// generateGoAsmHeader type-checks the Go package in pkgDir for goos and
// goarch, writes its go_asm.h equivalent into dir, and returns dir. An
// empty goos means the ambient one. The caller owns the directory and its
// removal.
func generateGoAsmHeader(pkgDir, goos, goarch, dir string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
imp := newSourceImporter(goos, goarch)
if imp.sizes == nil {
return "", fmt.Errorf("go_asm.h: unknown GOARCH %q", goarch)
}
bp, err := imp.ctxt.ImportDir(pkgDir, 0)
if err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
files, errs := imp.parse(bp)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %s", goarch, pkgDir, errorList(errs))
}
_, info, errs := imp.checkPackage(bp, files)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: package does not type-check: %s", goarch, pkgDir, errorList(errs))
}
var b strings.Builder
fmt.Fprintf(&b, "// generated by gasm from package %s (GOOS %s, GOARCH %s)\n\n", bp.Name, goos, goarch)
// Files in the build's own order and declarations in source order: the
// same walk the compiler's reader makes, so the header reads the same
// way the toolchain's does. Order carries no meaning to the assembler
// (defines form a table), only to a human diffing against one.
for _, f := range files {
for _, decl := range f.Decls {
gd, ok := decl.(*ast.GenDecl)
if !ok {
continue
}
for _, spec := range gd.Specs {
switch gd.Tok {
case token.CONST:
vs, ok := spec.(*ast.ValueSpec)
if !ok {
continue
}
for _, name := range vs.Names {
emitConst(&b, info.Defs[name], name.Name)
}
case token.TYPE:
ts, ok := spec.(*ast.TypeSpec)
if !ok {
continue
}
emitStruct(&b, imp.sizes, info.Defs[ts.Name], ts.Name.Name)
}
}
}
}
if err := os.MkdirAll(dir, 0o755); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
out := filepath.Join(dir, "go_asm.h")
if err := os.WriteFile(out, []byte(b.String()), 0o644); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
return dir, nil
}
// emitConst writes one const define, skipping what the toolchain skips:
// blank names, and float and complex values the assembler has no syntax for.
func emitConst(b *strings.Builder, obj types.Object, name string) {
c, ok := obj.(*types.Const)
if !ok || name == "_" {
return
}
switch c.Val().Kind() {
case constant.Float, constant.Complex, constant.Unknown:
return
}
fmt.Fprintf(b, "#define const_%s %s\n", name, c.Val().ExactString())
}
// emitStruct writes one named struct type's size and field offsets,
// skipping what the toolchain skips: blank names, non-struct types, and
// generic types, whose size depends on their instantiation.
func emitStruct(b *strings.Builder, sizes types.Sizes, obj types.Object, name string) {
tn, ok := obj.(*types.TypeName)
if !ok || name == "_" {
return
}
t := types.Unalias(tn.Type())
// Generic types are spelled *types.Named with a type-parameter list;
// a plain struct type or an instantiated one carries none.
if named, ok := t.(*types.Named); ok && named.TypeParams().Len() > 0 {
return
}
st, ok := t.Underlying().(*types.Struct)
if !ok {
return
}
fmt.Fprintf(b, "#define %s__size %d\n", name, sizes.Sizeof(t))
fields := make([]*types.Var, st.NumFields())
for i := range st.NumFields() {
fields[i] = st.Field(i)
}
for i, off := range sizes.Offsetsof(fields) {
fld := fields[i]
if fld.Name() == "_" {
continue
}
fmt.Fprintf(b, "#define %s_%s %d\n", name, fld.Name(), off)
}
}
// errorList renders at most three errors, enough to say what is wrong
// without burying the diagnostic the caller actually reads.
func errorList(errs []error) string {
if len(errs) > 3 {
errs = errs[:3]
}
msgs := make([]string, len(errs))
for i, err := range errs {
msgs[i] = err.Error()
}
return strings.Join(msgs, "; ")
}
// sourceImporter type-checks imported packages from source with the target
// architecture's sizes. go/importer's "source" importer pins the host
// GOARCH, which would lay out imported types (internal/cpu, internal/abi)
// for the wrong target on a cross-architecture header, so the recursion is
// carried here with one build context and one sizes instance per
// architecture.
type sourceImporter struct {
fset *token.FileSet
ctxt *build.Context
sizes types.Sizes
pkgs map[string]*types.Package
}
// newSourceImporter returns the importer for one target GOOS and GOARCH.
// Cgo is disabled so the file set is deterministic and independent of the
// host's C toolchain: cgo-tagged files drop out of the build exactly as
// they do from a CGO_ENABLED=0 build, whose assembly is what gasm targets.
func newSourceImporter(goos, goarch string) *sourceImporter {
ctxt := new(build.Context)
*ctxt = build.Default
ctxt.GOOS = goos
ctxt.GOARCH = goarch
ctxt.CgoEnabled = false
return &sourceImporter{
fset: token.NewFileSet(),
ctxt: ctxt,
sizes: types.SizesFor("gc", goarch),
pkgs: map[string]*types.Package{},
}
}
// Import type-checks one imported package and memoises it. "unsafe" must
// resolve to go/types' own package, never to the source in GOROOT/src/unsafe:
// the source declares Sizeof and Offsetof as ordinary functions over
// ArbitraryType, and checking against that signature rejects half the
// unsafe arithmetic the gc compiler accepts, which is exactly the divergence
// srcimporter guards against the same way.
func (im *sourceImporter) Import(path string) (*types.Package, error) {
if path == "unsafe" {
return types.Unsafe, nil
}
if p, ok := im.pkgs[path]; ok {
return p, nil
}
bp, err := im.ctxt.Import(path, "", 0)
if err != nil {
return nil, err
}
files, errs := im.parse(bp)
if len(errs) > 0 {
return nil, errors.New(errorList(errs))
}
pkg, _, _ := im.checkPackage(bp, files)
im.pkgs[path] = pkg
return pkg, nil
}
// parse reads the build package's Go files. Import-level failures (no Go
// files for the target, unreadable files) come back as errors, and the
// type-check decides the rest.
func (im *sourceImporter) parse(bp *build.Package) ([]*ast.File, []error) {
if len(bp.GoFiles) == 0 {
return nil, []error{fmt.Errorf("no Go source files for GOOS=%s GOARCH=%s", im.ctxt.GOOS, im.ctxt.GOARCH)}
}
var (
files []*ast.File
errs []error
)
for _, name := range bp.GoFiles {
f, err := parser.ParseFile(im.fset, filepath.Join(bp.Dir, name), nil, parser.SkipObjectResolution)
if err != nil {
errs = append(errs, err)
continue
}
files = append(files, f)
}
return files, errs
}
// checkPackage type-checks one package's files with the importer's sizes,
// recording every error: a header from a package that does not type-check
// could silently mis-state an offset, so the caller refuses the header
// rather than trusting it. The returned Defs map backs the root package's
// emission walk; imports only need the checked package itself.
func (im *sourceImporter) checkPackage(bp *build.Package, files []*ast.File) (*types.Package, *types.Info, []error) {
var errs []error
conf := &types.Config{
Importer: im,
Sizes: im.sizes,
Error: func(err error) { errs = append(errs, err) },
}
info := &types.Info{Defs: map[*ast.Ident]types.Object{}}
pkg, _ := conf.Check(bp.ImportPath, im.fset, files, info)
return pkg, info, errs
}
// asmhdrCache generates one go_asm.h per package directory and target
// architecture under one temp root, for callers that assemble many files
// (the corpus audit). Failures are cached too: a package that does not
// type-check must not be re-checked once per file.
type asmhdrCache struct {
root string
dirs map[string]string // "pkgDir\x00goos\x00goarch" -> directory holding go_asm.h
errs map[string]error
}
func newAsmhdrCache() (*asmhdrCache, error) {
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return nil, err
}
return &asmhdrCache{root: root, dirs: map[string]string{}, errs: map[string]error{}}, nil
}
// dirFor returns the directory holding the generated go_asm.h for pkgDir
// under goos and goarch, generating it on first use. An empty goos means
// the ambient one, resolved here so that one package cannot generate twice
// under an explicit and an implicit spelling of the same GOOS.
func (c *asmhdrCache) dirFor(pkgDir, goos, goarch string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
key := pkgDir + "\x00" + goos + "\x00" + goarch
if dir, ok := c.dirs[key]; ok {
return dir, nil
}
if err, ok := c.errs[key]; ok {
return "", err
}
dir := filepath.Join(c.root, fmt.Sprintf("h%d_%s_%s", len(c.dirs), goos, goarch))
if _, err := generateGoAsmHeader(pkgDir, goos, goarch, dir); err != nil {
c.errs[key] = err
return "", err
}
c.dirs[key] = dir
return dir, nil
}
// close removes the temp root.
func (c *asmhdrCache) close() { os.RemoveAll(c.root) }
// ensureGoAsmHeader prepares the include directory a file that includes
// go_asm.h needs: the generated header for the package in path's directory,
// for the file's target GOOS and architecture. It reports a usage error
// when the architecture cannot be determined, and passes through the
// generator's diagnostics, which name the package.
func ensureGoAsmHeader(path string, target arch.Arch, goos string, cache *asmhdrCache) (string, func(), error) {
if path == "-" {
return "", nil, errors.New("cannot generate go_asm.h for standard input (no package directory)")
}
if target == arch.Unknown {
return "", nil, errors.New("a file that includes go_asm.h needs a target architecture: name the file _<arch>.s or pass -GOARCH")
}
if cache != nil {
dir, err := cache.dirFor(filepath.Dir(path), goos, goarchName(target))
return dir, func() {}, err
}
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return "", nil, err
}
dir, err := generateGoAsmHeader(filepath.Dir(path), goos, goarchName(target), root)
if err != nil {
os.RemoveAll(root)
return "", nil, err
}
return dir, func() { os.RemoveAll(root) }, nil
}
+430
View File
@@ -0,0 +1,430 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
)
// writePkg lays out a minimal Go package in a temp directory.
func writePkg(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, src := range files {
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
// generateFor generates the header for dir and returns its text. An empty
// goos means the ambient one.
func generateFor(t *testing.T, dir, goos, goarch string) string {
t.Helper()
hdrDir, err := generateGoAsmHeader(dir, goos, goarch, t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(%q, %s, %s): %v", dir, goos, goarch, err)
}
b, err := os.ReadFile(filepath.Join(hdrDir, "go_asm.h"))
if err != nil {
t.Fatal(err)
}
return string(b)
}
func TestGenerateGoAsmHeaderShape(t *testing.T) {
dir := writePkg(t, map[string]string{"sample.go": `package sample
const bufSize = 1024
const (
a = iota * 8
b
c
)
const (
strConst = "hello"
boolConst = true
floatConst = 1.5
_ = "the blank identifier is skipped"
)
const shift = 1 << 20
type reader struct {
r int64
w int64
_ [4]byte
name string
}
type scalar int
type aliased struct {
k uint32
v uint32
}
type alias = aliased
`})
hdr := generateFor(t, dir, "", "amd64")
want := []string{
"#define const_bufSize 1024",
// iota resolves through go/types, one define per name.
"#define const_a 0",
"#define const_b 8",
"#define const_c 16",
`#define const_strConst "hello"`,
"#define const_boolConst true",
// Floats are the toolchain's own skip, as are blank names.
"#define const_shift 1048576",
// The blank field still occupies its bytes: the pad after w runs to
// the string's 8-byte alignment.
"#define reader__size 40",
"#define reader_r 0",
"#define reader_w 8",
"#define reader_name 24",
// Non-struct named types carry no defines; aliases to structs do.
"#define aliased__size 8",
"#define aliased_k 0",
"#define aliased_v 4",
"#define alias__size 8",
"#define alias_k 0",
"#define alias_v 4",
}
for _, w := range want {
if !strings.Contains(hdr, w+"\n") {
t.Errorf("header misses %q\ngot:\n%s", w, hdr)
}
}
for _, banned := range []string{"#define const_floatConst", "#define _ ", "#define scalar"} {
if strings.Contains(hdr, banned) {
t.Errorf("header must not carry %s\ngot:\n%s", banned, hdr)
}
}
}
func TestGenerateGoAsmHeaderPerArch(t *testing.T) {
dir := writePkg(t, map[string]string{
"common.go": `package perarch
type layout struct {
a int32
p uintptr
}
`,
// The build-tagged file set is part of the contract: a per-arch
// package is exactly how internal/cpu declares its layouts.
"const_amd64.go": `//go:build amd64
package perarch
const flavour = 1
`,
"const_arm64.go": `//go:build arm64
package perarch
const flavour = 2
`,
})
amd64 := generateFor(t, dir, "", "amd64")
arm64 := generateFor(t, dir, "", "arm64")
if !strings.Contains(amd64, "#define const_flavour 1\n") {
t.Errorf("amd64 header misses const_flavour 1:\n%s", amd64)
}
if !strings.Contains(arm64, "#define const_flavour 2\n") {
t.Errorf("arm64 header misses const_flavour 2:\n%s", arm64)
}
if strings.Contains(arm64, "#define const_flavour 1\n") {
t.Errorf("arm64 header must not carry the amd64 file's value")
}
// SizesFor makes the layout the target's: uintptr is 4 bytes wide on
// 386 and 8 on amd64, which must move p and grow the struct.
if !strings.Contains(amd64, "#define layout__size 16\n") || !strings.Contains(amd64, "#define layout_p 8\n") {
t.Errorf("amd64 layout wrong:\n%s", amd64)
}
w386 := generateFor(t, dir, "", "386")
if !strings.Contains(w386, "#define layout__size 8\n") || !strings.Contains(w386, "#define layout_p 4\n") {
t.Errorf("386 layout wrong:\n%s", w386)
}
}
// TestGenerateGoAsmHeaderGOOS pins the GOOS half of the target: only the
// platform's own files type-check into the header, which is why
// sys_darwin_arm64.s cannot assemble against a linux-generated one.
func TestGenerateGoAsmHeaderGOOS(t *testing.T) {
dir := writePkg(t, map[string]string{
"common.go": `package goosaware
type shared struct {
a int32
}
`,
"plat_darwin.go": `//go:build darwin
package goosaware
type platform struct {
trampoline_numer int64
}
`,
"plat_windows.go": `//go:build windows
package goosaware
type platform struct {
callbackArgs__size int32
}
`,
})
darwin := generateFor(t, dir, "darwin", "arm64")
if !strings.Contains(darwin, "#define platform__size 8\n") || !strings.Contains(darwin, "#define platform_trampoline_numer 0\n") {
t.Errorf("darwin header misses the darwin layout:\n%s", darwin)
}
if strings.Contains(darwin, "callbackArgs") {
t.Errorf("darwin header must not carry the windows layout:\n%s", darwin)
}
windows := generateFor(t, dir, "windows", "arm64")
if !strings.Contains(windows, "#define platform_callbackArgs__size 0\n") {
t.Errorf("windows header misses the windows layout:\n%s", windows)
}
if strings.Contains(windows, "trampoline_numer") {
t.Errorf("windows header must not carry the darwin layout:\n%s", windows)
}
// The ambient GOOS is neither of the two, so only shared's defines are
// emitted; the shared type keeps its layout there.
ambient := generateFor(t, dir, "", "arm64")
if !strings.Contains(ambient, "#define shared__size 4\n") {
t.Errorf("ambient header misses the shared layout:\n%s", ambient)
}
if strings.Contains(ambient, "#define platform_") {
t.Errorf("ambient header must not carry either platform layout:\n%s", ambient)
}
}
func TestGoosFromFilename(t *testing.T) {
for path, want := range map[string]string{
"/x/sys_darwin_arm64.s": "darwin",
"/x/sys_windows_arm64.s": "windows",
"/x/asm_linux_amd64.s": "linux",
"/x/rt0_darwin_arm64.s": "darwin",
"/x/vgetrandom_zos_s390x.s": "zos",
"/x/rt0_js_wasm.s": "js",
"/x/memmove_amd64.s": "",
"/x/vlop_arm.s": "",
"/x/stubs.s": "",
} {
if got := goosFromFilename(path); got != want {
t.Errorf("goosFromFilename(%q) = %q, want %q", path, got, want)
}
}
}
func TestGenerateGoAsmHeaderErrors(t *testing.T) {
t.Run("type error", func(t *testing.T) {
dir := writePkg(t, map[string]string{"bad.go": `package bad
const x = undefinedIdent
`})
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail for a package that does not type-check")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
if !strings.Contains(err.Error(), "type-check") {
t.Errorf("error must say the package does not type-check: %v", err)
}
})
t.Run("no go files", func(t *testing.T) {
dir := t.TempDir()
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail without Go files")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
})
}
func TestNeedsGoAsmHeader(t *testing.T) {
yes := "#include \"go_asm.h\"\n#include \"textflag.h\"\n"
no := "#include \"textflag.h\"\n#include \"funcdata.h\"\n"
if !needsGoAsmHeader(yes) {
t.Error("needsGoAsmHeader(missing on a go_asm.h include)")
}
if needsGoAsmHeader(no) {
t.Error("needsGoAsmHeader claims other headers need generation")
}
}
func TestGoAsmHeaderResolved(t *testing.T) {
dir := t.TempDir()
if goAsmHeaderResolved(dir, nil) {
t.Error("resolved with no header anywhere")
}
other := t.TempDir()
if goAsmHeaderResolved(dir, []string{other}) {
t.Error("resolved with an empty -I directory")
}
if err := os.WriteFile(filepath.Join(dir, "go_asm.h"), nil, 0o644); err != nil {
t.Fatal(err)
}
if !goAsmHeaderResolved(dir, nil) {
t.Error("not resolved with the header in the package directory")
}
}
func TestOtherGOOSFile(t *testing.T) {
for path, want := range map[string]bool{
"/x/sys_windows_amd64.s": true,
"/x/rt0_js_wasm.s": true,
"/x/sys_darwin_arm64.s": true,
"/x/sys_linux_amd64.s": false,
"/x/time_linux_amd64.s": false,
"/x/memmove_amd64.s": false,
"/x/generic.s": false,
} {
if got := otherGOOSFile(path); got != want {
t.Errorf("otherGOOSFile(%q) = %v, want %v", path, got, want)
}
}
}
// TestRunCorpusAuditGoAsm covers the audit wiring end to end: a package
// beside its kernel, the kernel living off the generated defines, and the
// histogram recording a generation failure as its own reason.
func TestRunCorpusAuditGoAsm(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("pkg.go", `package corpus
const pageSize = 4096
type header struct {
magic uint64
flags uint64
}
`)
write("kern_amd64.s", "#include \"go_asm.h\"\nTEXT \xc2\xb7f(SB), NOSPLIT, $0-16\n\tMOVQ\t$const_pageSize, AX\n\tMOVQ\t$header__size, BX\n\tRET\n")
// The defines live in the file's own package; a kernel in a directory
// without Go files has no package to generate from.
if err := os.MkdirAll(filepath.Join(dir, "sub"), 0o755); err != nil {
t.Fatal(err)
}
write(filepath.Join("sub", "lonely_arm64.s"), "#include \"go_asm.h\"\nTEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
get := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
if a := get("amd64"); a.attempted != 1 || a.assembled != 1 {
t.Errorf("amd64 = %d/%d, want 1/1", a.assembled, a.attempted)
}
// lonely_arm64.s is an arm64 file whose package cannot be generated.
if a := get("arm64"); a.attempted != 1 || a.assembled != 0 {
t.Errorf("arm64 = %d/%d, want 0/1", a.assembled, a.attempted)
}
if r := get("arm64").reasons["go_asm.h generation failed"]; r != 1 {
t.Errorf("arm64 go_asm.h failure count = %d, want 1", r)
}
}
// TestRunCorpusAuditGOOS covers the filename-derived GOOS end to end: a
// kernel whose name names darwin must have its header type-checked with
// GOOS=darwin, so the darwin-only constant it offsets with is defined. The
// operand mirrors sys_darwin_arm64.s's trampoline, where a missing define
// leaves an unexpanded symbol in the offset and fails.
func TestRunCorpusAuditGOOS(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("pkg.go", "package corpus\n")
write("plat_darwin.go", "//go:build darwin\n\npackage corpus\n\nconst trampolineNumer = 8\n")
write("kern_darwin_arm64.s", "#include \"go_asm.h\"\n"+
"GLOBL timebase<>(SB), NOPTR, $16\n"+
"TEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n"+
"\tMOVD\ttimebase<>+const_trampolineNumer(SB), R0\n"+
"\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
var arm *corpusTally
for i, tg := range stats.targets {
if tg.name == "arm64" {
arm = stats.tallies[i]
}
}
if arm == nil {
t.Fatal("no arm64 tally")
}
if arm.attempted != 1 || arm.assembled != 1 {
t.Errorf("arm64 = %d/%d, want 1/1; reasons: %v", arm.assembled, arm.attempted, arm.reasons)
}
}
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
// the runtime package of the ambient toolchain, whose header the toolchain's
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
// the whole package. The GOROOT comes from the go command itself, so the
// test follows whatever toolchain the host provides.
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
if testing.Short() {
t.Skip("type-checks the whole runtime package")
}
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
t.Skipf("no Go toolchain: %v", err)
}
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
}
b, err := os.ReadFile(dir + "/go_asm.h")
if err != nil {
t.Fatal(err)
}
hdr := string(b)
for _, want := range []string{
"#define const_hashSize 8\n",
"#define const_avxSupported 1\n",
"#define const_pageSize 8192\n",
"#define g_stackguard0 16\n",
"#define m__size ",
} {
if !strings.Contains(hdr, want) {
t.Errorf("runtime header misses %q", want)
}
}
}
+221 -21
View File
@@ -5,6 +5,7 @@ package main
import (
"fmt"
"maps"
"os"
"os/exec"
"path/filepath"
@@ -37,7 +38,7 @@ import (
// construction and are excluded from the diff; the other architectures list
// their conditional branches outright.
func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]", `
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
@@ -54,14 +55,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
the encodability backlog by frequency rather than by table order. With
-list the report also prints every failing file with its reason, per
architecture.
`)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
if err := fs.Parse(args); err != nil {
return err
}
if *corpus {
return cmdAuditCorpus(fs.Args())
return cmdAuditCorpus(fs.Args(), dirs, *list)
}
archName := "amd64"
switch n := len(fs.Args()); {
@@ -286,6 +292,11 @@ func probeShapes(a arch.Arch) []string {
"V1.B16, [V2.B16], V3.B16", "V1.B8, [V2.B16, V3.B16], V4.B8",
"$4, V1.B16, V2.B16, V3.B16", "$15, V1", "V1, V2, p2",
"R0, R1, $1, $4, p2",
// The landing-pad kind, the compiler's PCDATA
// bookkeeping and the four-operand bitfield
// insert/extract family, as the toolchain's own
// testdata spells them.
"C", "$1, $0", "$0, R1, $1, R2",
}
case arch.RISCV:
return []string{
@@ -305,6 +316,9 @@ func probeShapes(a arch.Arch) []string {
"X5, X6, p2", "R5, R6, p2",
"X5, E8, M8, TA, MA, X6", "$4, E32, M1, TA, MA, X1",
"(X5), X6, V1, V2",
// The CSR immediate forms the toolchain's testdata spells:
// immediate, CSR name, destination.
"$2, TIME, X5",
"",
}
case arch.LOONG64:
@@ -321,6 +335,12 @@ func probeShapes(a arch.Arch) []string {
"V1, V2, V3", "X1, X2, X3", "V1, V2", "X1, X2", "V1", "X1",
// The vector compare-to-flag forms land in an FCC register.
"V1, FCC0", "X1, FCC0",
// The compiler's bookkeeping pair and the raw spellings the
// toolchain's own testdata carries: JIRL rd, rj, offset (the
// form RET lowers to), the prefetch with a 32-bit address and
// hint, and the byte-shuffle quads.
"$1, $0", "R1, R5, 0", "0(R7), $5, $0", "(R7), $5, $0",
"V1, V2, V3, V4", "X1, X2, X3, X4",
"",
}
}
@@ -386,17 +406,30 @@ type corpusTally struct {
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
fails []corpusFailure // every failure, in file order, for --list
}
func (t *corpusTally) fail(path, reason string) {
// corpusFailure is one failed attempt, recorded for the --list report.
type corpusFailure struct {
path string
reason string
detail string
}
func (t *corpusTally) fail(path string, err error) {
reason := corpusReason(err)
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
}
// cmdAuditCorpus implements audit-instructions --corpus.
func cmdAuditCorpus(args []string) error {
// cmdAuditCorpus implements audit-instructions --corpus. The include
// directories carry #include resolution over a corpus whose files refer to
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
// needs.
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
@@ -410,11 +443,31 @@ func cmdAuditCorpus(args []string) error {
}
root = filepath.Join(strings.TrimSpace(string(out)), "src")
}
stats, err := runCorpusAudit(root)
// The toolchain's shipped headers (funcdata.h and friends) define the
// macros GOROOT files include; a corpus audit measures those files, so
// the header directory joins the search path automatically. go_asm.h
// is compiler-generated per package, so it is not resolved from here:
// files that include it get one generated per target architecture,
// which runCorpusAudit arranges.
if out, err := exec.Command("go", "env", "GOROOT").Output(); err == nil {
pkgInclude := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
if fi, err := os.Stat(pkgInclude); err == nil && fi.IsDir() {
seen := false
for _, d := range dirs {
if d == pkgInclude {
seen = true
}
}
if !seen {
dirs = append(dirs, pkgInclude)
}
}
}
stats, err := runCorpusAudit(root, dirs)
if err != nil {
return err
}
printCorpusStats(stats)
printCorpusStats(stats, list)
return nil
}
@@ -435,13 +488,19 @@ type corpusStats struct {
// set, even when gasm does not support the architecture.
var goPortSuffixes = []string{
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
"s390x", "wasm",
"mips64le", "mipsle", "mips64x", "mipsx", "ppc64", "ppc64le",
"ppc64x", "riscv", "riscv64", "s390x", "wasm",
}
// otherPortFile reports whether the file's name carries a Go-architecture
// suffix gasm does not support.
// otherPortFile reports whether the file belongs to a build no supported
// target ever compiles: either its name carries a Go-architecture suffix
// gasm does not support, or, for a file with no architecture suffix at all,
// it names another GOOS, which go/build drops from the file set
// (rt0_js_wasm.s is a javascript build, not a generic one).
func otherPortFile(path string) bool {
if otherGOOSFile(path) {
return true
}
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
@@ -454,7 +513,67 @@ func otherPortFile(path string) bool {
return false
}
func runCorpusAudit(root string) (*corpusStats, error) {
// goOSNames are the GOOS values go/build recognises in file names.
var goOSNames = map[string]bool{
"aix": true, "android": true, "darwin": true, "dragonfly": true,
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
"js": true, "linux": true, "nacl": true, "netbsd": true,
"openbsd": true, "plan9": true, "solaris": true, "wasip1": true,
"windows": true, "zos": true,
}
// resolveGOOS validates a -GOOS flag value, mirroring the architecture
// check's surface: a usage error naming what the tool accepts.
func resolveGOOS(name string) (string, error) {
lower := strings.ToLower(name)
if goOSNames[lower] {
return lower, nil
}
return "", &usageError{fmt.Errorf("unknown GOOS %q: want one of %s", name, strings.Join(slices.Sorted(maps.Keys(goOSNames)), ", "))}
}
// goosFromFilename returns the GOOS the file's name carries, by go/build's
// goodOSArchFile rule: the GOOS segment sits last, or last before the
// architecture segment (sys_darwin_arm64.s, vlop_arm.s carries none). An
// empty result means the name names no GOOS and the ambient one applies.
func goosFromFilename(path string) string {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
base = strings.TrimSuffix(base, ".s")
// go/build ignores everything before the first underscore, so a GOOS
// segment is only ever looked for from there on.
i := strings.IndexByte(base, '_')
if i < 0 {
return ""
}
segs := strings.Split(base[i:], "_")
if n := len(segs); n >= 2 && goOSNames[segs[n-2]] && slices.Contains(goPortSuffixes, segs[n-1]) {
return segs[n-2]
}
if goOSNames[segs[len(segs)-1]] {
return segs[len(segs)-1]
}
return ""
}
// otherGOOSFile reports whether the file's name names a GOOS other than the
// host's, by go/build's file-name rules.
func otherGOOSFile(path string) bool {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
for seg := range strings.SplitSeq(strings.TrimSuffix(base, ".s"), "_") {
if goOSNames[seg] && seg != runtime.GOOS {
return true
}
}
return false
}
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
return nil, err
@@ -474,12 +593,26 @@ func runCorpusAudit(root string) (*corpusStats, error) {
// its name allows assembles it.
full, generic, otherPort := 0, 0, 0
// Header generation is created on first use, so a corpus with no
// go_asm.h includes never pays for a temp directory.
var hdr *asmhdrCache
defer func() {
if hdr != nil {
hdr.close()
}
}()
for _, path := range files {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
// The GOOS the header generation type-checks under follows the
// file's name when the name carries one; the ambient GOOS is the
// honest guess otherwise (a build tag naming another GOOS is
// invisible to a file-name rule).
goos := goosFromFilename(path)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
@@ -490,10 +623,11 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
} else if otherPortFile(path) {
// A file named for a Go port gasm does not support (arm,
// 386, s390x, ...) is compiled by no supported-arch build,
// so it is neither generic nor a per-arch attempt: counting
// it as generic would make the headline unreachably low
// for reasons no supported target can fix.
// 386, s390x, ...) or for another GOOS is compiled by no
// supported-arch build, so it is neither generic nor a
// per-arch attempt: counting it as generic would make the
// headline unreachably low for reasons no supported target
// can fix.
otherPort++
} else {
generic++
@@ -502,19 +636,76 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
}
// A file that includes go_asm.h parses against a per-target header:
// the defines differ per architecture (internal/cpu's layout, for
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
// for another), so the parse cannot be shared the way a
// header-free file's can. A generation failure is a failure for
// every target, named for the package rather than a bare "include
// not found". A header already resolvable in the package
// directory or the -I list is left alone.
if len(wanted) > 0 && needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
if hdr == nil {
if hdr, err = newAsmhdrCache(); err != nil {
return nil, err
}
}
pkgDir := filepath.Dir(path)
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
if err != nil {
ok = false
t.fail(path, err)
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: append(slices.Clone(dirs), hdrDir),
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
if len(errs) > 0 {
ok = false
t.fail(path, errs[0])
continue
}
if _, err := assembleFile(tg.a, f, goos); err != nil {
ok = false
t.fail(path, err)
continue
}
t.assembled++
}
if ok && len(wanted) > 0 {
full++
}
continue
}
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
// The parse carries the target's platform predefines, so it
// cannot be shared across targets the way a header-free file's
// could: a #ifdef GOARCH_arm block must be live on arm64 and
// dead everywhere else.
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: dirs,
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
_, err = assembleFile(tg.a, f, goos)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -536,7 +727,7 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
func printCorpusStats(s *corpusStats, list bool) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
// The rate is over the files a supported build would attempt: the
// other ports' files sit in the count for completeness but can never
@@ -551,6 +742,13 @@ func printCorpusStats(s *corpusStats) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
if !list {
continue
}
for _, f := range t.fails {
fmt.Printf(" FAIL %s\n", f.path)
fmt.Printf(" %s: %s\n", f.reason, f.detail)
}
}
}
@@ -558,6 +756,8 @@ func printCorpusStats(s *corpusStats) {
func corpusReason(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "go_asm.h for GOARCH"):
return "go_asm.h generation failed"
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
return "instruction not encodable"
case strings.Contains(msg, "undefined label"):
+1 -1
View File
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
if len(errs) > 0 {
return 1
}
img, err := assembleFile(target, f)
img, err := assembleFile(target, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 1
+94 -20
View File
@@ -240,6 +240,16 @@ func readSource(path string) (string, error) {
return string(b), err
}
// includeDirs collects repeatable -I flags: the directories searched for
// #include files during macro expansion and include splicing.
type includeDirs []string
func (d *includeDirs) String() string { return strings.Join(*d, ",") }
func (d *includeDirs) Set(v string) error {
*d = append(*d, v)
return nil
}
func cmdTokens(args []string) int {
fs := newCommand("tokens", "gasm tokens <file>", `
Print the lexical token stream of FILE: position, token kind and text, one
@@ -476,7 +486,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
}
func cmdAsm(args []string) int {
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>", `
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>", `
Assemble FILE without the Go toolchain: every TEXT function is encoded to
machine code and printed as a hex dump. Supported architectures: amd64
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
@@ -493,14 +503,26 @@ system toolchain; goobj emits the Go toolchain's own object format, which
cmd/link consumes directly (it requires -p, the package path, and the
installed Go toolchain: the object preamble is captured from go tool asm
and the format version from go version).
A file that includes go_asm.h gets that header generated automatically from
the package it lives in (the .go files beside it, type-checked for the
target architecture, the toolchain's own defines), so GOROOT assembly
assembles without a compiler. -GOOS selects the type-checking GOOS for
that header: a GOOS-specific file (sys_darwin_arm64.s) needs its platform's
defines, which a header from the ambient GOOS silently omits. A package
that has no Go files for the target or does not type-check is a hard error
naming the package.
`)
out := fs.String("o", "", "write the output to this file")
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
goosName := fs.String("GOOS", "", "operating system for go_asm.h generation: a GOOS go/build recognises (default: the host's)")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
fs.Parse(args)
if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>")
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>")
return 2
}
// The format is validated before anything else, so a bogus value exits 2
@@ -521,12 +543,38 @@ and the format version from go version).
}
targetArch = a
}
goos := ""
if *goosName != "" {
g, err := resolveGOOS(*goosName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
return 2
}
goos = g
}
src, err := readSource(path)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm:", err)
return 1
}
f, errs := parser.Parse(path, src)
// A file that includes go_asm.h cannot assemble without the package's
// defines, and without a compiler nothing else has generated them, so
// gasm produces the equivalent itself: automatic, because the compiler
// behaves the same way and a flag would only ever be forgotten. A
// generation failure is fatal and names the package: assembling against
// a missing header would fail later with a bare "undefined" instead.
// A go_asm.h that already resolves (placed by hand, or passed with -I)
// is left alone.
if needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
hdrDir, cleanup, err := ensureGoAsmHeader(path, targetArch, goos, nil)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm asm:", err)
return 1
}
defer cleanup()
dirs = append(dirs, hdrDir)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
@@ -534,7 +582,7 @@ and the format version from go version).
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, goos)
if err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1
@@ -634,7 +682,7 @@ and the format version from go version).
// cmdDiff compares the machine code of two assembly files.
func cmdDiff(args []string) int {
set := newCommand("diff", "gasm diff [-GOARCH arch] <file1.s> <file2.s>", `
set := newCommand("diff", "gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>", `
Compare the machine code produced by assembling two files.
Shows which functions differ and the byte-level differences.
Useful for verifying that two implementations produce identical code,
@@ -645,9 +693,11 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
`)
mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names")
archName := set.String("GOARCH", "", "target architecture for both files: amd64, arm64, riscv64 or loong64")
var dirs includeDirs
set.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
set.Parse(args)
if set.NArg() != 2 {
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>")
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>")
return 2
}
path1, path2 := set.Arg(0), set.Arg(1)
@@ -675,12 +725,12 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
}
// Assemble both files.
img1, err := assemblePath(path1, forced)
img1, err := assemblePath(path1, forced, dirs)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err)
return 1
}
img2, err := assemblePath(path2, forced)
img2, err := assemblePath(path2, forced, dirs)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err)
return 1
@@ -739,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
return 1
}
// platformPredefines mirrors the go command's assembler invocation, which
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
// exactly those names, so an assembler without them cannot see the platform
// definitions at all.
func platformPredefines(goarch, goos string) map[string]string {
return map[string]string{
"GOARCH_" + goarch: "1",
"GOOS_" + goos: "1",
}
}
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
// and nothing else.
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
goos := fileGoos
if goos == "" {
goos = runtime.GOOS
}
return platformPredefines(goarch, goos)
}
// assembleFile assembles a parsed file for the given architecture and returns the image.
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
switch targetArch {
case arch.AMD64:
return asm.AssembleFile(f)
return asm.AssembleFile(f, asm.WithGOOS(goos))
case arch.RISCV:
return asm.AssembleFileRISCV(f)
case arch.ARM64:
@@ -755,25 +828,26 @@ func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
}
}
// assemblePath reads, parses and assembles a file (used by cmdDiff). A
// non-Unknown forced architecture overrides the file-name suffix.
func assemblePath(path string, forced arch.Arch) (*asm.Image, error) {
// assemblePath reads, preprocesses, parses and assembles a file (used by
// cmdDiff). A non-Unknown forced architecture overrides the file-name
// suffix.
func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image, error) {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return nil, fmt.Errorf("parse errors")
}
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
return assembleFile(target, f, "")
}
// printByteDiff shows the first few byte differences between two code blocks.
@@ -869,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
if len(errs) > 0 {
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
+1 -1
View File
@@ -403,7 +403,7 @@ func TestRunCorpusAudit(t *testing.T) {
write("generic.s", "#include \"textflag.h\"\nTEXT ·g(SB), NOSPLIT, $0-0\n\tRET\n")
write("broken.s", "#include \"textflag.h\"\nTEXT ·b(SB), NOSPLIT, $0-0\n\tJMP nowhere\n\tRET\n")
stats, err := runCorpusAudit(dir)
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
+109
View File
@@ -0,0 +1,109 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"path/filepath"
"strings"
"testing"
)
// writeTree writes a directory of files and returns its root.
func writeTree(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, content := range files {
path := filepath.Join(dir, name)
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
// TestAsmMacroAndIncludeEndToEnd drives `gasm asm` over a source with an
// in-file parameterised macro and an include resolved through -I, and checks
// the assembled bytes came from the expansion (the loop body counts six
// increments, two per expanded iteration).
func TestAsmMacroAndIncludeEndToEnd(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
dir := writeTree(t, map[string]string{
"inc/consts.h": "#define NITER 3\n",
"main_amd64.s": "#include \"textflag.h\"\n" +
"#include \"consts.h\"\n" +
"#define STEP(r) ADDQ $1, r; ADDQ $1, r\n" +
"TEXT ·f(SB), NOSPLIT, $0-8\n" +
"\tXORQ AX, AX\n" +
"\tMOVQ $NITER, CX\n" +
"loop:\n" +
"\tSTEP(AX)\n" +
"\tDECQ CX\n" +
"\tJNZ loop\n" +
"\tMOVQ AX, ret+0(FP)\n" +
"\tRET\n",
})
stdout, stderr, code := capture(func() int {
return cmdAsm([]string{"-I", filepath.Join(dir, "inc"), "-GOARCH", "amd64", filepath.Join(dir, "main_amd64.s")})
})
if code != 0 {
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
}
// The macro expanded to two ADDQ $1 encodings in the static body; the
// iteration count lives in the runtime loop.
if n := strings.Count(stdout, "83 c0 01"); n != 2 {
t.Errorf("found %d ADDQ $1 encodings in the image, want 2:\n%s", n, stdout)
}
}
// TestAsmIncludeResolutionOrder pins the -I search order end to end: the
// including file's directory wins over the -I directories.
func TestAsmIncludeResolutionOrder(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
dir := writeTree(t, map[string]string{
"src/main_amd64.s": "#include \"textflag.h\"\n" +
"#include \"vals.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\tMOVQ $VAL, AX\n" +
"\tRET\n",
"src/vals.h": "#define VAL 1\n",
"late/vals.h": "#define VAL 2\n",
"early/vals.h": "#define VAL 3\n",
})
stdout, stderr, code := capture(func() int {
return cmdAsm([]string{"-I", filepath.Join(dir, "early"), "-I", filepath.Join(dir, "late"),
"-GOARCH", "amd64", filepath.Join(dir, "src", "main_amd64.s")})
})
if code != 0 {
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
}
// VAL came from src/vals.h, not from either -I directory: the image
// loads the immediate 1.
if !strings.Contains(stdout, "b8 01 00 00 00") {
t.Errorf("expected the source-directory VAL (immediate 1) in:\n%s", stdout)
}
}
// TestAsmMissingIncludeIsAnError pins the diagnostic for an include that
// resolves nowhere on the assembly path.
func TestAsmMissingIncludeIsAnError(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
path := writeTemp(t, "main_amd64.s", "#include \"textflag.h\"\n#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n")
_, stderr, code := capture(func() int { return cmdAsm([]string{"-GOARCH", "amd64", path}) })
if code == 0 {
t.Fatal("gasm asm accepted a file whose include resolves nowhere")
}
if !strings.Contains(stderr, `#include "nothere.h"`) {
t.Errorf("stderr does not name the failing include: %s", stderr)
}
}
+14
View File
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
### Optional preprocessing
With `Options{Expand: true}` the parser runs a pre-parse pass
(`preproc.go`) that splices `#include` files (the source directory, then the
`-I` directories), expands object and parameterised `#define` macros,
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
folds constant expressions left in operands. The go command's platform
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
language server read the raw file. The command layer adds the go_asm.h
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
defines type-checked out of its Go files for the target architecture and
GOOS, with no compiler in the loop.
### `arch`
Register files are generated programmatically (the regular `R8`-`R15`,
+26 -5
View File
@@ -141,14 +141,16 @@ gasm lint kernel_amd64.s
## asm
```text
Usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>
Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>
```
| Flag | Default | Effect |
|---|---|---|
| `-format` | `raw` | output format: `raw` (concatenated image), `elf` or `goobj` (Go object) |
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
| `-p` | empty | package path for `--format goobj`, qualifying the exported symbols |
| `-GOARCH` | empty | target architecture: `amd64`, `arm64`, `riscv64` or `loong64`; overrides the file-name suffix |
| `-GOOS` | empty | operating system for the generated `go_asm.h`: any GOOS `go/build` recognises in file names; default is the host's |
| `-o` | empty | write the output to this file instead of a hex dump on stdout |
Supported architectures: amd64 (VEX/AVX2 and EVEX/AVX-512 included), arm64,
@@ -162,6 +164,20 @@ system toolchain; `goobj` emits the Go toolchain's own object format, which
installed: the object preamble is captured from `go tool asm` and the format
version from `go version`. `raw` and `elf` need no toolchain at all.
A file that includes `go_asm.h` gets that header generated from the Go
files beside it, type-checked for the target. `-GOOS` selects the
type-checking GOOS for that header, because a GOOS-specific file needs its
platform's defines: `sys_darwin_arm64.s` fails against the ambient GOOS
(`machTimebaseInfo_numer` is missing from a linux type-check) and assembles
with `-GOOS darwin`.
Assembly preprocessing matches the toolchain's: `#define` macros (object and
parameterised) expand at the point of use, `#undef`, `#ifdef`, `#ifndef`,
`#else` and `#endif` behave as in `go tool asm`, `;` separates statements,
and `#include "file"` splices the named file in, resolved against the source
directory and then each `-I` directory in order. `textflag.h` is the one
header that is not spliced: gasm consumes its flag names natively.
```sh
gasm asm hello_amd64.s
```
@@ -305,12 +321,13 @@ gasm debug --func add --cover hello_amd64.s
## diff
```text
Usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>
Usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>
```
| Flag | Default | Effect |
|---|---|---|
| `-GOARCH` | empty | target architecture for both files, overriding the file-name suffixes |
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
| `-map` | empty | comma-separated `old=new` pairs to match functions with different names |
Functions are paired by exact name unless `--map` says otherwise, so
@@ -348,7 +365,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
## audit-instructions
```text
Usage: gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
```
Compare the gasm encoder for the given architecture (default amd64) against the
@@ -381,10 +398,14 @@ With `--corpus` the audit changes shape: it assembles every `.s` file under
DIR (default `GOROOT/src`) with the gasm encoder only, no toolchain probing.
A file whose name carries a recognisable `_arch` suffix is attempted for that
architecture; a file without one is attempted for all four, exactly as a
`GOARCH` build would compile it. The report gives the headline number (files
`GOARCH` build would compile it, and a name that names a GOOS
(`sys_darwin_arm64.s`) type-checks its generated `go_asm.h` for that GOOS.
The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
drive the encodability backlog by frequency rather than by table order. With
`--list` the report additionally prints every failing file with its failure
reason, per architecture. A run
over GOROOT takes under a second.
```sh
+657
View File
@@ -0,0 +1,657 @@
# The GOOBJ object file format
This document is a complete specification of GOOBJ, the object file format
that the Go toolchain's assembler, compiler and linker exchange, written for
implementers of independent producers and consumers. It documents the format
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
No comparable document exists upstream. The format is defined only by the
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
internal interface with no stability promise, and it can change in any
release. This specification was therefore produced by reverse engineering
that source and by parsing real objects produced by `go tool asm` and
`go tool compile`, byte for byte, against the layout described here. Within
gasm-devkit it is kept honest by the differential tests in `asm/goobj_test.go`
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
the toolchain's own products and feed gasm objects to `go build`.
Every numeric value in this document, every block index, structure size, flag
bit, type code and relocation number, was read from the Go 1.27.1 source at
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
`cmd/internal/objabi`, and exercised against assembled objects.
## Containers
The unit this document specifies is the **object**: one package's worth of
symbols, relocations and data. An object is never consumed naked. Two
wrappers exist in practice, and the linker dispatches on the first bytes of
the file.
**The bare object**, written by `go tool asm`:
```text
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
"!\n"
<GOOBJ blob>
```
The first line is the toolchain configuration string, produced by
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
version, an optional architecture qualifier such as `GOAMD64=v1`, and
`X:` followed by the enabled experiments, comma separated. The linker requires
this line to match its own configuration exactly and rejects the file
otherwise; the `-f` linker flag waives the check. Header lines may be
followed by export data delimited by `$$` markers; the header region always
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
immediately after that line.
**The package archive**, written by the compiler output pipeline and consumed
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
data in a `__.PKGDEF` member and one or more objects as further members, each
carrying the bare-object structure above. `go tool pack` creates and
inspects such archives.
| Consumer | Role |
|---|---|
| `cmd/asm` | writes objects from `.s` files |
| `cmd/compile` | writes objects from Go source |
| `cmd/link` | reads objects and archives, produces executables |
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
## Conventions
- All integers are **little endian**.
- There is **no alignment or padding** anywhere in the file; structures follow
one another byte by byte.
- Every offset stored in the file is **relative to the first byte of the GOOBJ
blob**, not to the start of the container.
- The blob opens with a 96 byte header that carries the byte offset of every
block. A block's length is the difference between its own offset and the
next block's, so the offset array is the only index the format needs.
### Layout overview
```mermaid
flowchart TB
A[Container header line and ! terminator] --> B[File header, 96 bytes]
B --> C[String table, implicit region]
C --> D[Autolib]
D --> E[PkgIndex]
E --> F[Files]
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
G --> H[RefFlags]
H --> I[Hash64 and Hash]
I --> J[RelocIndex, AuxIndex, DataIndex]
J --> K[Relocs]
K --> L[Aux]
L --> M[Data]
M --> N[RefNames]
N --> O[BlkEnd marks the end of the blob]
```
## The file header
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
offsets.
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
| 16 | 4 | Flags | Bit field, see below. |
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
Header flags:
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | ObjFlagShared | built with `-shared` |
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
| 4 | 16 | ObjFlagStd | standard library package |
### Block indices
The offset array is indexed by these constants, in file order:
| Index | Constant | Contents |
|---|---|---|
| 0 | BlkAutolib | imported packages |
| 1 | BlkPkgIndex | referenced packages, indexed |
| 2 | BlkFile | source file names |
| 3 | BlkSymdef | symbol definitions, package scope |
| 4 | BlkHashed64def | short hashed definitions |
| 5 | BlkHasheddef | hashed definitions |
| 6 | BlkNonpkgdef | non-package definitions |
| 7 | BlkNonpkgref | non-package references |
| 8 | BlkRefFlags | flags of referenced symbols |
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
| 10 | BlkHash | 16 byte hashes for hashed definitions |
| 11 | BlkRelocIndex | per symbol relocation start index |
| 12 | BlkAuxIndex | per symbol aux start index |
| 13 | BlkDataIndex | per symbol data offset |
| 14 | BlkReloc | relocations |
| 15 | BlkAux | aux symbol entries |
| 16 | BlkData | symbol payloads |
| 17 | BlkRefName | names of referenced symbols, for tools |
| 18 | BlkEnd | no contents; its offset is the end of the blob |
## The string table
There is no block index for strings. The table occupies the implicit region
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
string offset in the file points into that region. The writer de-duplicates:
each distinct string is stored once, in first-use order, and the empty string
is always the first entry, so its reference is length 0 and offset 96.
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
offset of the bytes. The bytes are stored raw, with no terminator.
## Symbol references and the package index
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
lives in:
| Value | Constant | SymIdx indexes |
|---|---|---|
| 0 | PkgIdxInvalid | never valid in a written file |
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
Assignment rules, as the toolchain performs them:
- Every definition a package exports to the linker by index lands in Symdefs
with PkgIdxSelf. The compiler puts its functions and data here; the
assembler puts only its file-local static symbols here, everything else by
name, see below.
- External package references take indices 1, 2, 3, in order of first
reference during assembly; the package names go into PkgIndex at those
indices, entry 0 is the empty package and is never referenced.
- References to the compiler's builtin functions become PkgIdxBuiltin with
SymIdx set to the builtin's index.
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
A producer must emit the definitions it made in NonPkgDefs and the pure
references in NonPkgRefs.
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
is referenced by name, PkgIdxNone, **except** file-local static symbols,
whose names carry `<>` and which are referenced by index. The compiler also
forces references by name for symbols marked `//go:linkname` and for any
symbol with the DUPOK attribute, which the linker de-duplicates by name.
## Symbol definition entries
The five definition and reference arrays (block indices 3 to 7) share one
element layout, 21 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Name | string reference |
| 8 | 2 | ABI | see table below |
| 10 | 1 | Type | symbol kind, see the kind table |
| 11 | 1 | Flag | bit field, see below |
| 12 | 1 | Flag2 | second bit field, see below |
| 13 | 4 | Siz | payload size in bytes, `uint32` |
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
The Name is a real string reference for hand-written symbols. The auxiliary
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
entries, have empty names: length 0, and their identity is only via the Aux
entries that point at them by index.
### The ABI field
| Value | Meaning |
|---|---|
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
| 1 | ABIInternal, the register ABI of compiler-generated functions |
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
### The Flag byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
| 1 | 2 | SymFlagLocal | file-local |
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
| 3 | 8 | SymFlagLeaf | leaf function |
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
amd64 the assembler itself marks any function whose frame is below
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
### The Flag2 byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
| 3 | 8 | SymFlagPkgInit | package initialisation function |
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
### The Type byte: symbol kinds
Values of `objabi.SymKind`, in numeric order:
| Value | Name | Meaning |
|---|---|---|
| 0 | Sxxx | invalid zero value |
| 1 | STEXT | executable code |
| 2 | STEXTFIPS | executable code, FIPS section |
| 3 | SRODATA | read only data |
| 4 | SRODATAFIPS | read only data, FIPS section |
| 5 | SNOPTRDATA | data without pointers |
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
| 7 | SDATA | data, may contain pointers |
| 8 | SDATAFIPS | data, FIPS section |
| 9 | SBSS | zero initialised data |
| 10 | SNOPTRBSS | zero initialised data without pointers |
| 11 | STLSBSS | thread local zero initialised data |
| 12 | SDWARFCUINFO | DWARF compile unit information |
| 13 | SDWARFCONST | DWARF constants |
| 14 | SDWARFFCN | DWARF function entry |
| 15 | SDWARFABSFCN | DWARF absolute function entry |
| 16 | SDWARFTYPE | DWARF type information |
| 17 | SDWARFVAR | DWARF variable information |
| 18 | SDWARFRANGE | DWARF range lists |
| 19 | SDWARFLOC | DWARF location lists |
| 20 | SDWARFLINES | DWARF line programs |
| 21 | SDWARFADDR | DWARF address table |
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
| 23 | SCOVERAGE_COUNTER | coverage counter |
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
## Referenced symbol flags (RefFlags)
Element size 10 bytes, one per referenced external indexed symbol that
carries a non-zero Flag2:
| Offset | Size | Field |
|---|---|---|
| 0 | 8 | Sym, a SymRef into another package |
| 8 | 1 | Flag, always 0 in current writers |
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
The linker uses these to preserve reachability of interface conversions
across package boundaries. Entries with no flags are omitted entirely.
## Hashes
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
a hash at all: the writer copies the **first 8 bytes of the symbol's
payload**. Only symbols whose content-hash section byte is 0 may use the
short form.
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
input, from `cmd/internal/obj/objfile.go`:
1. the payload size, little endian `uint64`;
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
symbols, otherwise 0;
3. for text symbols, the symbol name, which keeps distinct functions from
merging;
4. the payload with trailing zero bytes trimmed;
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
low type byte `uint8`, addend `int64`, followed by an encoding of the
target: tag byte 0 then the target's short hash, tag 1 then its full
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
for PkgIdxSelf and imported packages, no tag, then the package path
and the symbol index.
Two symbols with equal hashes are interchangeable at link time, which is what
makes content addressing work. A producer that computes these hashes wrongly
produces objects that link but de-duplicate wrongly; gasm verifies them by
byte comparison against `go tool asm`.
## The index arrays
Three arrays of `uint32`, one element per **defined** symbol plus one final
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
defined symbols, each array holds N + 1 entries, and the entry at N is the
total.
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
entry i + 1 minus entry i is its count.
- AuxIndex: the same construction over BlkAux.
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
the count is the difference of neighbours.
The toolchain writes relocations grouped per symbol in definition order, and
sorts each symbol's relocations by their Off field first. A producer that
skips the sort produces objects the linker still accepts, but that no longer
compare byte-for-byte with the toolchain's output.
## Relocations
Element size 23 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
| 4 | 1 | Siz | patch width in bytes |
| 5 | 2 | Type | relocation type, `uint16`, see the table |
| 7 | 8 | Add | addend, `int64` |
| 15 | 8 | Sym | target SymRef |
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
flavour the type prescribes is the linker's job; the object only records the
request. A size 0 relocation patches nothing and exists purely as a marker
for the linker's reachability analysis.
### Relocation types
Values of `objabi.RelocType`. The assembler and compiler emit the generic
ones plus their own architecture's family; the rest exist for other ports and
for the linker itself.
| Value | Name | Meaning |
|---|---|---|
| 1 | R_ADDR | absolute address |
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
| 6 | R_SIZE | size of the referenced symbol |
| 7 | R_CALL | direct call, PC relative |
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
| 9 | R_CALLARM64 | arm64: BL |
| 10 | R_CALLIND | indirect call marker |
| 11 | R_CALLPOWER | ppc64: call |
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
| 13 | R_CONST | constant value of the symbol |
| 14 | R_PCREL | PC relative displacement |
| 15 | R_TLS_LE | thread local, local exec offset |
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
| 17 | R_GOTOFF | offset from the GOT base |
| 18 | R_PLT0 | PLT sequence, first instruction |
| 19 | R_PLT1 | PLT sequence, second instruction |
| 20 | R_PLT2 | PLT sequence, third instruction |
| 21 | R_USEFIELD | field reachability marker |
| 22 | R_USETYPE | type reachability marker, no bytes patched |
| 23 | R_USEIFACE | interface conversion marker, size 0 |
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
| 27 | R_KEEP | keeps the target alive if the source survives |
| 28 | R_POWER_TOC | ppc64: TOC relative |
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
| 99 | R_WASMIMPORT | wasm: import module and name indices |
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
| 102 | R_INITORDER | orders inittask records, patches nothing |
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
type's value with bit 15 set. The linker strips the bit before dispatch.
## Aux symbol entries
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
auxiliary symbols to a definition; the arrays run per symbol in the order
given by AuxIndex.
| Value | Name | Attaches |
|---|---|---|
| 0 | AuxGotype | the Go type of a data symbol |
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
| 3 | AuxDwarfInfo | DWARF debug info for the function |
| 4 | AuxDwarfLoc | DWARF location lists |
| 5 | AuxDwarfRanges | DWARF range lists |
| 6 | AuxDwarfLines | DWARF line program |
| 7 | AuxPcsp | pc-value table: SP adjustments |
| 8 | AuxPcfile | pc-value table: source file indices |
| 9 | AuxPcline | pc-value table: line numbers |
| 10 | AuxPcinline | pc-value table: inlining tree positions |
| 11 | AuxPcdata | one pc-value table per live variable slot |
| 12 | AuxWasmImport | wasm import description |
| 13 | AuxWasmType | wasm export type description |
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
payload would be empty. A function assembled from `.s` source by Go 1.27.1
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
same set.
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
resolve to definitions in the package's compiled Go code when there is any.
## Symbol payloads (BlkData)
The payloads of all defined symbols, in definition order, concatenated with
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
its machine code, with the stack-split preamble and any morestack block
already included. A data symbol's payload is the bytes laid down by its DATA
directives, zero filled to its declared size. If a symbol was created from an
embedded file, the file's bytes follow the payload and count towards its
DataIndex extent; assembly producers never write this extension.
### The FuncInfo payload
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
little endian:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
| 4 | 4 | Locals | frame size in bytes |
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
| 12 | 4 | StartLine | source line of the TEXT declaration |
| 16 | 4 | NumFile | count of file indices that follow |
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
| then | 24 × NumInlTree | InlTree | nodes, see below |
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
`int32` parent PC.
The assembler derives FuncID from the symbol name through
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
specially gets that classification even when defined in assembly; an ordinary
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
### The pc-value tables
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
pc-value tables, each a sequence of value deltas and PC deltas:
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
raw delta divided by the architecture's minimum instruction length;
- the table ends with a final PC delta to the end of the function followed by
a zero byte.
The first value applies from function entry. The encoding is the one
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
final runtime pclntable carries.
### The DWARF payloads
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
symbols whose payloads are DWARF byte streams. The object format treats them
as opaque: the linker concatenates them into the final `.debug_*` sections
and resolves the relocations recorded inside them. The compiler produces
DWARF content per its own generation; gasm produces DWARF5 streams in
`asm/goobj_dwarf.go`.
## Builtins
Frequently referenced runtime functions are referenced by index rather than
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
ABI 0. Builtin names never enter the string table. The mapping only applies
while the object is not linked against shared libraries, and a linkname'd
symbol never counts as a builtin even when its name matches.
## Fingerprints
The 8 byte fingerprint identifies one build of a package. The compiler fills
it with a hash of the package's export data; the assembler leaves it zero.
The linker checks a package's fingerprint against the fingerprints its
importers recorded in their Autolib entries and rejects a mismatched build,
which is how stale objects are caught.
## What a producer must do
The checklist a third-party writer must satisfy for `go build` to accept its
objects, in one place:
1. Write the container exactly: the `go object` line matching the target
toolchain's configuration string, the `!\n` terminator, then the blob.
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
3. Deduplicate the string table, keep the empty string at offset 96, and
reference it everywhere a name appears.
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
5. Sort relocations by offset within each symbol.
6. Fill Siz with the true payload length, set Align for every
content-addressable symbol, and keep symbols under 2 GB.
7. Reference symbols by the package-index rules. An assembly producer
references everything outside the object by name, PkgIdxNone,
except its own file-local statics and the builtins; PkgIdxSelf is
reserved for definitions in this object. Assembly TEXT symbols
carry ABI 0.
8. Compute the content hashes exactly as the toolchain does, or emit no
hashed definitions at all.
## How gasm-devkit implements and verifies it
The writer lives in `asm/goobj.go`, which carries the shared container and the
amd64 relocation emission, with per-architecture relocation emitters in
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
resolution in `asm/goobj_resolve.go` and DWARF generation in
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
that `go build` consumes in place of the toolchain's own.
Verification is differential and continuous:
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
output for the same source, byte for byte;
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
objects and runs them;
- `gasm verify` keeps the machine code itself identical to the toolchain's,
which is the precondition for the object comparison to be meaningful.
## Versioning and drift
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
a toolchain release changes the format, it changes that string first, and the
linker refuses blobs whose magic it does not know. The watch points for a new
release are, in order: the magic, the block index list, the Aux type list,
the tail of the relocation table, the FuncInfo layout, and the builtin table
count. gasm's tests fail against any of these changes, which is the mechanism
that keeps this document and the writer current.
The authoritative sources, for the release this document covers:
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
document;
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
- `cmd/internal/objabi/reloctype.go`: relocation types;
- `cmd/internal/objabi/symkind.go`: symbol kinds;
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.
+124
View File
@@ -0,0 +1,124 @@
# AMD64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
gasm's encoder, whose output is compared byte for byte with the toolchain's
and executed on real hardware (`gasm verify`). The complete mnemonic
inventory lives in the generated appendix
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
the conventions.
## Registers
| Group | Names | Notes |
|---|---|---|
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
| Mask | `K0` to `K7` | AVX-512 opmask |
| System | `TLS` | the thread pointer, see below |
Roles the calling convention fixes, which assembly must respect and can rely
on:
- `SP` is the hardware stack pointer; the virtual frame pointer of the
common language is the pseudo-register SP of OPERANDS.md, a different
spelling with a different meaning.
- `BP` is callee-save. The assembler inserts the save and restore whenever
the function has a non-zero frame, so using BP as a general register
interferes with sampling profilers that walk the frame chain.
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
the closure context; `R12` and `R13` are the register ABI's scratch pair
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
uses. An ABI0 assembly function called from Go sees none of these live
across the call, but runtime assembly reads them directly.
- The legacy spellings for the goroutine pointer are the macros of
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
`0(r)(TLS*1)`, the segment base riding the index field.
## Addressing
The common forms of OPERANDS.md, with the amd64 specifics:
```text
offset(base) MOVQ 16(BX), AX
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB) MOVQ ·table(SB), CX
```
- Global references assemble as absolute addresses and produce R_ADDR
relocations; branch targets produce R_PCREL.
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
index position, exists for the gather and scatter families.
- There are no segment overrides in source; the one segment-flavoured form
is the TLS base in the index field shown above.
## The frame and the split check
The assembler manages the frame, not the programmer:
- It inserts the `BP` save and restore for any non-zero frame.
- It inserts the stack-split check for any function that is not NoSplit:
the check compares SP against the guard, and on exhaustion calls
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
the small compare; frames at or below 4096 bytes, StackBig, use the
adjusted form; larger frames compare in two steps.
- On amd64 the assembler marks a function NoSplit itself when the frame is
under StackSmall and the body calls nothing that needs stack: such a
function carries the NoSplit flag in the object without the source ever
writing NOSPLIT.
Results and arguments are stack-only in ABI0: the caller's frame carries
them at FP offsets, per the Go prototype.
## Instructions
The inventory counts 1654 recognised mnemonics today, of which the encoder
emits 1113; both numbers are generated in the appendix, and the gap is the
encoder backlog that `gasm audit-instructions` measures. The families:
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
`MOVLQSX` and their siblings, which the compiler's output leans on;
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
string primitives `MOVS` and `STOS`.
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
pair `ADCX` and `ADOX`; `CRC32`.
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
immediates are not encodable on this target, so the assembler
materialises them: the constant lands in a synthesised read-only pool,
and a positive zero collapses to `XORPS` of the register with itself,
exactly as the toolchain does.
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
integer and floating operations, shuffles, lane extracts and inserts and
the imm8-controlled forms.
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
opmask operations on `K0` to `K7`, gathers and scatters, and the
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
`STMXCSR`, the prefetch family.
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
words into the stream for encodings the assembler does not know; `ADJSP`
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
are compiler-side names the table recognises but an encoder need not
emit.
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
gasm reports it as an explicit error, never as wrong bytes, and the
`unencodable-instruction` lint flags it at edit time.
## Relocations
The relocations an amd64 object carries, all specified in
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
emits in the four-byte flavour.
+122
View File
@@ -0,0 +1,122 @@
# ARM64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
## Registers
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
platform-reserved register and the Go toolchain never addresses it.
## Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the
mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
| `MOVD` | ldr, str, stur, 64-bit |
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
| `MOVWU` | ldr, 32-bit zero extending |
| `MOVH` | ldrsh, strh, sturh |
| `MOVHU` | ldrh |
| `MOVB` | ldrsb, strb, sturb |
| `MOVBU` | ldrb |
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
## Addressing
```text
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
```
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
## Operand order and the special forms
Most instructions appear in left-to-right assignment order: `ADD R11,
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
`<Rm>, <Ra>, <Rn>, <Rd>`.
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
`$<lsb>, <Rn>, $<width>, <Rd>`.
- The conditional compare and select families carry the condition as the
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
`FCCMPD AL, F8, F26, $0`.
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
last: `STLXR ZR, (R15), R16`.
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
Shifted and extended register operands ride the register: `R19>>30`,
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
## Conditions, branches and names
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
per-condition names such as `BEQ`. Both spellings exist; the canonical
names are what the generated inventory lists.
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
`JMP (R3)` and `CALL (R17)`.
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
an alias of `HINT $0`.
- `umov` is written as `VMOV`.
## Constants
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
assembler.
- Large integer constants: `MOV` materialises any 64-bit constant, the
closest-instruction way.
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
64-bit halves for a 128-bit value:
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
## SIMD
Floating-point and SIMD instructions mostly carry a `V` prefix
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
## Relocations
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
`R_CALLARM64` for direct calls, all specified in
[GOOBJ.md](../GOOBJ.md).
+148
View File
@@ -0,0 +1,148 @@
# Directives: TEXT, DATA, GLOBL and the annotations
Layer 1, the common language, with the flag vocabulary both layers share.
Verified against `go tool asm` of Go 1.27.1, against the shipped headers
`textflag.h` and `funcdata.h` in `$GOROOT/pkg/include`, and against gasm's
parser. Where gasm extends a directive, the extension says so and is marked.
Six directives exist. Three define things: TEXT, DATA, GLOBL. Three
annotate: FUNCDATA, PCDATA, PCALIGN.
## TEXT
```text
// func Add(a, b int64) int64
TEXT ·Add(SB), NOSPLIT, $0-24
...instructions...
RET
```
```text
TEXT symbol(SB), [flags,] $framesize[-argsize]
```
- The symbol is an `·Name(SB)` reference into the current package, or a
fully qualified name.
- The optional flag argument is a constant expression, normally an OR of the
names from `textflag.h`, the table below. Without `#include "textflag.h"`
the names are not macros and the assembler reports the misleading error
`illegal or missing addressing mode for symbol NOSPLIT`: include the
header first.
- `$framesize-argsize` is two constants, not a subtraction: the local frame
size in bytes, and the caller's argument area in bytes. The argument size
may be omitted entirely, `$16`, which marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`); a frame size may be negative only in the generated ABI
wrappers.
- A function whose last instruction is not a branch cannot fall through into
the next TEXT: the toolchain appends a jump to itself, so end functions
with `RET` deliberately.
- One TEXT per symbol; redeclaring is an error. The TEXT line also fixes the
function's source line for traceback: it is the line number that pcln
reports for the function's start.
The framesize and argsize fields do real work: the framesize drives the
stack-split preamble (RUNTIME.md carries the contract), and both travel into
the FuncInfo record of the object (GOOBJ.md carries its layout).
### The flag table
Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
| Name | Value | Applies to | Meaning |
|---|---|---|---|
| NOPROF | 1 | both | do not profile; deprecated |
| DUPOK | 2 | both | the linker may keep one of several duplicates |
| NOSPLIT | 4 | TEXT | no stack-split preamble |
| RODATA | 8 | data | put the data in a read-only section |
| NOPTR | 16 | data | the data contains no pointers |
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
| TLSBSS | 256 | data | a thread local word in BSS |
| NOFRAME | 512 | TEXT | no frame setup; only valid with a frame size of 0 |
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
Rules with teeth:
- `NOSPLIT` removes the split check, so the frame plus everything the
function calls must fit in the stack segment that remains. It exists to
protect the splitting code itself; reaching for it to save two instructions
is how stack overflows corrupt memory. On amd64 the assembler additionally
marks small leaf functions NoSplit itself and omits the check, so the
absence of the preamble is not proof the flag was written.
- A TEXT whose symbol is declared `ABIInternal` must carry NOSPLIT: the
assembler rejects it otherwise, because it cannot generate
the split path for a register-ABI function.
- `RODATA` implies NOPTR for the garbage collector.
## DATA
```text
DATA ·table+0(SB)/8, $0x0102030405060708
DATA ·msg+0(SB)/14, $"hello, world\n"
GLOBL ·msg(SB), RODATA, $14
```
```text
DATA symbol+offset(SB)/width, value
```
- `width` is exactly 1, 2, 4 or 8: the initialiser is written into the data
image at `symbol+offset` in that many bytes.
- The value is an integer or character constant of the width, or a string
literal whose byte length equals the width exactly; escapes count. Long
data is written as successive DATA lines at increasing offsets; bytes the
directives never name are zero.
- Every symbol initialised with DATA ends with a GLOBL line declaring its
total size, after all of its DATA lines.
A symbol containing pointers cannot be defined in assembly, because the
collector cannot see into it: define it in Go and refer to it by name. As a
rule, data that is not read-only belongs in Go.
Extension, gasm only: a DATA initialiser may name a symbol,
`DATA ·fn+0(SB)/8, $·handler(SB)`, which gasm lays down as an absolute
relocation on that field. The toolchain offers no ground truth for this
form; gasm's behaviour is verified by linking and execution.
## GLOBL
```text
GLOBL symbol(SB), [flags,] $size
```
Declares the symbol global with its total size in bytes. The useful flags
are RODATA, NOPTR, DUPOK and TLSBSS from the table above. Uninitialised
bytes are zero, which makes GLOBL with no DATA the language's BSS.
## FUNCDATA and PCDATA
```text
FUNCDATA $functypeid, symbol(SB)
PCDATA $pctypeid, $value
```
The compiler's annotations for the garbage collector and traceback, named by
the ids in `funcdata.h`: FUNCDATA 0 to 7 (args pointer maps, locals pointer
maps, stack objects, inline tree, open-coded defer info, argument info,
argument liveness, wrap info), PCDATA 0 to 4 (unsafe point, stack map index,
inline tree index, argument liveness index, panic bounds). Assembly code
normally reaches them only through the macro forms in `funcdata.h`, which
RUNTIME.md explains. Outside the macros, hand-written PCDATA is meaningless:
the values are pc-value tables the compiler builds from its own view of the
program.
## PCALIGN
```text
PCALIGN $32
```
Pads the code so that the next instruction lands on the given boundary,
which must be a power of two and at least the target's instruction
alignment. Supported on amd64, arm64, ppc64, loong64 and riscv64. The
padding instructions are the target's NOP encoding, so the bytes between
functions differ from what the instruction stream alone would produce, which
matters to anyone comparing encodings byte for byte.
File diff suppressed because it is too large Load Diff
+570
View File
@@ -0,0 +1,570 @@
# ARM64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/arm64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADC` | ADC (64-bit) |
| `ADCS` | ADCS (64-bit) |
| `ADCSW` | ADCS (32-bit) |
| `ADCW` | ADC (32-bit) |
| `ADD` | ADD (64-bit) |
| `ADDS` | ADDS (64-bit) |
| `ADDSW` | ADDS (32-bit) |
| `ADDW` | ADD (32-bit) |
| `ADR` | Address of label/page |
| `ADRP` | Address of label/page |
| `AESD` | AES round |
| `AESE` | AES round |
| `AESIMC` | AES round |
| `AESMC` | AES round |
| `AND` | AND (64-bit) |
| `ANDS` | ANDS (64-bit) |
| `ANDSW` | ANDS (32-bit) |
| `ANDW` | AND (32-bit) |
| `ASR` | ASR shift |
| `ASRW` | ASR shift (32-bit) |
| `AT` | |
| `AUTIA1716` | |
| `AUTIASP` | |
| `AUTIB1716` | |
| `AUTIBSP` | |
| `BCC` | Conditional branch |
| `BCS` | Conditional branch |
| `BEQ` | Conditional branch |
| `BFI` | |
| `BFIW` | |
| `BFM` | |
| `BFMW` | |
| `BFXIL` | Bitfield extract |
| `BFXILW` | |
| `BGE` | Conditional branch |
| `BGT` | Conditional branch |
| `BHI` | Conditional branch |
| `BHS` | Conditional branch |
| `BIC` | BIC (64-bit) |
| `BICS` | BICS (64-bit) |
| `BICSW` | BICS (32-bit) |
| `BICW` | BIC (32-bit) |
| `BLE` | Conditional branch |
| `BLO` | Conditional branch |
| `BLS` | Conditional branch |
| `BLT` | Conditional branch |
| `BMI` | Conditional branch |
| `BNE` | Conditional branch |
| `BPL` | Conditional branch |
| `BRK` | Breakpoint |
| `BTI` | |
| `BVC` | Conditional branch |
| `BVS` | Conditional branch |
| `CASAD` | |
| `CASALB` | |
| `CASALD` | |
| `CASALH` | |
| `CASALW` | |
| `CASAW` | |
| `CASB` | |
| `CASD` | |
| `CASH` | |
| `CASLD` | |
| `CASLW` | |
| `CASPD` | |
| `CASPW` | |
| `CASW` | |
| `CBNZ` | Compare/test and branch |
| `CBNZW` | Compare/test and branch (32-bit) |
| `CBZ` | Compare/test and branch |
| `CBZW` | Compare/test and branch (32-bit) |
| `CCMN` | Conditional compare |
| `CCMNW` | Conditional compare |
| `CCMP` | Conditional compare |
| `CCMPW` | Conditional compare |
| `CINC` | Conditional select |
| `CINCW` | Conditional select (32-bit) |
| `CINV` | Conditional select |
| `CINVW` | Conditional select (32-bit) |
| `CLREX` | |
| `CLS` | Bit manipulation |
| `CLSW` | Bit manipulation |
| `CLZ` | Bit manipulation |
| `CLZW` | Bit manipulation |
| `CMN` | CMN (64-bit) |
| `CMNW` | CMN (32-bit) |
| `CMP` | CMP (64-bit) |
| `CMPW` | CMP (32-bit) |
| `CNEG` | Conditional select |
| `CNEGW` | Conditional select (32-bit) |
| `CRC32B` | |
| `CRC32CB` | |
| `CRC32CH` | |
| `CRC32CW` | |
| `CRC32CX` | |
| `CRC32H` | |
| `CRC32W` | |
| `CRC32X` | |
| `CSEL` | Conditional select |
| `CSELW` | Conditional select (32-bit) |
| `CSET` | Conditional select |
| `CSETM` | Conditional select |
| `CSETMW` | Conditional select (32-bit) |
| `CSETW` | Conditional select (32-bit) |
| `CSINC` | Conditional select |
| `CSINCW` | Conditional select (32-bit) |
| `CSINV` | Conditional select |
| `CSINVW` | Conditional select (32-bit) |
| `CSNEG` | Conditional select |
| `CSNEGW` | Conditional select (32-bit) |
| `DC` | Data cache maintenance |
| `DCPS1` | |
| `DCPS2` | |
| `DCPS3` | |
| `DMB` | Barrier |
| `DRPS` | |
| `DSB` | Barrier |
| `DWORD` | |
| `EON` | EON (64-bit) |
| `EONW` | EON (32-bit) |
| `EOR` | EOR (64-bit) |
| `EORW` | EOR (32-bit) |
| `ERET` | |
| `EXTR` | Bitfield extract |
| `EXTRW` | |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | |
| `FADDS` | |
| `FCCMPD` | |
| `FCCMPED` | |
| `FCCMPES` | |
| `FCCMPS` | |
| `FCMPD` | |
| `FCMPED` | |
| `FCMPES` | |
| `FCMPS` | |
| `FCSELD` | |
| `FCSELS` | |
| `FCVTDH` | |
| `FCVTDS` | |
| `FCVTHD` | |
| `FCVTHS` | |
| `FCVTSD` | |
| `FCVTSH` | |
| `FCVTZSD` | |
| `FCVTZSDW` | |
| `FCVTZSS` | |
| `FCVTZSSW` | |
| `FCVTZUD` | |
| `FCVTZUDW` | |
| `FCVTZUS` | |
| `FCVTZUSW` | |
| `FDIVD` | |
| `FDIVS` | |
| `FLDPD` | Register-pair load or store |
| `FLDPQ` | |
| `FLDPS` | |
| `FMADDD` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXNMD` | |
| `FMAXNMS` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINNMD` | |
| `FMINNMS` | |
| `FMINS` | |
| `FMOVD` | Move / load / store |
| `FMOVQ` | |
| `FMOVS` | Move / load / store |
| `FMSUBD` | |
| `FMSUBS` | |
| `FMULD` | |
| `FMULS` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNMADDD` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBS` | |
| `FNMULD` | |
| `FNMULS` | |
| `FRINTAD` | |
| `FRINTAS` | |
| `FRINTID` | |
| `FRINTIS` | |
| `FRINTMD` | |
| `FRINTMS` | |
| `FRINTND` | |
| `FRINTNS` | |
| `FRINTPD` | |
| `FRINTPS` | |
| `FRINTXD` | |
| `FRINTXS` | |
| `FRINTZD` | |
| `FRINTZS` | |
| `FSQRTD` | |
| `FSQRTS` | |
| `FSTPD` | Register-pair load or store |
| `FSTPQ` | |
| `FSTPS` | |
| `FSUBD` | |
| `FSUBS` | |
| `HINT` | |
| `HLT` | |
| `HVC` | Exception generation |
| `IC` | |
| `ISB` | Barrier |
| `LDADDAB` | |
| `LDADDAD` | |
| `LDADDAH` | |
| `LDADDALB` | |
| `LDADDALD` | |
| `LDADDALH` | |
| `LDADDALW` | |
| `LDADDAW` | |
| `LDADDB` | |
| `LDADDD` | |
| `LDADDH` | |
| `LDADDLB` | |
| `LDADDLD` | |
| `LDADDLH` | |
| `LDADDLW` | |
| `LDADDW` | |
| `LDAR` | Atomic memory operation |
| `LDARB` | Atomic memory operation |
| `LDARH` | Atomic memory operation |
| `LDARW` | Atomic memory operation |
| `LDAXP` | |
| `LDAXPW` | |
| `LDAXR` | Atomic memory operation |
| `LDAXRB` | Atomic memory operation |
| `LDAXRH` | Atomic memory operation |
| `LDAXRW` | Atomic memory operation |
| `LDCLRAB` | |
| `LDCLRAD` | |
| `LDCLRAH` | |
| `LDCLRALB` | |
| `LDCLRALD` | |
| `LDCLRALH` | |
| `LDCLRALW` | |
| `LDCLRAW` | |
| `LDCLRB` | |
| `LDCLRD` | |
| `LDCLRH` | |
| `LDCLRLB` | |
| `LDCLRLD` | |
| `LDCLRLH` | |
| `LDCLRLW` | |
| `LDCLRW` | |
| `LDEORAB` | |
| `LDEORAD` | |
| `LDEORAH` | |
| `LDEORALB` | |
| `LDEORALD` | |
| `LDEORALH` | |
| `LDEORALW` | |
| `LDEORAW` | |
| `LDEORB` | |
| `LDEORD` | |
| `LDEORH` | |
| `LDEORLB` | |
| `LDEORLD` | |
| `LDEORLH` | |
| `LDEORLW` | |
| `LDEORW` | |
| `LDORAB` | |
| `LDORAD` | |
| `LDORAH` | |
| `LDORALB` | |
| `LDORALD` | |
| `LDORALH` | |
| `LDORALW` | |
| `LDORAW` | |
| `LDORB` | |
| `LDORD` | |
| `LDORH` | |
| `LDORLB` | |
| `LDORLD` | |
| `LDORLH` | |
| `LDORLW` | |
| `LDORW` | |
| `LDP` | Register-pair load or store |
| `LDPSW` | |
| `LDPW` | Register-pair load or store |
| `LDXP` | |
| `LDXPW` | |
| `LDXR` | |
| `LDXRB` | |
| `LDXRH` | |
| `LDXRW` | |
| `LSL` | LSL shift |
| `LSLW` | LSL shift (32-bit) |
| `LSR` | LSR shift |
| `LSRW` | LSR shift (32-bit) |
| `MADD` | Multiply / multiply-accumulate |
| `MADDW` | |
| `MNEG` | Multiply / multiply-accumulate |
| `MNEGW` | |
| `MOVB` | Move / load / store |
| `MOVBU` | Move / load / store |
| `MOVD` | Move / load / store |
| `MOVH` | Move / load / store |
| `MOVHU` | Move / load / store |
| `MOVK` | Move wide constant |
| `MOVKW` | Move wide constant |
| `MOVN` | Move wide constant |
| `MOVNW` | Move wide constant |
| `MOVP` | |
| `MOVPD` | |
| `MOVPQ` | |
| `MOVPS` | |
| `MOVPSW` | |
| `MOVPW` | |
| `MOVW` | Move / load / store |
| `MOVWU` | Move / load / store |
| `MOVZ` | Move wide constant |
| `MOVZW` | Move wide constant |
| `MRS` | System register access |
| `MSR` | System register access |
| `MSUB` | Multiply / multiply-accumulate |
| `MSUBW` | |
| `MUL` | Multiply / multiply-accumulate |
| `MULW` | |
| `MVN` | MVN (64-bit) |
| `MVNW` | MVN (32-bit) |
| `NEG` | NEG (64-bit) |
| `NEGS` | |
| `NEGSW` | |
| `NEGW` | NEG (32-bit) |
| `NGC` | NGC (64-bit) |
| `NGCS` | |
| `NGCSW` | |
| `NGCW` | NGC (32-bit) |
| `NOOP` | |
| `ORN` | ORN (64-bit) |
| `ORNW` | ORN (32-bit) |
| `ORR` | ORR (64-bit) |
| `ORRW` | ORR (32-bit) |
| `PACIASP` | |
| `PACIBSP` | |
| `PRFM` | Memory prefetch |
| `PRFUM` | |
| `RBIT` | Bit manipulation |
| `RBITW` | Bit manipulation |
| `REM` | |
| `REMW` | |
| `REV` | Bit manipulation |
| `REV16` | Bit manipulation |
| `REV16W` | |
| `REV32` | Bit manipulation |
| `REVW` | Bit manipulation |
| `ROR` | ROR shift |
| `RORW` | ROR shift (32-bit) |
| `SBC` | SBC (64-bit) |
| `SBCS` | SBCS (64-bit) |
| `SBCSW` | SBCS (32-bit) |
| `SBCW` | SBC (32-bit) |
| `SBFIZ` | |
| `SBFIZW` | |
| `SBFM` | Bitfield extract |
| `SBFMW` | |
| `SBFX` | Bitfield extract |
| `SBFXW` | |
| `SCVTFD` | |
| `SCVTFS` | |
| `SCVTFWD` | |
| `SCVTFWS` | |
| `SDIV` | Divide |
| `SDIVW` | Divide |
| `SEV` | |
| `SEVL` | |
| `SHA1C` | SHA round |
| `SHA1H` | SHA round |
| `SHA1M` | SHA round |
| `SHA1P` | SHA round |
| `SHA1SU0` | SHA round |
| `SHA1SU1` | SHA round |
| `SHA256H` | SHA round |
| `SHA256H2` | SHA round |
| `SHA256SU0` | SHA round |
| `SHA256SU1` | SHA round |
| `SHA512H` | SHA round |
| `SHA512H2` | SHA round |
| `SHA512SU0` | SHA round |
| `SHA512SU1` | SHA round |
| `SMADDL` | Multiply / multiply-accumulate |
| `SMC` | Exception generation |
| `SMNEGL` | |
| `SMSUBL` | Multiply / multiply-accumulate |
| `SMULH` | Multiply / multiply-accumulate |
| `SMULL` | Multiply / multiply-accumulate |
| `STLR` | Atomic memory operation |
| `STLRB` | Atomic memory operation |
| `STLRH` | Atomic memory operation |
| `STLRW` | Atomic memory operation |
| `STLXP` | |
| `STLXPW` | |
| `STLXR` | |
| `STLXRB` | |
| `STLXRH` | |
| `STLXRW` | |
| `STP` | Register-pair load or store |
| `STPW` | Register-pair load or store |
| `STXP` | |
| `STXPW` | |
| `STXR` | Atomic memory operation |
| `STXRB` | Atomic memory operation |
| `STXRH` | Atomic memory operation |
| `STXRW` | Atomic memory operation |
| `SUB` | SUB (64-bit) |
| `SUBS` | SUBS (64-bit) |
| `SUBSW` | SUBS (32-bit) |
| `SUBW` | SUB (32-bit) |
| `SVC` | Exception generation |
| `SWPAB` | |
| `SWPAD` | |
| `SWPAH` | |
| `SWPALB` | |
| `SWPALD` | |
| `SWPALH` | |
| `SWPALW` | |
| `SWPAW` | |
| `SWPB` | |
| `SWPD` | |
| `SWPH` | |
| `SWPLB` | |
| `SWPLD` | |
| `SWPLH` | |
| `SWPLW` | |
| `SWPW` | |
| `SXTB` | |
| `SXTBW` | |
| `SXTH` | |
| `SXTHW` | |
| `SXTW` | |
| `SYS` | |
| `SYSL` | |
| `TBNZ` | Compare/test and branch |
| `TBZ` | Compare/test and branch |
| `TLBI` | |
| `TST` | TST (64-bit) |
| `TSTW` | TST (32-bit) |
| `UBFIZ` | |
| `UBFIZW` | |
| `UBFM` | Bitfield extract |
| `UBFMW` | |
| `UBFX` | Bitfield extract |
| `UBFXW` | |
| `UCVTFD` | |
| `UCVTFS` | |
| `UCVTFWD` | |
| `UCVTFWS` | |
| `UDIV` | Divide |
| `UDIVW` | Divide |
| `UMADDL` | Multiply / multiply-accumulate |
| `UMNEGL` | |
| `UMSUBL` | Multiply / multiply-accumulate |
| `UMULH` | Multiply / multiply-accumulate |
| `UMULL` | Multiply / multiply-accumulate |
| `UREM` | |
| `UREMW` | |
| `UXTB` | |
| `UXTBW` | |
| `UXTH` | |
| `UXTHW` | |
| `UXTW` | |
| `VADD` | NEON SIMD vector operation |
| `VADDP` | |
| `VADDV` | NEON SIMD vector operation |
| `VAND` | NEON SIMD vector operation |
| `VBCAX` | Three-way XOR / rotate crypto vector operation |
| `VBIF` | NEON SIMD vector operation |
| `VBIT` | |
| `VBSL` | NEON SIMD vector operation |
| `VCMEQ` | |
| `VCMTST` | |
| `VCNT` | NEON SIMD vector operation |
| `VDUP` | NEON SIMD vector operation |
| `VEOR` | NEON SIMD vector operation |
| `VEOR3` | Three-way XOR / rotate crypto vector operation |
| `VEXT` | NEON SIMD vector operation |
| `VFMLA` | NEON SIMD vector operation |
| `VFMLS` | NEON SIMD vector operation |
| `VLD1` | NEON SIMD vector operation |
| `VLD1R` | |
| `VLD2` | NEON SIMD vector operation |
| `VLD2R` | |
| `VLD3` | NEON SIMD vector operation |
| `VLD3R` | |
| `VLD4` | NEON SIMD vector operation |
| `VLD4R` | |
| `VMOV` | NEON SIMD vector operation |
| `VMOVD` | |
| `VMOVI` | NEON SIMD vector operation |
| `VMOVQ` | NEON SIMD vector operation |
| `VMOVS` | |
| `VORR` | NEON SIMD vector operation |
| `VPMULL` | |
| `VPMULL2` | |
| `VRAX1` | Three-way XOR / rotate crypto vector operation |
| `VRBIT` | |
| `VREV16` | NEON SIMD vector operation |
| `VREV32` | NEON SIMD vector operation |
| `VREV64` | NEON SIMD vector operation |
| `VSHL` | NEON SIMD vector operation |
| `VSLI` | |
| `VSRI` | |
| `VST1` | NEON SIMD vector operation |
| `VST2` | NEON SIMD vector operation |
| `VST3` | NEON SIMD vector operation |
| `VST4` | NEON SIMD vector operation |
| `VSUB` | NEON SIMD vector operation |
| `VTBL` | NEON SIMD vector operation |
| `VTBX` | NEON SIMD vector operation |
| `VTRN1` | NEON SIMD vector operation |
| `VTRN2` | NEON SIMD vector operation |
| `VUADDLV` | |
| `VUADDW` | |
| `VUADDW2` | |
| `VUMAX` | |
| `VUMIN` | |
| `VUSHLL` | |
| `VUSHLL2` | |
| `VUSHR` | NEON SIMD vector operation |
| `VUSRA` | |
| `VUXTL` | |
| `VUXTL2` | |
| `VUZP1` | NEON SIMD vector operation |
| `VUZP2` | NEON SIMD vector operation |
| `VXAR` | Three-way XOR / rotate crypto vector operation |
| `VZIP1` | NEON SIMD vector operation |
| `VZIP2` | NEON SIMD vector operation |
| `WFE` | |
| `WFI` | |
| `WORD` | |
| `YIELD` | |
| `B` | Unconditional branch |
| `BL` | Branch with link |
Recognised: 554 mnemonics.
+830
View File
@@ -0,0 +1,830 @@
# LoongArch 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/loong64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ABSD` | |
| `ABSF` | |
| `ADD` | Integer add (word) |
| `ADDD` | Add doubleword |
| `ADDF` | |
| `ADDV` | |
| `ADDV16` | |
| `ADDVU` | |
| `ADDW` | Add word |
| `ALSLV` | |
| `ALSLW` | |
| `ALSLWU` | |
| `AMADDDBV` | |
| `AMADDDBW` | |
| `AMADDV` | |
| `AMADDW` | |
| `AMANDDBV` | |
| `AMANDDBW` | |
| `AMANDV` | |
| `AMANDW` | |
| `AMCASB` | |
| `AMCASDBB` | |
| `AMCASDBH` | |
| `AMCASDBV` | |
| `AMCASDBW` | |
| `AMCASH` | |
| `AMCASV` | |
| `AMCASW` | |
| `AMMAXDBV` | |
| `AMMAXDBVU` | |
| `AMMAXDBW` | |
| `AMMAXDBWU` | |
| `AMMAXV` | |
| `AMMAXVU` | |
| `AMMAXW` | |
| `AMMAXWU` | |
| `AMMINDBV` | |
| `AMMINDBVU` | |
| `AMMINDBW` | |
| `AMMINDBWU` | |
| `AMMINV` | |
| `AMMINVU` | |
| `AMMINW` | |
| `AMMINWU` | |
| `AMORDBV` | |
| `AMORDBW` | |
| `AMORV` | |
| `AMORW` | |
| `AMSWAPB` | |
| `AMSWAPDBB` | |
| `AMSWAPDBH` | |
| `AMSWAPDBV` | |
| `AMSWAPDBW` | |
| `AMSWAPH` | |
| `AMSWAPV` | |
| `AMSWAPW` | |
| `AMXORDBV` | |
| `AMXORDBW` | |
| `AMXORV` | |
| `AMXORW` | |
| `AND` | Bitwise AND |
| `ANDN` | |
| `BEQ` | Branch if equal |
| `BFPF` | |
| `BFPT` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGTZ` | |
| `BITREV4B` | |
| `BITREV8B` | |
| `BITREVV` | |
| `BITREVW` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BREAK` | Breakpoint |
| `BSTRINSV` | |
| `BSTRINSW` | |
| `BSTRPICKV` | |
| `BSTRPICKW` | |
| `CLOV` | |
| `CLOW` | |
| `CLZV` | |
| `CLZW` | |
| `CMPEQD` | |
| `CMPEQF` | |
| `CMPGED` | |
| `CMPGEF` | |
| `CMPGTD` | |
| `CMPGTF` | |
| `CPUCFG` | |
| `CRCCWBW` | |
| `CRCCWHW` | |
| `CRCCWVW` | |
| `CRCCWWW` | |
| `CRCWBW` | |
| `CRCWHW` | |
| `CRCWVW` | |
| `CRCWWW` | |
| `CTOV` | |
| `CTOW` | |
| `CTZV` | |
| `CTZW` | |
| `DBAR` | Barrier |
| `DIV` | Divide (word) |
| `DIVD` | Divide doubleword |
| `DIVF` | |
| `DIVU` | |
| `DIVV` | |
| `DIVVU` | |
| `DIVW` | Divide word |
| `DIVWU` | |
| `EXTWB` | |
| `EXTWH` | |
| `FCLASSD` | |
| `FCLASSF` | |
| `FCOPYSGD` | |
| `FCOPYSGF` | |
| `FFINTDV` | |
| `FFINTDW` | |
| `FFINTFV` | |
| `FFINTFW` | |
| `FLOGBD` | |
| `FLOGBF` | |
| `FMADDD` | |
| `FMADDF` | |
| `FMAXAD` | |
| `FMAXAF` | |
| `FMAXD` | |
| `FMAXF` | |
| `FMINAD` | |
| `FMINAF` | |
| `FMIND` | |
| `FMINF` | |
| `FMSUBD` | |
| `FMSUBF` | |
| `FNMADDD` | |
| `FNMADDF` | |
| `FNMSUBD` | |
| `FNMSUBF` | |
| `FSCALEBD` | |
| `FSCALEBF` | |
| `FSEL` | |
| `FTINTRMVD` | |
| `FTINTRMVF` | |
| `FTINTRMWD` | |
| `FTINTRMWF` | |
| `FTINTRNEVD` | |
| `FTINTRNEVF` | |
| `FTINTRNEWD` | |
| `FTINTRNEWF` | |
| `FTINTRPVD` | |
| `FTINTRPVF` | |
| `FTINTRPWD` | |
| `FTINTRPWF` | |
| `FTINTRZVD` | |
| `FTINTRZVF` | |
| `FTINTRZWD` | |
| `FTINTRZWF` | |
| `FTINTVD` | |
| `FTINTVF` | |
| `FTINTWD` | |
| `FTINTWF` | |
| `JIRL` | Jump indirect with link |
| `LL` | |
| `LLV` | |
| `LU12IW` | |
| `LU32ID` | |
| `LU52ID` | |
| `LUI` | |
| `MASKEQZ` | |
| `MASKNEZ` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVDF` | |
| `MOVDV` | |
| `MOVDW` | |
| `MOVF` | |
| `MOVFD` | |
| `MOVFV` | |
| `MOVFW` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVV` | |
| `MOVVD` | |
| `MOVVF` | |
| `MOVVP` | |
| `MOVW` | |
| `MOVWD` | |
| `MOVWF` | |
| `MOVWP` | |
| `MOVWU` | |
| `MUL` | Multiply (word) |
| `MULD` | Multiply doubleword |
| `MULF` | |
| `MULH` | |
| `MULHU` | |
| `MULHV` | |
| `MULHVU` | |
| `MULV` | |
| `MULVU` | |
| `MULW` | Multiply word |
| `MULWVW` | |
| `MULWVWU` | |
| `NEGD` | |
| `NEGF` | |
| `NEGV` | |
| `NEGW` | |
| `NOOP` | |
| `NOR` | Bitwise NOR |
| `OR` | Bitwise OR |
| `ORN` | |
| `PCADDU12I` | |
| `PCALAU12I` | |
| `PRELD` | |
| `PRELDX` | |
| `RDTIMED` | |
| `RDTIMEHW` | |
| `RDTIMELW` | |
| `REM` | |
| `REMU` | |
| `REMV` | |
| `REMVU` | |
| `REMW` | |
| `REMWU` | |
| `REVB2H` | |
| `REVB2W` | |
| `REVB4H` | |
| `REVBV` | |
| `REVH2W` | |
| `REVHV` | |
| `RFE` | |
| `ROTR` | Rotate right |
| `ROTRV` | |
| `SC` | |
| `SCV` | |
| `SGT` | |
| `SGTU` | |
| `SLL` | Shift left logical |
| `SLLV` | |
| `SQRTD` | |
| `SQRTF` | |
| `SRA` | Shift right arithmetic |
| `SRAV` | |
| `SRL` | Shift right logical |
| `SRLV` | |
| `SUB` | Subtract (word) |
| `SUBD` | Subtract doubleword |
| `SUBF` | |
| `SUBV` | |
| `SUBVU` | |
| `SUBW` | Subtract word |
| `SYSCALL` | System call |
| `TEQ` | |
| `TNE` | |
| `TRUNCDV` | |
| `TRUNCDW` | |
| `TRUNCFV` | |
| `TRUNCFW` | |
| `VADDB` | |
| `VADDBU` | |
| `VADDD` | |
| `VADDF` | |
| `VADDH` | |
| `VADDHU` | |
| `VADDQ` | |
| `VADDV` | |
| `VADDVU` | |
| `VADDW` | |
| `VADDWEVHB` | |
| `VADDWEVHBU` | |
| `VADDWEVQV` | |
| `VADDWEVQVU` | |
| `VADDWEVVW` | |
| `VADDWEVVWU` | |
| `VADDWEVWH` | |
| `VADDWEVWHU` | |
| `VADDWODHB` | |
| `VADDWODHBU` | |
| `VADDWODQV` | |
| `VADDWODQVU` | |
| `VADDWODVW` | |
| `VADDWODVWU` | |
| `VADDWODWH` | |
| `VADDWODWHU` | |
| `VADDWU` | |
| `VANDB` | |
| `VANDNV` | |
| `VANDV` | |
| `VBITCLRB` | |
| `VBITCLRH` | |
| `VBITCLRV` | |
| `VBITCLRW` | |
| `VBITREVB` | |
| `VBITREVH` | |
| `VBITREVV` | |
| `VBITREVW` | |
| `VBITSETB` | |
| `VBITSETH` | |
| `VBITSETV` | |
| `VBITSETW` | |
| `VDIVB` | |
| `VDIVBU` | |
| `VDIVD` | |
| `VDIVF` | |
| `VDIVH` | |
| `VDIVHU` | |
| `VDIVV` | |
| `VDIVVU` | |
| `VDIVW` | |
| `VDIVWU` | |
| `VEXTRINSB` | |
| `VEXTRINSH` | |
| `VEXTRINSV` | |
| `VEXTRINSW` | |
| `VFCLASSD` | |
| `VFCLASSF` | |
| `VFRECIPD` | |
| `VFRECIPF` | |
| `VFRINTD` | |
| `VFRINTF` | |
| `VFRINTRMD` | |
| `VFRINTRMF` | |
| `VFRINTRNED` | |
| `VFRINTRNEF` | |
| `VFRINTRPD` | |
| `VFRINTRPF` | |
| `VFRINTRZD` | |
| `VFRINTRZF` | |
| `VFRSQRTD` | |
| `VFRSQRTF` | |
| `VFSQRTD` | |
| `VFSQRTF` | |
| `VILVHB` | |
| `VILVHH` | |
| `VILVHV` | |
| `VILVHW` | |
| `VILVLB` | |
| `VILVLH` | |
| `VILVLV` | |
| `VILVLW` | |
| `VMADDB` | |
| `VMADDH` | |
| `VMADDV` | |
| `VMADDW` | |
| `VMADDWEVHB` | |
| `VMADDWEVHBU` | |
| `VMADDWEVHBUB` | |
| `VMADDWEVQV` | |
| `VMADDWEVQVU` | |
| `VMADDWEVQVUV` | |
| `VMADDWEVVW` | |
| `VMADDWEVVWU` | |
| `VMADDWEVVWUW` | |
| `VMADDWEVWH` | |
| `VMADDWEVWHU` | |
| `VMADDWEVWHUH` | |
| `VMADDWODHB` | |
| `VMADDWODHBU` | |
| `VMADDWODHBUB` | |
| `VMADDWODQV` | |
| `VMADDWODQVU` | |
| `VMADDWODQVUV` | |
| `VMADDWODVW` | |
| `VMADDWODVWU` | |
| `VMADDWODVWUW` | |
| `VMADDWODWH` | |
| `VMADDWODWHU` | |
| `VMADDWODWHUH` | |
| `VMODB` | |
| `VMODBU` | |
| `VMODH` | |
| `VMODHU` | |
| `VMODV` | |
| `VMODVU` | |
| `VMODW` | |
| `VMODWU` | |
| `VMOVQ` | |
| `VMSUBB` | |
| `VMSUBH` | |
| `VMSUBV` | |
| `VMSUBW` | |
| `VMUHB` | |
| `VMUHBU` | |
| `VMUHH` | |
| `VMUHHU` | |
| `VMUHV` | |
| `VMUHVU` | |
| `VMUHW` | |
| `VMUHWU` | |
| `VMULB` | |
| `VMULD` | |
| `VMULF` | |
| `VMULH` | |
| `VMULV` | |
| `VMULW` | |
| `VMULWEVHB` | |
| `VMULWEVHBU` | |
| `VMULWEVHBUB` | |
| `VMULWEVQV` | |
| `VMULWEVQVU` | |
| `VMULWEVQVUV` | |
| `VMULWEVVW` | |
| `VMULWEVVWU` | |
| `VMULWEVVWUW` | |
| `VMULWEVWH` | |
| `VMULWEVWHU` | |
| `VMULWEVWHUH` | |
| `VMULWODHB` | |
| `VMULWODHBU` | |
| `VMULWODHBUB` | |
| `VMULWODQV` | |
| `VMULWODQVU` | |
| `VMULWODQVUV` | |
| `VMULWODVW` | |
| `VMULWODVWU` | |
| `VMULWODVWUW` | |
| `VMULWODWH` | |
| `VMULWODWHU` | |
| `VMULWODWHUH` | |
| `VNEGB` | |
| `VNEGH` | |
| `VNEGV` | |
| `VNEGW` | |
| `VNORB` | |
| `VNORV` | |
| `VORB` | |
| `VORNV` | |
| `VORV` | |
| `VPCNTB` | |
| `VPCNTH` | |
| `VPCNTV` | |
| `VPCNTW` | |
| `VPERMIW` | |
| `VROTRB` | |
| `VROTRH` | |
| `VROTRV` | |
| `VROTRW` | |
| `VSADDB` | |
| `VSADDBU` | |
| `VSADDH` | |
| `VSADDHU` | |
| `VSADDV` | |
| `VSADDVU` | |
| `VSADDW` | |
| `VSADDWU` | |
| `VSEQB` | |
| `VSEQH` | |
| `VSEQV` | |
| `VSEQW` | |
| `VSETALLNEB` | |
| `VSETALLNEH` | |
| `VSETALLNEV` | |
| `VSETALLNEW` | |
| `VSETANYEQB` | |
| `VSETANYEQH` | |
| `VSETANYEQV` | |
| `VSETANYEQW` | |
| `VSETEQV` | |
| `VSETNEV` | |
| `VSHUF4IB` | |
| `VSHUF4IH` | |
| `VSHUF4IV` | |
| `VSHUF4IW` | |
| `VSHUFB` | |
| `VSHUFH` | |
| `VSHUFV` | |
| `VSHUFW` | |
| `VSLLB` | |
| `VSLLH` | |
| `VSLLV` | |
| `VSLLW` | |
| `VSLTB` | |
| `VSLTBU` | |
| `VSLTH` | |
| `VSLTHU` | |
| `VSLTV` | |
| `VSLTVU` | |
| `VSLTW` | |
| `VSLTWU` | |
| `VSRAB` | |
| `VSRAH` | |
| `VSRAV` | |
| `VSRAW` | |
| `VSRLB` | |
| `VSRLH` | |
| `VSRLV` | |
| `VSRLW` | |
| `VSSUBB` | |
| `VSSUBBU` | |
| `VSSUBH` | |
| `VSSUBHU` | |
| `VSSUBV` | |
| `VSSUBVU` | |
| `VSSUBW` | |
| `VSSUBWU` | |
| `VSUBB` | |
| `VSUBBU` | |
| `VSUBD` | |
| `VSUBF` | |
| `VSUBH` | |
| `VSUBHU` | |
| `VSUBQ` | |
| `VSUBV` | |
| `VSUBVU` | |
| `VSUBW` | |
| `VSUBWEVHB` | |
| `VSUBWEVHBU` | |
| `VSUBWEVQV` | |
| `VSUBWEVQVU` | |
| `VSUBWEVVW` | |
| `VSUBWEVVWU` | |
| `VSUBWEVWH` | |
| `VSUBWEVWHU` | |
| `VSUBWODHB` | |
| `VSUBWODHBU` | |
| `VSUBWODQV` | |
| `VSUBWODQVU` | |
| `VSUBWODVW` | |
| `VSUBWODVWU` | |
| `VSUBWODWH` | |
| `VSUBWODWHU` | |
| `VSUBWU` | |
| `VXORB` | |
| `VXORV` | |
| `WORD` | |
| `XOR` | Bitwise XOR |
| `XVADDB` | |
| `XVADDBU` | |
| `XVADDD` | |
| `XVADDF` | |
| `XVADDH` | |
| `XVADDHU` | |
| `XVADDQ` | |
| `XVADDV` | |
| `XVADDVU` | |
| `XVADDW` | |
| `XVADDWEVHB` | |
| `XVADDWEVHBU` | |
| `XVADDWEVQV` | |
| `XVADDWEVQVU` | |
| `XVADDWEVVW` | |
| `XVADDWEVVWU` | |
| `XVADDWEVWH` | |
| `XVADDWEVWHU` | |
| `XVADDWODHB` | |
| `XVADDWODHBU` | |
| `XVADDWODQV` | |
| `XVADDWODQVU` | |
| `XVADDWODVW` | |
| `XVADDWODVWU` | |
| `XVADDWODWH` | |
| `XVADDWODWHU` | |
| `XVADDWU` | |
| `XVANDB` | |
| `XVANDNV` | |
| `XVANDV` | |
| `XVBITCLRB` | |
| `XVBITCLRH` | |
| `XVBITCLRV` | |
| `XVBITCLRW` | |
| `XVBITREVB` | |
| `XVBITREVH` | |
| `XVBITREVV` | |
| `XVBITREVW` | |
| `XVBITSETB` | |
| `XVBITSETH` | |
| `XVBITSETV` | |
| `XVBITSETW` | |
| `XVDIVB` | |
| `XVDIVBU` | |
| `XVDIVD` | |
| `XVDIVF` | |
| `XVDIVH` | |
| `XVDIVHU` | |
| `XVDIVV` | |
| `XVDIVVU` | |
| `XVDIVW` | |
| `XVDIVWU` | |
| `XVEXTRINSB` | |
| `XVEXTRINSH` | |
| `XVEXTRINSV` | |
| `XVEXTRINSW` | |
| `XVFCLASSD` | |
| `XVFCLASSF` | |
| `XVFRECIPD` | |
| `XVFRECIPF` | |
| `XVFRINTD` | |
| `XVFRINTF` | |
| `XVFRINTRMD` | |
| `XVFRINTRMF` | |
| `XVFRINTRNED` | |
| `XVFRINTRNEF` | |
| `XVFRINTRPD` | |
| `XVFRINTRPF` | |
| `XVFRINTRZD` | |
| `XVFRINTRZF` | |
| `XVFRSQRTD` | |
| `XVFRSQRTF` | |
| `XVFSQRTD` | |
| `XVFSQRTF` | |
| `XVILVHB` | |
| `XVILVHH` | |
| `XVILVHV` | |
| `XVILVHW` | |
| `XVILVLB` | |
| `XVILVLH` | |
| `XVILVLV` | |
| `XVILVLW` | |
| `XVMADDB` | |
| `XVMADDH` | |
| `XVMADDV` | |
| `XVMADDW` | |
| `XVMADDWEVHB` | |
| `XVMADDWEVHBU` | |
| `XVMADDWEVHBUB` | |
| `XVMADDWEVQV` | |
| `XVMADDWEVQVU` | |
| `XVMADDWEVQVUV` | |
| `XVMADDWEVVW` | |
| `XVMADDWEVVWU` | |
| `XVMADDWEVVWUW` | |
| `XVMADDWEVWH` | |
| `XVMADDWEVWHU` | |
| `XVMADDWEVWHUH` | |
| `XVMADDWODHB` | |
| `XVMADDWODHBU` | |
| `XVMADDWODHBUB` | |
| `XVMADDWODQV` | |
| `XVMADDWODQVU` | |
| `XVMADDWODQVUV` | |
| `XVMADDWODVW` | |
| `XVMADDWODVWU` | |
| `XVMADDWODVWUW` | |
| `XVMADDWODWH` | |
| `XVMADDWODWHU` | |
| `XVMADDWODWHUH` | |
| `XVMODB` | |
| `XVMODBU` | |
| `XVMODH` | |
| `XVMODHU` | |
| `XVMODV` | |
| `XVMODVU` | |
| `XVMODW` | |
| `XVMODWU` | |
| `XVMOVQ` | |
| `XVMSUBB` | |
| `XVMSUBH` | |
| `XVMSUBV` | |
| `XVMSUBW` | |
| `XVMUHB` | |
| `XVMUHBU` | |
| `XVMUHH` | |
| `XVMUHHU` | |
| `XVMUHV` | |
| `XVMUHVU` | |
| `XVMUHW` | |
| `XVMUHWU` | |
| `XVMULB` | |
| `XVMULD` | |
| `XVMULF` | |
| `XVMULH` | |
| `XVMULV` | |
| `XVMULW` | |
| `XVMULWEVHB` | |
| `XVMULWEVHBU` | |
| `XVMULWEVHBUB` | |
| `XVMULWEVQV` | |
| `XVMULWEVQVU` | |
| `XVMULWEVQVUV` | |
| `XVMULWEVVW` | |
| `XVMULWEVVWU` | |
| `XVMULWEVVWUW` | |
| `XVMULWEVWH` | |
| `XVMULWEVWHU` | |
| `XVMULWEVWHUH` | |
| `XVMULWODHB` | |
| `XVMULWODHBU` | |
| `XVMULWODHBUB` | |
| `XVMULWODQV` | |
| `XVMULWODQVU` | |
| `XVMULWODQVUV` | |
| `XVMULWODVW` | |
| `XVMULWODVWU` | |
| `XVMULWODVWUW` | |
| `XVMULWODWH` | |
| `XVMULWODWHU` | |
| `XVMULWODWHUH` | |
| `XVNEGB` | |
| `XVNEGH` | |
| `XVNEGV` | |
| `XVNEGW` | |
| `XVNORB` | |
| `XVNORV` | |
| `XVORB` | |
| `XVORNV` | |
| `XVORV` | |
| `XVPCNTB` | |
| `XVPCNTH` | |
| `XVPCNTV` | |
| `XVPCNTW` | |
| `XVPERMIQ` | |
| `XVPERMIV` | |
| `XVPERMIW` | |
| `XVROTRB` | |
| `XVROTRH` | |
| `XVROTRV` | |
| `XVROTRW` | |
| `XVSADDB` | |
| `XVSADDBU` | |
| `XVSADDH` | |
| `XVSADDHU` | |
| `XVSADDV` | |
| `XVSADDVU` | |
| `XVSADDW` | |
| `XVSADDWU` | |
| `XVSEQB` | |
| `XVSEQH` | |
| `XVSEQV` | |
| `XVSEQW` | |
| `XVSETALLNEB` | |
| `XVSETALLNEH` | |
| `XVSETALLNEV` | |
| `XVSETALLNEW` | |
| `XVSETANYEQB` | |
| `XVSETANYEQH` | |
| `XVSETANYEQV` | |
| `XVSETANYEQW` | |
| `XVSETEQV` | |
| `XVSETNEV` | |
| `XVSHUF4IB` | |
| `XVSHUF4IH` | |
| `XVSHUF4IV` | |
| `XVSHUF4IW` | |
| `XVSHUFB` | |
| `XVSHUFH` | |
| `XVSHUFV` | |
| `XVSHUFW` | |
| `XVSLLB` | |
| `XVSLLH` | |
| `XVSLLV` | |
| `XVSLLW` | |
| `XVSLTB` | |
| `XVSLTBU` | |
| `XVSLTH` | |
| `XVSLTHU` | |
| `XVSLTV` | |
| `XVSLTVU` | |
| `XVSLTW` | |
| `XVSLTWU` | |
| `XVSRAB` | |
| `XVSRAH` | |
| `XVSRAV` | |
| `XVSRAW` | |
| `XVSRLB` | |
| `XVSRLH` | |
| `XVSRLV` | |
| `XVSRLW` | |
| `XVSSUBB` | |
| `XVSSUBBU` | |
| `XVSSUBH` | |
| `XVSSUBHU` | |
| `XVSSUBV` | |
| `XVSSUBVU` | |
| `XVSSUBW` | |
| `XVSSUBWU` | |
| `XVSUBB` | |
| `XVSUBBU` | |
| `XVSUBD` | |
| `XVSUBF` | |
| `XVSUBH` | |
| `XVSUBHU` | |
| `XVSUBQ` | |
| `XVSUBV` | |
| `XVSUBVU` | |
| `XVSUBW` | |
| `XVSUBWEVHB` | |
| `XVSUBWEVHBU` | |
| `XVSUBWEVQV` | |
| `XVSUBWEVQVU` | |
| `XVSUBWEVVW` | |
| `XVSUBWEVVWU` | |
| `XVSUBWEVWH` | |
| `XVSUBWEVWHU` | |
| `XVSUBWODHB` | |
| `XVSUBWODHBU` | |
| `XVSUBWODQV` | |
| `XVSUBWODQVU` | |
| `XVSUBWODVW` | |
| `XVSUBWODVWU` | |
| `XVSUBWODWH` | |
| `XVSUBWODWHU` | |
| `XVSUBWU` | |
| `XVXORB` | |
| `XVXORV` | |
| `JAL` | |
Recognised: 814 mnemonics.
+991
View File
@@ -0,0 +1,991 @@
# RISC-V 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/riscv/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | Call subroutine |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | Unconditional jump |
| `NOP` | |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADD` | Integer add |
| `ADDI` | Add immediate |
| `ADDIW` | Add immediate (32-bit) |
| `ADDUW` | |
| `ADDW` | Add (32-bit) |
| `AMOADDD` | Atomic add doubleword |
| `AMOADDW` | Atomic add word |
| `AMOANDD` | |
| `AMOANDW` | |
| `AMOMAXD` | |
| `AMOMAXUD` | |
| `AMOMAXUW` | |
| `AMOMAXW` | |
| `AMOMIND` | |
| `AMOMINUD` | |
| `AMOMINUW` | |
| `AMOMINW` | |
| `AMOORD` | |
| `AMOORW` | |
| `AMOSWAPD` | Atomic swap doubleword |
| `AMOSWAPW` | Atomic swap word |
| `AMOXORD` | |
| `AMOXORW` | |
| `AND` | Bitwise AND |
| `ANDI` | AND immediate |
| `ANDN` | |
| `AUIPC` | Add upper immediate to PC |
| `BCLR` | |
| `BCLRI` | |
| `BEQ` | Branch if equal |
| `BEQZ` | |
| `BEXT` | |
| `BEXTI` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGT` | |
| `BGTU` | |
| `BGTZ` | |
| `BINV` | |
| `BINVI` | |
| `BLE` | |
| `BLEU` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BNEZ` | |
| `BSET` | |
| `BSETI` | |
| `CADD` | |
| `CADDI` | |
| `CADDI16SP` | |
| `CADDI4SPN` | |
| `CADDIW` | |
| `CADDW` | |
| `CAND` | |
| `CANDI` | |
| `CBEQZ` | |
| `CBNEZ` | |
| `CEBREAK` | |
| `CFLD` | |
| `CFLDSP` | |
| `CFSD` | |
| `CFSDSP` | |
| `CJ` | |
| `CJALR` | |
| `CJR` | |
| `CLD` | |
| `CLDSP` | |
| `CLI` | |
| `CLUI` | |
| `CLW` | |
| `CLWSP` | |
| `CLZ` | |
| `CLZW` | |
| `CMV` | |
| `CNOP` | |
| `COR` | |
| `CPOP` | |
| `CPOPW` | |
| `CSD` | |
| `CSDSP` | |
| `CSLLI` | |
| `CSRAI` | |
| `CSRLI` | |
| `CSRRC` | |
| `CSRRCI` | |
| `CSRRS` | |
| `CSRRSI` | |
| `CSRRW` | |
| `CSRRWI` | |
| `CSUB` | |
| `CSUBW` | |
| `CSW` | |
| `CSWSP` | |
| `CTZ` | |
| `CTZW` | |
| `CXOR` | |
| `CZEROEQZ` | |
| `CZERONEZ` | |
| `DIV` | Divide |
| `DIVU` | Divide unsigned |
| `DIVUW` | |
| `DIVW` | Divide (32-bit) |
| `DRET` | |
| `EBREAK` | Breakpoint |
| `ECALL` | Environment call |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | FP add (double) |
| `FADDQ` | |
| `FADDS` | FP add (single) |
| `FCLASSD` | |
| `FCLASSQ` | |
| `FCLASSS` | |
| `FCVTDL` | |
| `FCVTDLU` | |
| `FCVTDQ` | |
| `FCVTDS` | |
| `FCVTDW` | |
| `FCVTDWU` | |
| `FCVTLD` | |
| `FCVTLQ` | |
| `FCVTLS` | |
| `FCVTLUD` | |
| `FCVTLUQ` | |
| `FCVTLUS` | |
| `FCVTQD` | |
| `FCVTQL` | |
| `FCVTQLU` | |
| `FCVTQS` | |
| `FCVTQW` | |
| `FCVTQWU` | |
| `FCVTSD` | |
| `FCVTSL` | |
| `FCVTSLU` | |
| `FCVTSQ` | |
| `FCVTSW` | |
| `FCVTSWU` | |
| `FCVTWD` | |
| `FCVTWQ` | |
| `FCVTWS` | |
| `FCVTWUD` | |
| `FCVTWUQ` | |
| `FCVTWUS` | |
| `FDIVD` | FP divide (double) |
| `FDIVQ` | |
| `FDIVS` | FP divide (single) |
| `FENCE` | Memory barrier |
| `FEQD` | |
| `FEQQ` | |
| `FEQS` | |
| `FLD` | FP load doubleword |
| `FLED` | |
| `FLEQ` | |
| `FLES` | |
| `FLQ` | |
| `FLTD` | |
| `FLTQ` | |
| `FLTS` | |
| `FLW` | FP load word |
| `FMADDD` | |
| `FMADDQ` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXQ` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINQ` | |
| `FMINS` | |
| `FMSUBD` | |
| `FMSUBQ` | |
| `FMSUBS` | |
| `FMULD` | FP multiply (double) |
| `FMULQ` | |
| `FMULS` | FP multiply (single) |
| `FMVDX` | |
| `FMVSX` | |
| `FMVWX` | |
| `FMVXD` | |
| `FMVXS` | |
| `FMVXW` | |
| `FNED` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNES` | |
| `FNMADDD` | |
| `FNMADDQ` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBQ` | |
| `FNMSUBS` | |
| `FSD` | FP store doubleword |
| `FSGNJD` | |
| `FSGNJND` | |
| `FSGNJNQ` | |
| `FSGNJNS` | |
| `FSGNJQ` | |
| `FSGNJS` | |
| `FSGNJXD` | |
| `FSGNJXQ` | |
| `FSGNJXS` | |
| `FSQ` | |
| `FSQRTD` | |
| `FSQRTQ` | |
| `FSQRTS` | |
| `FSUBD` | FP subtract (double) |
| `FSUBQ` | |
| `FSUBS` | FP subtract (single) |
| `FSW` | FP store word |
| `JAL` | Jump and link |
| `JALR` | Jump and link register |
| `LB` | Load byte |
| `LBU` | Load byte unsigned |
| `LD` | Load doubleword |
| `LH` | Load halfword |
| `LHU` | Load halfword unsigned |
| `LRD` | Load-reserved doubleword |
| `LRW` | Load-reserved word |
| `LUI` | Load upper immediate |
| `LW` | Load word |
| `LWU` | Load word unsigned |
| `MAX` | |
| `MAXU` | |
| `MIN` | |
| `MINU` | |
| `MOV` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVF` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVW` | |
| `MOVWU` | |
| `MRET` | |
| `MUL` | Multiply |
| `MULH` | Multiply high |
| `MULHSU` | Multiply high signed/unsigned |
| `MULHU` | Multiply high unsigned |
| `MULW` | Multiply (32-bit) |
| `NEG` | |
| `NEGW` | |
| `NOT` | |
| `OR` | Bitwise OR |
| `ORCB` | |
| `ORI` | OR immediate |
| `ORN` | |
| `RDCYCLE` | |
| `RDINSTRET` | |
| `RDTIME` | |
| `REM` | Remainder |
| `REMU` | Remainder unsigned |
| `REMUW` | |
| `REMW` | |
| `REV8` | |
| `ROL` | |
| `ROLW` | |
| `ROR` | |
| `RORI` | |
| `RORIW` | |
| `RORW` | |
| `SB` | Store byte |
| `SBREAK` | |
| `SCALL` | |
| `SCD` | Store-conditional doubleword |
| `SCW` | Store-conditional word |
| `SD` | Store doubleword |
| `SEQZ` | |
| `SEXTB` | |
| `SEXTH` | |
| `SFENCEVMA` | |
| `SH` | Store halfword |
| `SH1ADD` | |
| `SH1ADDUW` | |
| `SH2ADD` | |
| `SH2ADDUW` | |
| `SH3ADD` | |
| `SH3ADDUW` | |
| `SLL` | Shift left logical |
| `SLLI` | Shift left logical immediate |
| `SLLIUW` | |
| `SLLIW` | |
| `SLLW` | |
| `SLT` | Set if less than |
| `SLTI` | Set if less than immediate |
| `SLTIU` | Set if less than unsigned immediate |
| `SLTU` | Set if less than unsigned |
| `SNEZ` | |
| `SRA` | Shift right arithmetic |
| `SRAI` | Shift right arithmetic immediate |
| `SRAIW` | |
| `SRAW` | |
| `SRET` | |
| `SRL` | Shift right logical |
| `SRLI` | Shift right logical immediate |
| `SRLIW` | |
| `SRLW` | |
| `SUB` | Integer subtract |
| `SUBW` | Subtract (32-bit) |
| `SW` | Store word |
| `VAADDUVV` | |
| `VAADDUVX` | |
| `VAADDVV` | |
| `VAADDVX` | |
| `VADCVIM` | |
| `VADCVVM` | |
| `VADCVXM` | |
| `VADDVI` | |
| `VADDVV` | |
| `VADDVX` | |
| `VANDVI` | |
| `VANDVV` | |
| `VANDVX` | |
| `VASUBUVV` | |
| `VASUBUVX` | |
| `VASUBVV` | |
| `VASUBVX` | |
| `VCOMPRESSVM` | |
| `VCPOPM` | |
| `VDIVUVV` | |
| `VDIVUVX` | |
| `VDIVVV` | |
| `VDIVVX` | |
| `VFABSV` | |
| `VFADDVF` | |
| `VFADDVV` | |
| `VFCLASSV` | |
| `VFCVTFXUV` | |
| `VFCVTFXV` | |
| `VFCVTRTZXFV` | |
| `VFCVTRTZXUFV` | |
| `VFCVTXFV` | |
| `VFCVTXUFV` | |
| `VFDIVVF` | |
| `VFDIVVV` | |
| `VFIRSTM` | |
| `VFMACCVF` | |
| `VFMACCVV` | |
| `VFMADDVF` | |
| `VFMADDVV` | |
| `VFMAXVF` | |
| `VFMAXVV` | |
| `VFMERGEVFM` | |
| `VFMINVF` | |
| `VFMINVV` | |
| `VFMSACVF` | |
| `VFMSACVV` | |
| `VFMSUBVF` | |
| `VFMSUBVV` | |
| `VFMULVF` | |
| `VFMULVV` | |
| `VFMVFS` | |
| `VFMVSF` | |
| `VFMVVF` | |
| `VFNCVTFFW` | |
| `VFNCVTFXUW` | |
| `VFNCVTFXW` | |
| `VFNCVTRODFFW` | |
| `VFNCVTRTZXFW` | |
| `VFNCVTRTZXUFW` | |
| `VFNCVTXFW` | |
| `VFNCVTXUFW` | |
| `VFNEGV` | |
| `VFNMACCVF` | |
| `VFNMACCVV` | |
| `VFNMADDVF` | |
| `VFNMADDVV` | |
| `VFNMSACVF` | |
| `VFNMSACVV` | |
| `VFNMSUBVF` | |
| `VFNMSUBVV` | |
| `VFRDIVVF` | |
| `VFREC7V` | |
| `VFREDMAXVS` | |
| `VFREDMINVS` | |
| `VFREDOSUMVS` | |
| `VFREDUSUMVS` | |
| `VFRSQRT7V` | |
| `VFRSUBVF` | |
| `VFSGNJNVF` | |
| `VFSGNJNVV` | |
| `VFSGNJVF` | |
| `VFSGNJVV` | |
| `VFSGNJXVF` | |
| `VFSGNJXVV` | |
| `VFSLIDE1DOWNVF` | |
| `VFSLIDE1UPVF` | |
| `VFSQRTV` | |
| `VFSUBVF` | |
| `VFSUBVV` | |
| `VFWADDVF` | |
| `VFWADDVV` | |
| `VFWADDWF` | |
| `VFWADDWV` | |
| `VFWCVTFFV` | |
| `VFWCVTFXUV` | |
| `VFWCVTFXV` | |
| `VFWCVTRTZXFV` | |
| `VFWCVTRTZXUFV` | |
| `VFWCVTXFV` | |
| `VFWCVTXUFV` | |
| `VFWMACCVF` | |
| `VFWMACCVV` | |
| `VFWMSACVF` | |
| `VFWMSACVV` | |
| `VFWMULVF` | |
| `VFWMULVV` | |
| `VFWNMACCVF` | |
| `VFWNMACCVV` | |
| `VFWNMSACVF` | |
| `VFWNMSACVV` | |
| `VFWREDOSUMVS` | |
| `VFWREDUSUMVS` | |
| `VFWSUBVF` | |
| `VFWSUBVV` | |
| `VFWSUBWF` | |
| `VFWSUBWV` | |
| `VIDV` | |
| `VIOTAM` | |
| `VL1RE16V` | |
| `VL1RE32V` | |
| `VL1RE64V` | |
| `VL1RE8V` | |
| `VL1RV` | |
| `VL2RE16V` | |
| `VL2RE32V` | |
| `VL2RE64V` | |
| `VL2RE8V` | |
| `VL2RV` | |
| `VL4RE16V` | |
| `VL4RE32V` | |
| `VL4RE64V` | |
| `VL4RE8V` | |
| `VL4RV` | |
| `VL8RE16V` | |
| `VL8RE32V` | |
| `VL8RE64V` | |
| `VL8RE8V` | |
| `VL8RV` | |
| `VLE16FFV` | |
| `VLE16V` | |
| `VLE32FFV` | |
| `VLE32V` | |
| `VLE64FFV` | |
| `VLE64V` | |
| `VLE8FFV` | |
| `VLE8V` | |
| `VLMV` | |
| `VLOXEI16V` | |
| `VLOXEI32V` | |
| `VLOXEI64V` | |
| `VLOXEI8V` | |
| `VLOXSEG2EI16V` | |
| `VLOXSEG2EI32V` | |
| `VLOXSEG2EI64V` | |
| `VLOXSEG2EI8V` | |
| `VLOXSEG3EI16V` | |
| `VLOXSEG3EI32V` | |
| `VLOXSEG3EI64V` | |
| `VLOXSEG3EI8V` | |
| `VLOXSEG4EI16V` | |
| `VLOXSEG4EI32V` | |
| `VLOXSEG4EI64V` | |
| `VLOXSEG4EI8V` | |
| `VLOXSEG5EI16V` | |
| `VLOXSEG5EI32V` | |
| `VLOXSEG5EI64V` | |
| `VLOXSEG5EI8V` | |
| `VLOXSEG6EI16V` | |
| `VLOXSEG6EI32V` | |
| `VLOXSEG6EI64V` | |
| `VLOXSEG6EI8V` | |
| `VLOXSEG7EI16V` | |
| `VLOXSEG7EI32V` | |
| `VLOXSEG7EI64V` | |
| `VLOXSEG7EI8V` | |
| `VLOXSEG8EI16V` | |
| `VLOXSEG8EI32V` | |
| `VLOXSEG8EI64V` | |
| `VLOXSEG8EI8V` | |
| `VLSE16V` | |
| `VLSE32V` | |
| `VLSE64V` | |
| `VLSE8V` | |
| `VLSEG2E16FFV` | |
| `VLSEG2E16V` | |
| `VLSEG2E32FFV` | |
| `VLSEG2E32V` | |
| `VLSEG2E64FFV` | |
| `VLSEG2E64V` | |
| `VLSEG2E8FFV` | |
| `VLSEG2E8V` | |
| `VLSEG3E16FFV` | |
| `VLSEG3E16V` | |
| `VLSEG3E32FFV` | |
| `VLSEG3E32V` | |
| `VLSEG3E64FFV` | |
| `VLSEG3E64V` | |
| `VLSEG3E8FFV` | |
| `VLSEG3E8V` | |
| `VLSEG4E16FFV` | |
| `VLSEG4E16V` | |
| `VLSEG4E32FFV` | |
| `VLSEG4E32V` | |
| `VLSEG4E64FFV` | |
| `VLSEG4E64V` | |
| `VLSEG4E8FFV` | |
| `VLSEG4E8V` | |
| `VLSEG5E16FFV` | |
| `VLSEG5E16V` | |
| `VLSEG5E32FFV` | |
| `VLSEG5E32V` | |
| `VLSEG5E64FFV` | |
| `VLSEG5E64V` | |
| `VLSEG5E8FFV` | |
| `VLSEG5E8V` | |
| `VLSEG6E16FFV` | |
| `VLSEG6E16V` | |
| `VLSEG6E32FFV` | |
| `VLSEG6E32V` | |
| `VLSEG6E64FFV` | |
| `VLSEG6E64V` | |
| `VLSEG6E8FFV` | |
| `VLSEG6E8V` | |
| `VLSEG7E16FFV` | |
| `VLSEG7E16V` | |
| `VLSEG7E32FFV` | |
| `VLSEG7E32V` | |
| `VLSEG7E64FFV` | |
| `VLSEG7E64V` | |
| `VLSEG7E8FFV` | |
| `VLSEG7E8V` | |
| `VLSEG8E16FFV` | |
| `VLSEG8E16V` | |
| `VLSEG8E32FFV` | |
| `VLSEG8E32V` | |
| `VLSEG8E64FFV` | |
| `VLSEG8E64V` | |
| `VLSEG8E8FFV` | |
| `VLSEG8E8V` | |
| `VLSSEG2E16V` | |
| `VLSSEG2E32V` | |
| `VLSSEG2E64V` | |
| `VLSSEG2E8V` | |
| `VLSSEG3E16V` | |
| `VLSSEG3E32V` | |
| `VLSSEG3E64V` | |
| `VLSSEG3E8V` | |
| `VLSSEG4E16V` | |
| `VLSSEG4E32V` | |
| `VLSSEG4E64V` | |
| `VLSSEG4E8V` | |
| `VLSSEG5E16V` | |
| `VLSSEG5E32V` | |
| `VLSSEG5E64V` | |
| `VLSSEG5E8V` | |
| `VLSSEG6E16V` | |
| `VLSSEG6E32V` | |
| `VLSSEG6E64V` | |
| `VLSSEG6E8V` | |
| `VLSSEG7E16V` | |
| `VLSSEG7E32V` | |
| `VLSSEG7E64V` | |
| `VLSSEG7E8V` | |
| `VLSSEG8E16V` | |
| `VLSSEG8E32V` | |
| `VLSSEG8E64V` | |
| `VLSSEG8E8V` | |
| `VLUXEI16V` | |
| `VLUXEI32V` | |
| `VLUXEI64V` | |
| `VLUXEI8V` | |
| `VLUXSEG2EI16V` | |
| `VLUXSEG2EI32V` | |
| `VLUXSEG2EI64V` | |
| `VLUXSEG2EI8V` | |
| `VLUXSEG3EI16V` | |
| `VLUXSEG3EI32V` | |
| `VLUXSEG3EI64V` | |
| `VLUXSEG3EI8V` | |
| `VLUXSEG4EI16V` | |
| `VLUXSEG4EI32V` | |
| `VLUXSEG4EI64V` | |
| `VLUXSEG4EI8V` | |
| `VLUXSEG5EI16V` | |
| `VLUXSEG5EI32V` | |
| `VLUXSEG5EI64V` | |
| `VLUXSEG5EI8V` | |
| `VLUXSEG6EI16V` | |
| `VLUXSEG6EI32V` | |
| `VLUXSEG6EI64V` | |
| `VLUXSEG6EI8V` | |
| `VLUXSEG7EI16V` | |
| `VLUXSEG7EI32V` | |
| `VLUXSEG7EI64V` | |
| `VLUXSEG7EI8V` | |
| `VLUXSEG8EI16V` | |
| `VLUXSEG8EI32V` | |
| `VLUXSEG8EI64V` | |
| `VLUXSEG8EI8V` | |
| `VMACCVV` | |
| `VMACCVX` | |
| `VMADCVI` | |
| `VMADCVIM` | |
| `VMADCVV` | |
| `VMADCVVM` | |
| `VMADCVX` | |
| `VMADCVXM` | |
| `VMADDVV` | |
| `VMADDVX` | |
| `VMANDMM` | |
| `VMANDNMM` | |
| `VMAXUVV` | |
| `VMAXUVX` | |
| `VMAXVV` | |
| `VMAXVX` | |
| `VMCLRM` | |
| `VMERGEVIM` | |
| `VMERGEVVM` | |
| `VMERGEVXM` | |
| `VMFEQVF` | |
| `VMFEQVV` | |
| `VMFGEVF` | |
| `VMFGEVV` | |
| `VMFGTVF` | |
| `VMFGTVV` | |
| `VMFLEVF` | |
| `VMFLEVV` | |
| `VMFLTVF` | |
| `VMFLTVV` | |
| `VMFNEVF` | |
| `VMFNEVV` | |
| `VMINUVV` | |
| `VMINUVX` | |
| `VMINVV` | |
| `VMINVX` | |
| `VMMVM` | |
| `VMNANDMM` | |
| `VMNORMM` | |
| `VMNOTM` | |
| `VMORMM` | |
| `VMORNMM` | |
| `VMSBCVV` | |
| `VMSBCVVM` | |
| `VMSBCVX` | |
| `VMSBCVXM` | |
| `VMSBFM` | |
| `VMSEQVI` | |
| `VMSEQVV` | |
| `VMSEQVX` | |
| `VMSETM` | |
| `VMSGEUVI` | |
| `VMSGEUVV` | |
| `VMSGEVI` | |
| `VMSGEVV` | |
| `VMSGTUVI` | |
| `VMSGTUVV` | |
| `VMSGTUVX` | |
| `VMSGTVI` | |
| `VMSGTVV` | |
| `VMSGTVX` | |
| `VMSIFM` | |
| `VMSLEUVI` | |
| `VMSLEUVV` | |
| `VMSLEUVX` | |
| `VMSLEVI` | |
| `VMSLEVV` | |
| `VMSLEVX` | |
| `VMSLTUVI` | |
| `VMSLTUVV` | |
| `VMSLTUVX` | |
| `VMSLTVI` | |
| `VMSLTVV` | |
| `VMSLTVX` | |
| `VMSNEVI` | |
| `VMSNEVV` | |
| `VMSNEVX` | |
| `VMSOFM` | |
| `VMULHSUVV` | |
| `VMULHSUVX` | |
| `VMULHUVV` | |
| `VMULHUVX` | |
| `VMULHVV` | |
| `VMULHVX` | |
| `VMULVV` | |
| `VMULVX` | |
| `VMV1RV` | |
| `VMV2RV` | |
| `VMV4RV` | |
| `VMV8RV` | |
| `VMVSX` | |
| `VMVVI` | |
| `VMVVV` | |
| `VMVVX` | |
| `VMVXS` | |
| `VMXNORMM` | |
| `VMXORMM` | |
| `VNCLIPUWI` | |
| `VNCLIPUWV` | |
| `VNCLIPUWX` | |
| `VNCLIPWI` | |
| `VNCLIPWV` | |
| `VNCLIPWX` | |
| `VNCVTXXW` | |
| `VNEGV` | |
| `VNMSACVV` | |
| `VNMSACVX` | |
| `VNMSUBVV` | |
| `VNMSUBVX` | |
| `VNOTV` | |
| `VNSRAWI` | |
| `VNSRAWV` | |
| `VNSRAWX` | |
| `VNSRLWI` | |
| `VNSRLWV` | |
| `VNSRLWX` | |
| `VORVI` | |
| `VORVV` | |
| `VORVX` | |
| `VREDANDVS` | |
| `VREDMAXUVS` | |
| `VREDMAXVS` | |
| `VREDMINUVS` | |
| `VREDMINVS` | |
| `VREDORVS` | |
| `VREDSUMVS` | |
| `VREDXORVS` | |
| `VREMUVV` | |
| `VREMUVX` | |
| `VREMVV` | |
| `VREMVX` | |
| `VRGATHEREI16VV` | |
| `VRGATHERVI` | |
| `VRGATHERVV` | |
| `VRGATHERVX` | |
| `VRSUBVI` | |
| `VRSUBVX` | |
| `VS1RV` | |
| `VS2RV` | |
| `VS4RV` | |
| `VS8RV` | |
| `VSADDUVI` | |
| `VSADDUVV` | |
| `VSADDUVX` | |
| `VSADDVI` | |
| `VSADDVV` | |
| `VSADDVX` | |
| `VSBCVVM` | |
| `VSBCVXM` | |
| `VSE16V` | |
| `VSE32V` | |
| `VSE64V` | |
| `VSE8V` | |
| `VSETIVLI` | |
| `VSETVL` | |
| `VSETVLI` | |
| `VSEXTVF2` | |
| `VSEXTVF4` | |
| `VSEXTVF8` | |
| `VSLIDE1DOWNVX` | |
| `VSLIDE1UPVX` | |
| `VSLIDEDOWNVI` | |
| `VSLIDEDOWNVX` | |
| `VSLIDEUPVI` | |
| `VSLIDEUPVX` | |
| `VSLLVI` | |
| `VSLLVV` | |
| `VSLLVX` | |
| `VSMULVV` | |
| `VSMULVX` | |
| `VSMV` | |
| `VSOXEI16V` | |
| `VSOXEI32V` | |
| `VSOXEI64V` | |
| `VSOXEI8V` | |
| `VSOXSEG2EI16V` | |
| `VSOXSEG2EI32V` | |
| `VSOXSEG2EI64V` | |
| `VSOXSEG2EI8V` | |
| `VSOXSEG3EI16V` | |
| `VSOXSEG3EI32V` | |
| `VSOXSEG3EI64V` | |
| `VSOXSEG3EI8V` | |
| `VSOXSEG4EI16V` | |
| `VSOXSEG4EI32V` | |
| `VSOXSEG4EI64V` | |
| `VSOXSEG4EI8V` | |
| `VSOXSEG5EI16V` | |
| `VSOXSEG5EI32V` | |
| `VSOXSEG5EI64V` | |
| `VSOXSEG5EI8V` | |
| `VSOXSEG6EI16V` | |
| `VSOXSEG6EI32V` | |
| `VSOXSEG6EI64V` | |
| `VSOXSEG6EI8V` | |
| `VSOXSEG7EI16V` | |
| `VSOXSEG7EI32V` | |
| `VSOXSEG7EI64V` | |
| `VSOXSEG7EI8V` | |
| `VSOXSEG8EI16V` | |
| `VSOXSEG8EI32V` | |
| `VSOXSEG8EI64V` | |
| `VSOXSEG8EI8V` | |
| `VSRAVI` | |
| `VSRAVV` | |
| `VSRAVX` | |
| `VSRLVI` | |
| `VSRLVV` | |
| `VSRLVX` | |
| `VSSE16V` | |
| `VSSE32V` | |
| `VSSE64V` | |
| `VSSE8V` | |
| `VSSEG2E16V` | |
| `VSSEG2E32V` | |
| `VSSEG2E64V` | |
| `VSSEG2E8V` | |
| `VSSEG3E16V` | |
| `VSSEG3E32V` | |
| `VSSEG3E64V` | |
| `VSSEG3E8V` | |
| `VSSEG4E16V` | |
| `VSSEG4E32V` | |
| `VSSEG4E64V` | |
| `VSSEG4E8V` | |
| `VSSEG5E16V` | |
| `VSSEG5E32V` | |
| `VSSEG5E64V` | |
| `VSSEG5E8V` | |
| `VSSEG6E16V` | |
| `VSSEG6E32V` | |
| `VSSEG6E64V` | |
| `VSSEG6E8V` | |
| `VSSEG7E16V` | |
| `VSSEG7E32V` | |
| `VSSEG7E64V` | |
| `VSSEG7E8V` | |
| `VSSEG8E16V` | |
| `VSSEG8E32V` | |
| `VSSEG8E64V` | |
| `VSSEG8E8V` | |
| `VSSRAVI` | |
| `VSSRAVV` | |
| `VSSRAVX` | |
| `VSSRLVI` | |
| `VSSRLVV` | |
| `VSSRLVX` | |
| `VSSSEG2E16V` | |
| `VSSSEG2E32V` | |
| `VSSSEG2E64V` | |
| `VSSSEG2E8V` | |
| `VSSSEG3E16V` | |
| `VSSSEG3E32V` | |
| `VSSSEG3E64V` | |
| `VSSSEG3E8V` | |
| `VSSSEG4E16V` | |
| `VSSSEG4E32V` | |
| `VSSSEG4E64V` | |
| `VSSSEG4E8V` | |
| `VSSSEG5E16V` | |
| `VSSSEG5E32V` | |
| `VSSSEG5E64V` | |
| `VSSSEG5E8V` | |
| `VSSSEG6E16V` | |
| `VSSSEG6E32V` | |
| `VSSSEG6E64V` | |
| `VSSSEG6E8V` | |
| `VSSSEG7E16V` | |
| `VSSSEG7E32V` | |
| `VSSSEG7E64V` | |
| `VSSSEG7E8V` | |
| `VSSSEG8E16V` | |
| `VSSSEG8E32V` | |
| `VSSSEG8E64V` | |
| `VSSSEG8E8V` | |
| `VSSUBUVV` | |
| `VSSUBUVX` | |
| `VSSUBVV` | |
| `VSSUBVX` | |
| `VSUBVV` | |
| `VSUBVX` | |
| `VSUXEI16V` | |
| `VSUXEI32V` | |
| `VSUXEI64V` | |
| `VSUXEI8V` | |
| `VSUXSEG2EI16V` | |
| `VSUXSEG2EI32V` | |
| `VSUXSEG2EI64V` | |
| `VSUXSEG2EI8V` | |
| `VSUXSEG3EI16V` | |
| `VSUXSEG3EI32V` | |
| `VSUXSEG3EI64V` | |
| `VSUXSEG3EI8V` | |
| `VSUXSEG4EI16V` | |
| `VSUXSEG4EI32V` | |
| `VSUXSEG4EI64V` | |
| `VSUXSEG4EI8V` | |
| `VSUXSEG5EI16V` | |
| `VSUXSEG5EI32V` | |
| `VSUXSEG5EI64V` | |
| `VSUXSEG5EI8V` | |
| `VSUXSEG6EI16V` | |
| `VSUXSEG6EI32V` | |
| `VSUXSEG6EI64V` | |
| `VSUXSEG6EI8V` | |
| `VSUXSEG7EI16V` | |
| `VSUXSEG7EI32V` | |
| `VSUXSEG7EI64V` | |
| `VSUXSEG7EI8V` | |
| `VSUXSEG8EI16V` | |
| `VSUXSEG8EI32V` | |
| `VSUXSEG8EI64V` | |
| `VSUXSEG8EI8V` | |
| `VWADDUVV` | |
| `VWADDUVX` | |
| `VWADDUWV` | |
| `VWADDUWX` | |
| `VWADDVV` | |
| `VWADDVX` | |
| `VWADDWV` | |
| `VWADDWX` | |
| `VWCVTUXXV` | |
| `VWCVTXXV` | |
| `VWMACCSUVV` | |
| `VWMACCSUVX` | |
| `VWMACCUSVX` | |
| `VWMACCUVV` | |
| `VWMACCUVX` | |
| `VWMACCVV` | |
| `VWMACCVX` | |
| `VWMULSUVV` | |
| `VWMULSUVX` | |
| `VWMULUVV` | |
| `VWMULUVX` | |
| `VWMULVV` | |
| `VWMULVX` | |
| `VWREDSUMUVS` | |
| `VWREDSUMVS` | |
| `VWSUBUVV` | |
| `VWSUBUVX` | |
| `VWSUBUWV` | |
| `VWSUBUWX` | |
| `VWSUBVV` | |
| `VWSUBVX` | |
| `VWSUBWV` | |
| `VWSUBWX` | |
| `VXORVI` | |
| `VXORVV` | |
| `VXORVX` | |
| `VZEXTVF2` | |
| `VZEXTVF4` | |
| `VZEXTVF8` | |
| `WFI` | |
| `WORD` | |
| `XNOR` | |
| `XOR` | Bitwise XOR |
| `XORI` | XOR immediate |
| `ZEXTH` | |
Recognised: 975 mnemonics.
+140
View File
@@ -0,0 +1,140 @@
# Language: lexicon, statements and expressions
Layer 1, the common language, the same on every target. Verified against
`go tool asm` of Go 1.27.1 and against gasm's parser, which is differentially
tested against the toolchain. The authoritative sources behind this page are
the assembler's lexer (`cmd/asm/internal/lex`), its parser
(`cmd/asm/internal/asm/parse.go`) and the toolchain's own test data.
## Source files and targets
An assembly source is a `.s` file. The Go build convention names a
target-specific file with the architecture suffix, `_amd64.s`, `_arm64.s`,
`_riscv64.s` or `_loong64.s`; files without a suffix are portable across
targets. The same assembler program assembles every target: `go tool asm`
picks the target from the `GOOS` and `GOARCH` environment variables, and gasm
from the file name suffix or the `--arch` flag.
## Character set and identifiers
Sources are ASCII text. An identifier is a sequence of ASCII letters, digits
and underscores, digits never first, with exactly two additions:
- U+00B7, the middle dot `·`, stands for the period in a symbol's
package-qualified name;
- U+2215, the division slash `∕`, stands for the slash in a package path.
The two substitutions exist because the parser treats a real period and a
real slash as punctuation. The syntax is otherwise uppercase throughout:
instructions, registers and directives are written in upper case. The one
inherited exception is the `g` register name on 32-bit ARM.
## Comments
Two comment forms, both Go's:
```text
// a line comment
/* a block comment */
```
A comment of the form `//go:build` or the legacy `+build` comment is not a
plain comment: the lexer reports it to the build system as a build
constraint.
## Statements
The grammar of one line, from the parser:
```text
{label:} WORD[.qualifier] [ arg {, arg} ] (';' | '\n')
```
- A **label** is an identifier followed by a colon. Labels are
function-local: two functions in one file may reuse the same name, and a
reference resolves within the function that contains it. A branch
instruction names its target with a bare label operand, and the assembler
resolves it PC-relative. The explicit forms `offset(PC)`, a constant
counting instructions from the branch, and `name(SB)`, a cross-function
static reference, appear as branch targets as well.
- **WORD** is the instruction or directive name, upper case. On the ARM
family the word may carry a dot qualifier selecting a condition or shift
mode, such as the condition suffixes on 32-bit ARM; the amd64, arm64,
riscv64 and loong64 assemblies carry no instruction qualifiers apart from
their own width suffixes, which are part of the mnemonic.
- **Arguments** are separated by commas, with no trailing comma.
- A statement ends at a newline or at a semicolon, so several statements fit
on one line separated by `;`. Blank lines are free.
The first word of a line is a directive if it is one of the directive names
(TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN) and an instruction otherwise.
Unknown instruction names are errors; the instruction set is the set the
toolchain itself defines per target, plus the common pseudo-instructions.
## Literals
| Form | Examples | Notes |
|---|---|---|
| Integer | `0`, `42`, `0x2a`, `0o52`, `0b101010`, `1_000` | decimal, hexadecimal, octal and binary forms with Go's digit separators |
| Character | `'a'`, `'\n'`, `'\x41'` | single quoted, Go escape rules |
| String | `"this program can only run\n"` | double quoted, Go escape rules; accepted where an operand takes raw bytes, in practice a DATA initialiser |
| Float | `1.5`, `1e9` | accepted by the lexer; only meaningful where the target's encoding takes a float operand |
## Expressions
Constant expressions may appear wherever a constant is expected: in
immediates after `$`, in memory offsets, in frame and data sizes. The
evaluator works on unsigned 64-bit values with Go's operator precedence, and
the parser states its grammar in exactly those terms:
```text
expr = term { '+' term | '-' term | '|' term | '^' term }
term = factor { '*' factor | '/' factor | '%' factor | '<<' factor | '>>' factor | '&' factor }
factor = const | '+' factor | '-' factor | '~' factor | '(' expr ')'
```
Two consequences are worth naming, because the arithmetic surprises people
who read it as C:
- Shifts bind at the multiplicative level, next to `*` and `&`, while `|`
and `^` bind at the additive level. `$x<<1|3` computes `(x<<1)|3`, which
differs from `x*2+3` whenever `x` is odd. Plan 9 arithmetic is Go
precedence applied to a byte-oriented language, not the C expression it
resembles.
- The evaluator is unsigned and guarded: division or modulo by zero is an
error, and so is dividing a value with the high bit set; shift counts must
be non-negative; and a right shift of a value with the high bit set is
rejected rather than sign-extended.
An address expression such as `(index*4)(base)` is evaluated at assembly
time only if every name in it is a constant; a name that resolves to a
symbol turns the expression into a relocation request, never into a folded
constant.
Named constants enter expressions through the preprocessor (`#define`,
`-D`) and, in Go-embedded packages, through the generated `go_asm.h`; see
PREPROCESSOR.md and RUNTIME.md.
## The common pseudo-instructions
A handful of instructions exist on every target, assembled by the assembler
itself rather than the encoder: `NOP`, which emits the target's no-operation
encoding, and the frame-management pseudo-instructions the compiler emits
(`FUNCDATA`, `PCDATA`) which DIRECTIVES.md specifies. Everything else is the
target's own instruction set, and the assembler knows only the instructions
the toolchain's compiler emits; a hand-written kernel wanting more lays the
encoding down with `BYTE` on amd64 or waits for the extended layer.
## Case study: three lines, decomposed
```text
B.EQ 1(PC) // arm64: condition qualifier on the mnemonic,
// target one instruction past the branch
JMP done // every target: bare label, function-local,
// resolved PC-relative
MOVQ $reader__size>>3, CX // amd64: expression over a go_asm.h constant
```
The first shows a qualifier and the explicit relative target form; the second
the ordinary label reference; the third an expression over a generated
constant. Labels are reusable between functions without conflict.
+94
View File
@@ -0,0 +1,94 @@
# LoongArch 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md).
## Registers
- General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors
`V0` to `V31` and LASX vectors `X0` to `X31`.
- Fixed roles from the toolchain's table: `R0` is the constant zero, `R1`
the return address, `R3` the stack pointer, `R22` the goroutine pointer,
`R29` the closure context and `R30` the assembler's temporary. `R12`,
`R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences:
usable in assembly, but saved before any call.
## Widths ride the mnemonic
| Suffix | Width |
|---|---|
| `B`, `BU` | 8-bit, 8-bit unsigned |
| `H`, `HU` | 16-bit, 16-bit unsigned |
| `W`, `WU` | 32-bit, 32-bit unsigned |
| `V` | 64-bit |
| `F`, `D` | 32-bit and 64-bit float |
| `V` prefix (LSX) | 128-bit vector |
| `XV` prefix (LASX) | 256-bit vector |
The MOV series is the load and store interface: `MOVB (R2), R3` loads a
byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and
`XVMOVQ (R2), X1` a 256-bit one.
## Operand order
Most instructions appear in left-to-right assignment order: `ADDV R11, R12,
R13` is `add.d R13, R12, R11`, and the two-operand form
`OR R5, R6` assigns into R6. Exceptions:
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
## Addressing
- Plain: `offset(Rbase)`.
- Base plus offset **register**, no scale: `(R4)(R5)`, as in
`MOVB (R4)(R5), R6`, the `ldx` family.
- The pointer loads and stores `MOVWP` and `MOVVP` take a source-level
16-bit offset that the encoder halves into the 14-bit field, writing
`MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`.
## Vector element syntax
The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves
with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a
general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one,
`VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index],
Vd.B16` replicates one lane. The broadcast-from-memory form takes the true
byte offset at source level, which the encoder rescales per arrangement.
The permute and extract families take their 8-bit control word first:
`VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`.
## Alignment
`PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048,
and this target additionally auto-aligns loop heads to 16 bytes.
## Atomics, barriers and prefetch
- The `AM` atomic family comes in plain and `_DB` flavours; the `_DB`
forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a
full data barrier. Within the AM family the destination and base
registers may not coincide and the destination may not equal the operand
register: one is an exception, the other silently unspecified.
- `DBAR` carries the graded hint encoding documented for LA664 and later,
with hint 0x700 as the read-after-read lightweight barrier; older cores
treat every hint as the full barrier.
- `PRELD offset(Rbase), $hint` prefetches with the documented hints (0
load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded
block descriptor.
- `ALSL`-family shift-and-add writes the desired shift amount in source and
encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4.
- `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the
pointer loads for GOT relative access.
## Relocations
`R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the
PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`,
`GOT` and `GOT64` high and low pairs, the aligned conditional jump forms
`R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64`
for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).
+114
View File
@@ -0,0 +1,114 @@
# Operands: grammar, pseudo-registers, addressing and symbols
Layer 1, the common language. Verified against `go tool asm` of Go 1.27.1 and
against gasm's parser. The operand grammar is the part of the language that
varies most between targets, so this page fixes the common grammar and the
pseudo-registers; the per architecture pages carry the register names and the
addressing quirks each target adds.
## The four operand kinds
Every operand is one of four kinds:
```text
R1 register
$4 immediate
label branch target or symbol
-8(BX)(DI*4) memory
```
**Operands go source first, destination last**: `MOVQ x+0(FP), AX` loads the
argument into AX. This is the opposite of Intel order and the same order as
AT&T, with the sigils removed: registers are bare names, immediates take
`$`, memory is `offset(base)`.
## Registers
A register operand is its bare name, with no prefix: `AX`, `X15`, `R14` on
amd64; `R0` to `R30`, `ZR`, `V0` to `V31` on arm64; `X0` to `X31`, `F0` to
`F31`, `V0` on riscv64; `R0` to `R31`, `F0` to `F31`, `V0` on loong64.
Sub-register and width selection rides the mnemonic, not the operand: the
amd64 family spells `MOVB`, `MOVW`, `MOVL`, `MOVQ`, and the arm64 family
suffices `B`, `H`, `S`, `D`, `Q` on the shared forms. Each architecture page
lists its registers and the reserved ones.
## Immediates
`$` introduces a constant: `$42`, `$-1`, `$0x2a`, `$'A'`, `$bufSize`. The
`$` applies to the whole constant expression that follows, so
`$(4*8+reader__size)` is one immediate. Without the `$`, a number in operand
position is an address, not a value; the classic error `ADDQ 1, AX` asks the
assembler for the byte at address 1.
The one place a `$` number is not an immediate is the frame and argument
size field of TEXT, `$16-24`, which is two separate constants and not a
subtraction; DIRECTIVES.md specifies it.
## Memory
```text
offset(base)
offset(base)(index*scale)
```
Both parts are optional where the target allows them: `(BX)` is the memory
at BX, `foo+16(SB)` is a global, and on amd64 `foo+32(SP)(R9*8)` adds a
scaled index. `offset` is a constant expression, optionally carrying a
symbol name. The extensions beyond `offset(base)` are where the targets
diverge, and each belongs to its architecture page: amd64 carries the
`index*scale` form with scale 1, 2, 4 or 8 and its own rules on which
registers may index; loong64 writes base plus index as `(R4)(R5)`; the ARM
family attaches shift amounts to the index register in its own spelling.
The address arithmetic is on **byte addresses**: the offset is added to the
base as it stands, whatever the operand width of the instruction. Loading
the third 8-byte word of an array at BX is `16(BX)`, not `2(BX)`.
## The four pseudo-registers
Four names denote locations no target register holds, and they mean the same
on every architecture:
- **FP**, the frame pointer: the arguments and results of the current
function, at positive offsets, in the order the Go prototype declares
them. Every FP reference must carry a name: `x+0(FP)`, and an unnamed
`0(FP)` is rejected. Results follow arguments; an unnamed result is called
`ret`.
- **SP**, the virtual stack pointer: the high end of the function's local
frame, so locals live at negative offsets, `x-8(SP)`. A reference without
a name and without a plus, `-8(SP)`, addresses the **hardware** stack
pointer instead: the two spellings are one character apart and mean
different registers. That is the sharpest edge in the language and the
source of the deepest bugs.
- **SB**, the static base: the origin of memory, used for globals and
cross-package symbols, always with a name: `foo(SB)`, `foo+4(SB)`.
- **PC**, the program counter: branch targets, and the explicit relative
form `1(PC)`.
## Symbol names
A symbol's full name is the package path, a period, and the base name. In
source, the period is written U+00B7 (`·`) and a slash in the path U+2215
(`∕`), because the parser treats the ASCII forms as punctuation. Inside the
package's own file, `·Name` is enough and is the preferred spelling, since
it survives a rename of the import path.
| Spelling | Meaning |
|---|---|
| `·Name(SB)` | this package's Name |
| `runtime·morestack(SB)` | another package's morestack |
| `sourcedock.dev∕petrbalvin∕pkg·Name(SB)` | fully qualified |
| `msg<>(SB)` | file-local, the static of this language; `<>` also makes the ABI field static in the object |
| `Name<ABIInternal>(SB)` | ABI-qualified reference, the ABI in angle brackets after the name |
The object file these symbols produce, with the index rules that decide what
is referenced by name and what by index, is specified in
[GOOBJ.md](../GOOBJ.md).
## What vet adds in Go
Inside a Go package, `go vet`'s asmdecl analyzer checks every FP offset and
name against the Go prototype, and checks the declared argument area against
the frame. That layer, the prototype requirement and `go_asm.h`, belongs to
RUNTIME.md; the grammar above is the whole of what the assembler itself
requires.
+79
View File
@@ -0,0 +1,79 @@
# Preprocessing: include, define and selection
Layer 1, the common language. Verified against the preprocessor inside
`go tool asm` of Go 1.27.1 (`cmd/asm/internal/lex`), whose directives are
`#define`, `#undef`, `#include`, `#ifdef`, `#ifndef`, `#else`, `#endif` and
`#line`, and against gasm's implementation, which is differentially tested
against the toolchain's.
Input runs through a simplified C preprocessor before the parser sees it.
The set is deliberately small: there is no `#if` with constant expressions
and no token pasting with `##`. `#line` is honoured, so it changes the
positions the assembler reports and records.
## #include
```text
#include "textflag.h"
#include "go_asm.h"
#include "defs_linux_amd64.h"
```
The search path, in order: the directory of the including file, then the
directories given by repeatable `-I` flags. The assembler seeds no default
of its own: a bare `go tool asm` invocation finds none of the standard
headers, and it is the `go` build system that passes `$GOROOT/pkg/include`
among the `-I` directories when it drives the build. That directory ships
`textflag.h`, `funcdata.h` and the per architecture register headers.
Includes nest; a file included twice through different paths is processed
twice, which is why headers guard their defines.
## #define and #undef
```text
#define bufSize 1024
#define MOVD(d, s) MOVQ s, d
#undef bufSize
```
- An object macro replaces its name with its token sequence at the point of
use.
- A parameterised macro takes its arguments in parentheses and substitutes
them into the body. Macro parameters compose with the rest of the
language: an argument used with an element suffix, as in `A.S4` on the
vector forms, substitutes correctly.
- Redefinition is an error; `#undef` first, or pick a new name.
- The `-D name[=value]` flag predefines an object macro from the command
line, repeatable, exactly as `#define` would; a `-D` without a value
defines the name as `1`.
- Expansion happens when the name is used, so a macro may expand to
instructions, operands or fragments of either, and a macro body may use
macros defined before it.
`textflag.h` and `funcdata.h` are themselves ordinary `#define` files: the
flag names and the runtime macros are preprocessor definitions, not language
keywords. That is why a missing include produces a parser error at the first
use of `NOSPLIT` rather than a complaint about the name.
## #ifdef, #ifndef, #else, #endif
```text
#ifdef GOOS_windows
#define SYSCALL_INT 0x2b
#endif
```
Selection is by defined-name only: `#ifdef`, `#ifndef`, `#else`, `#endif`,
nesting freely. There is no `#if defined(x) && y`, because the preprocessor
evaluates no expressions; reach that with a build-tag Go file generating a
header, which is exactly how the runtime's own `go_asm.h` and defs headers
are produced.
## What preprocessing does not cover
The preprocessor is textual and runs first, so it knows nothing of assembly
semantics: it does not check that a macro expansion is a legal instruction,
and it does not participate in the constant expression evaluator, which runs
later, in the parser. A constant folded with `#define` and a constant folded
in an operand expression end at the same value through different doors;
GOOBJ.md records both in the object identically.
+56
View File
@@ -0,0 +1,56 @@
# The Plan 9 assembly language
This directory is the reference for the Plan 9 assembly language as the Go
toolchain and gasm accept it, written to be complete enough to implement
against. It exists because no such reference exists upstream: Go documents
the language on a single page, and the rest of the knowledge lives in the
toolchain's source and in the practice of reading it.
Every page carries the same conformance statement: which layer of the system
it describes, which toolchain release it was verified against, and how the
claims were checked. Pages in this directory are verified against Go 1.27.1
and against gasm's own differential test suite, which compares gasm's
behaviour with `go tool asm` byte for byte and output for output.
## The three layers
The reference deliberately separates three layers, because their rules have
different owners and different lifetimes:
1. **The common language** (LANGUAGE, OPERANDS, DIRECTIVES,
PREPROCESSOR): the syntax, operands, directives and preprocessing, the
same on every target and meaningful without a Go runtime.
2. **The Go-embedded layer** (RUNTIME): everything that exists only because
the code runs inside a Go program: the ABI0 contract, generated wrappers,
`go_asm.h`, the garbage collector annotations and `go vet` checks.
3. **The standalone layer** (STANDALONE, planned with the standalone
compilation phase): using the language outside Go, through gasm's ELF
output and the extended instruction set, where the toolchain offers no
ground truth and execution testing is the only verification.
A rule stated in layer 1 holds on every target. A rule stated in layer 2
says which part of the Go machinery imposes it. Nothing in layer 3 changes
layers 1 or 2; it extends them.
## Pages
| Page | Layer | Contents |
|---|---|---|
| [LANGUAGE.md](LANGUAGE.md) | 1 | lexicon, statement structure, labels, literals, expressions |
| [OPERANDS.md](OPERANDS.md) | 1 | operand grammar, pseudo-registers, addressing modes, symbol naming |
| [DIRECTIVES.md](DIRECTIVES.md) | 1 | TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN and the function flags |
| [PREPROCESSOR.md](PREPROCESSOR.md) | 1 | `#include`, `#define`, `#ifdef` and friends, `-D`, `-I` |
| [RUNTIME.md](RUNTIME.md) | 2 | ABI0, prototypes, `go_asm.h`, `funcdata.h`, `go vet` |
| [AMD64.md](AMD64.md) | 1 | registers, addressing, the frame and split check, families, relocations |
| [ARM64.md](ARM64.md) | 1 | registers, the MOV load and store series, special operand orders, SIMD |
| [RISCV64.md](RISCV64.md) | 1 | registers and their constrained names, per class operand order, profiles, vector extension |
| [LOONG64.md](LOONG64.md) | 1 | registers, width suffixes, vector element syntax, atomics and barriers |
| INSTRUCTIONS-AMD64.md and the other three | 1 | generated per architecture inventory of every accepted mnemonic |
| STANDALONE.md | 3 | the language outside Go |
## Status
The common-language core, the Go-embedded layer, all four per-architecture
pages and the generated instruction appendices are written and verified.
STANDALONE.md lands with the standalone compilation phase. The object format
these pages feed is specified in [GOOBJ.md](../GOOBJ.md).
+104
View File
@@ -0,0 +1,104 @@
# RISC-V 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
## Registers
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
goroutine pointer, must be written `g` and may not be written `S11`; in
shared builds `X3` is off limits and must be written `GP`.
- The other integer registers may be written `Xn` or by their ABI names
(`A0`, `T0`, `S1`, and so on).
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
- `X26` is the closure pointer and `X31` is the assembler's own scratch
register: its value may be clobbered by instruction sequences the
assembler inserts, so hand-written code must not rely on it.
- There is no reserved frame pointer register on this target.
## Operand order
The ordering differs from the ISA manual, and per instruction class:
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
- **I-type arithmetic** keeps that shape with the immediate first:
`ADDI $1, X11, X12`.
- **Loads and stores** are source first, like every Plan 9 dialect:
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
the width; `MOVB` through `MOVD` spell it out.
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
X12 < X23, the reverse of the SLT operand order.
- **FMA** is rotated one place left so the destination comes last:
`FMADDS F1, F2, F3, F4`.
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
`ADD X10, X12, X12`.
Where an R-type instruction has an I-type sibling, the assembler picks the
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
## Names, suffixes and rounding
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
the FCSR.
## Constants
- `MOV` materialises any 64-bit integer constant, synthesising it from a
few arithmetic instructions where possible and otherwise loading it from
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
## Extensions and profiles
The default target profile is rva20u64, selected or raised with the
GORISCV64 environment variable. A short list of instructions outside the
default profile is synthesised by the assembler when the profile does not
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
## Fences and atomics
`FENCE` takes predecessor and successor sets in that order, uppercase
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
and the AMO instructions are not specifiable in source: the assembler sets
acquire and release on the AMO instructions, acquire on `LR` and release on
`SC`, always.
## Compressed instructions
The assembler converts 32-bit instructions to their compressed encodings
automatically; the conversion is a property of the emitted machine code, not
of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
source first like the scalar ones, with an optional stride or index register
second and the mask register, when present, always penultimate:
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
penultimate.
## Relocations
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
families and `R_RISCV_ADD32` and `SUB32`, all specified in
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
`R_DWTXTADDR_U4` flavour inside its DWARF records.
+120
View File
@@ -0,0 +1,120 @@
# The Go-embedded layer: ABI0, prototypes and the runtime contract
Layer 2: everything that exists only because the assembly runs inside a Go
program. Without a Go runtime this page does not apply; the language of
OPERANDS.md and DIRECTIVES.md still does. Verified against Go 1.27.1, against
the shipped `funcdata.h` header, and against the object files the toolchain
produces, which were parsed and checked field by field while writing
[GOOBJ.md](../GOOBJ.md).
## Hand-written assembly is ABI0
Go functions compiled from source use ABIInternal, the register-based
calling convention, which the toolchain documents as unstable and free to
change between releases. A `.s` function is written against ABI0, the stack
based convention: arguments and results live in the caller's frame at
positive FP offsets, byte-addressed, in declaration order, with no registers
assigned at all. The toolchain generates the wrapper that translates between
the two; a caller in Go calling an assembly function goes through it, and it
is marked `ABIWRAPPER` in the object. Hand-writing a bridge is never needed
and never correct.
## Every assembly function carries a Go prototype
```go
package add
func Add(x, y int64) int64
```
The body-less declaration is not optional, and not only for the linker: it
is what tells the garbage collector which arguments and results hold
pointers, and what `go vet` checks the assembly against. Even a function
nothing in Go calls gets one. Consequences:
- The FP operand names and offsets are checked by vet's asmdecl analyzer
against the prototype: `x+0(FP)` must name an argument that exists, at the
offset the prototype says. A file that assembles and links can still fail
vet.
- The declared argument area in `$framesize-argsize` is checked against the
prototype's size. An omitted argsize marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`), which is the normal spelling for functions with no Go
callers.
- `//go:noescape` on the declaration tells the compiler that a pointer
argument does not escape, for assembly that keeps the pointer beyond the
call.
## The frame, the stack and the collector
The runtime owns the stack and the pointer map, and assembly must hold up
its end of four rules:
1. **Arguments are initialised on entry; results are not.** A function whose
results hold live pointers across a call must zero them and then execute
`GO_RESULTS_INITIALIZED`. Designing functions that return no pointers
avoids the problem.
2. **A frame with calls and no local pointers says so** with
`NO_LOCAL_POINTERS`. A frame with local pointers that the runtime cannot
see is not allowed at all: assembly cannot describe a pointer-containing
local, so it must not have one. Data symbols containing pointers are the
same: define them in Go.
3. **The stack may move.** Stack growth copies the frame, so no pointer into
the frame may be held across a call, and the raw hardware SP register may
not be cached across a call either.
4. **The split check is not optional by default.** Without NOSPLIT, the
assembler inserts the stack-growth preamble, including the morestack
block for framed functions; NOSPLIT is a contract that the frame and
everything below it fit in the remaining stack segment. On amd64 the
assembler also marks small leaf functions NoSplit itself and skips the
preamble, so silence is not a promise.
The simplest safe shape is a leaf function with no local frame and no calls:
it needs no annotation beyond the prototype.
## go_asm.h: Go constants and layout in assembly
A package with `.s` files gets a generated header. Include it and use the
generated names instead of hard-coding layouts, which lie silently when the
Go side changes:
| Go declaration | Assembly name |
|---|---|
| `const bufSize = 1024` | `const_bufSize` |
| field `r` of `type reader struct` | `reader_r` |
| size of `type reader struct` | `reader__size` |
The constants arrive as macros, usable as immediates and offsets, computed
from the Go declarations. An ambiguous name, such as a struct that really
has a `_size` field, fails the generation with a redefinition error.
## funcdata.h: the runtime macros
`$GOROOT/pkg/include/funcdata.h` defines the PCDATA and FUNCDATA ids and the
three macros assembly normally uses instead:
| Macro | Expands to | Meaning |
|---|---|---|
| `GO_ARGS` | `FUNCDATA $FUNCDATA_ArgsPointerMaps, go_args_stackmap(SB)` | the Go prototype defines the argument pointer map |
| `GO_RESULTS_INITIALIZED` | `PCDATA $PCDATA_StackMapIndex, $1` | results are initialised; treat them as live from here |
| `NO_LOCAL_POINTERS` | `FUNCDATA $FUNCDATA_LocalsPointerMaps, no_pointers_stackmap(SB)` | the frame holds no pointers |
`GO_ARGS` is inserted implicitly by the assembler for any function whose
package-qualified name belongs to the current package, which is why most
assembly never writes it. `NOSPLIT` leaf functions that call nothing need
none of the three.
The underlying ids, for reading toolchain output rather than for writing
source: FUNCDATA 0 to 7 are args pointer maps, locals pointer maps, stack
objects, inline tree, open-coded defer info, argument info, argument
liveness and wrap info; PCDATA 0 to 4 are unsafe point, stack map index,
inline tree index, argument liveness index and panic bounds.
## What the runtime does with all of this
The object file records the annotations as aux symbols and FuncInfo records;
GOOBJ.md specifies the encoding. The linker assembles them into the runtime's
pclntable, which traceback and the collector consume. An assembly function
that misdeclares its frame is not a compile error and usually not a link
error: it is a wrong collector decision or a wrong traceback at runtime,
which is why the annotations are a contract and not documentation.
+18 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-asm \- assemble Plan 9 assembly without the Go toolchain
.SH SYNOPSIS
.B gasm asm [\-\-format raw|elf|goobj] [\-p pkg] [\-GOARCH arch] [\-o out] <file>
.B gasm asm [\-\-format raw|elf|goobj] [\-I dir] [\-p pkg] [\-GOARCH arch] [\-GOOS os] [\-o out] <file>
.SH DESCRIPTION
Assemble FILE without the Go toolchain: every TEXT function is encoded
to machine code and printed as a hex dump. Supported architectures:
@@ -42,11 +42,24 @@ need no toolchain at all.
Framed functions receive the stack-split guard and the trailing
morestack block, byte-identical to the toolchain's output, so split
functions link too.
.PP
A file that includes go_asm.h gets that header generated from the Go
files beside it, type-checked for the target.
.B \-GOOS
selects the type-checking GOOS for that header, because a GOOS-specific
file needs its platform's defines: sys_darwin_arm64.s fails against the
ambient GOOS (machTimebaseInfo_numer is missing from a linux type-check)
and assembles with
.BR "\-GOOS darwin" .
.SH OPTIONS
.TP
.B \-\-format \fIraw|elf|goobj\fR
Output format; the default is raw.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory.
.TP
.B \-p \fIpkg\fR
Package path for --format goobj, qualifying the exported symbols.
.TP
@@ -55,6 +68,10 @@ Target architecture: amd64, arm64, riscv64 or loong64; overrides the
file-name suffix, which is how the suffix-less majority of GOROOT's
files (cpu_x86.s, stub.s, ...) become assemblable.
.TP
.B \-GOOS \fIos\fR
Operating system for the generated go_asm.h: any GOOS go/build
recognises in file names; the default is the host's.
.TP
.B \-o \fIfile\fR
Write the output to this file instead of a hex dump on stdout.
.SH EXIT STATUS
+14 -2
View File
@@ -1,8 +1,8 @@
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-19" "gasm" "User Commands"
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-21" "gasm" "User Commands"
.SH NAME
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
.SH SYNOPSIS
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [amd64|arm64|riscv64|loong64]
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-\-list] [\-I dir] [amd64|arm64|riscv64|loong64]
.SH DESCRIPTION
Compare the gasm encoder for the given architecture (default amd64)
against
@@ -38,6 +38,18 @@ second.
.B \-\-corpus [\fIdir\fR]
Assemble a corpus of .s files and report pass rates and failure
reasons.
.TP
.B \-\-list
With
.BR \-\-corpus ,
print every failing file with its failure reason, per architecture,
instead of one representative file per reason.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory. A corpus run whose files include
toolchain headers (such as GOROOT/pkg/include) needs it, the same -I a
toolchain comparison takes.
.SH EXIT STATUS
The mnemonic-diff mode reports through its output and exits 0; a failed
probe or an unknown architecture exits non-zero.
+5 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-diff \- compare the machine code of two assembly files
.SH SYNOPSIS
.B gasm diff [\-GOARCH arch] <file1.s> <file2.s>
.B gasm diff [\-GOARCH arch] [\-I dir] <file1.s> <file2.s>
.SH DESCRIPTION
Compare the machine code produced by assembling two files. Shows which
functions differ and the byte-level differences. Useful for verifying
@@ -20,6 +20,10 @@ pairs two variants regardless of suffix.
Target architecture for both files: amd64, arm64, riscv64 or loong64;
overrides the file-name suffixes.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory.
.TP
.B \-\-map \fIspec\fR
Comma-separated old=new pairs to match functions with different names.
.SH EXIT STATUS
+23 -6
View File
@@ -241,6 +241,13 @@ func renderInstr(line []token.Token, width int) string {
if line[0].Kind != token.Ident {
return "\t" + mnem + " " + ops
}
// A statement separator belongs to the statement it ends: when the
// operands open with a ';', the alignment padding would land between
// the mnemonic and its own separator (REP ; MOVSQ), so such a line
// renders with a single space whatever the function's width.
if strings.HasPrefix(ops, ";") {
return "\t" + mnem + " " + ops
}
if width < len(mnem) {
width = len(mnem)
}
@@ -254,11 +261,12 @@ func renderPreproc(line []token.Token) string {
line[2].Kind == token.String {
return "#include " + line[2].Text
}
parts := make([]string, 0, len(line)-1)
for _, t := range line[1:] {
parts = append(parts, t.Text)
}
return "#" + strings.Join(parts, " ")
// The body of a directive, a macro definition included, is an ordinary
// token run: rendering it through renderOps applies the same punctuation
// rules as everywhere else, so a macro body keeps its canonical spelling
// ($v, (a, b), the ';' separators between statements) instead of being
// spread with a space between every token.
return "#" + renderOps(line[1:])
}
// renderOps re-spaces a run of operand tokens into canonical form. It never
@@ -348,6 +356,13 @@ func spaceBetween(prev, cur token.Token) bool {
return false
case token.Comma:
return false
case token.Semicolon:
// A ';' is a statement separator on the assembly path, not an
// operand: dropping it would fuse two statements into a line the
// assembler rejects, so it must survive as punctuation. It glues
// to the statement it ends and the next statement takes one space,
// matching the toolchain's listing style.
return false
case token.Star, token.Plus, token.Minus, token.Slash, token.Pipe:
return false
case token.LShift, token.RShift, token.Arrow, token.At:
@@ -376,7 +391,9 @@ func spaceBetween(prev, cur token.Token) bool {
return false
case token.LAngle, token.RAngle:
return false
case token.Comma:
case token.Comma, token.Semicolon:
// The statement after a ';' separator takes its own space, exactly
// like the operand after a comma.
return true
}
return true
+192
View File
@@ -5,9 +5,11 @@ package format
import (
"os"
"slices"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
@@ -298,6 +300,196 @@ func TestCRLFInputIsNormalisedToLF(t *testing.T) {
}
}
// TestSemicolonSeparators pins the treatment of ';' statement separators.
// The separator is load-bearing on the assembly path, where the parser reads
// semicolon-separated statements: a formatter that drops it fuses two
// statements into a line the assembler rejects, which is data corruption.
// Each row pins the canonical spelling, one space after the ';', tight
// before it, the way the toolchain's own sources and listings write it.
func TestSemicolonSeparators(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{
name: "between instructions, tight",
in: "TEXT ·f(SB), $0\nBYTE $0x48;BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "between instructions, spaced",
in: "TEXT ·f(SB), $0\nBYTE $0x48 ; BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "after a label",
in: "TEXT ·f(SB), $0\nlabel: BYTE $1; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\nlabel:\n\tBYTE $1; BYTE $2\n\tRET\n",
},
{
// The continuation-spliced macro shape of the runtime sources:
// the lexer makes one logical line of the backslash continuations.
name: "inside a macro body, continued",
in: "#define MOVLTOREG(v, off) \\\n\tMOVL $v, AX; \\\n\tMOVL AX, ret+off(FP)\n",
want: "#define MOVLTOREG(v, off) MOVL $v, AX; MOVL AX, ret+off(FP)\n",
},
{
name: "inside a macro body, one line",
in: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
want: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
},
{
name: "several separators in one line",
in: "TEXT ·f(SB), $0\nBYTE $1; BYTE $2; BYTE $3\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1; BYTE $2; BYTE $3\n\tRET\n",
},
{
name: "two separators back to back",
in: "TEXT ·f(SB), $0\nBYTE $1;; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;; BYTE $2\n\tRET\n",
},
{
name: "inside a line comment, untouched",
in: "TEXT ·f(SB), $0\n// keep; the; separators\nBYTE $1\nRET\n",
want: "TEXT ·f(SB), $0\n\t// keep; the; separators\n\tBYTE $1\n\tRET\n",
},
{
name: "after a statement, before a comment",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX; // tail\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX; // tail\n\tRET\n",
},
{
name: "last character on a line",
in: "TEXT ·f(SB), $0\nBYTE $1;\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;\n\tRET\n",
},
{
// The REP shape: a prefix-style zero-operand statement
// followed by the instruction it prefixes. The separator
// belongs to the statement it ends, so the function's
// alignment width (MOVSQ is the widest mnemonic here) must
// not open a gap before it: one space after the mnemonic
// whatever the neighbours' lengths.
name: "after a prefix-style statement",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX\nREP; MOVSQ\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX\n\tREP ; MOVSQ\n\tRET\n",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := Source(tc.in)
if got != tc.want {
t.Fatalf("formatting mismatch:\n--- got ---\n%q\n--- want ---\n%q", got, tc.want)
}
if again := Source(got); again != got {
t.Fatalf("not idempotent:\n%q", again)
}
if in, out := strings.Count(tc.in, ";"), strings.Count(got, ";"); in != out {
t.Fatalf("semicolon count changed: %d -> %d\n%s", in, out, got)
}
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
t.Fatalf("formatted output no longer parses: %v", errs)
}
})
}
}
// TestSemicolonStatementRoundTrip proves the formatter's contract on the
// path where ';' separates statements: parse the source the way the
// assembler does, format it, re-parse the formatted text and compare the
// statement sequence. Raw operand texts are token-joined, so they are
// insensitive to the whitespace a format pass chooses, and the comparison
// can only fail when a token is lost: dropping a ';' fuses two statements
// into one, exactly the corruption the released formatter committed.
func TestSemicolonStatementRoundTrip(t *testing.T) {
src := "#define MOVLTOREG(v, off) \\\n" +
"\tMOVL $v, AX; \\\n" +
"\tMOVL AX, ret+off(FP)\n" +
"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"BYTE $0x48; BYTE $0xc7\n" +
"first: BYTE $1; BYTE $2\n" +
"MOVLTOREG($42, 0)\n" +
"RET\n"
before, errs := parser.ParseWithOptions("in.s", src, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("source does not parse: %v", errs)
}
formatted := Source(src)
after, errs := parser.ParseWithOptions("in.s", formatted, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("formatted source does not parse: %v", errs)
}
want, got := stmtSignature(before), stmtSignature(after)
if !slices.Equal(got, want) {
t.Fatalf("statement sequence changed:\n--- before ---\n%q\n--- after ---\n%q", want, got)
}
if again := Source(formatted); again != formatted {
t.Fatalf("not idempotent:\n%q", again)
}
// The two BYTE statements on the first line must stay two: one fused
// statement here is the exact defect this package once shipped.
var bytes []string
for _, stmt := range stmtSignature(after) {
if rest, ok := strings.CutPrefix(stmt, "instr BYTE "); ok {
bytes = append(bytes, rest)
}
}
if want := []string{"$ 0x48", "$ 0xc7", "$ 1", "$ 2"}; !slices.Equal(bytes, want) {
t.Fatalf("BYTE statements after expansion = %q, want %q", bytes, want)
}
}
// stmtSignature flattens a parsed file into one string per declaration and
// statement, in source order. Every component is token-derived, so the
// signature is stable across format passes and moves only when a token is
// lost or gained.
func stmtSignature(f *ast.File) []string {
var out []string
for _, d := range f.Decls {
switch d := d.(type) {
case *ast.Text:
out = append(out, "text "+d.Name.Raw)
for _, s := range d.Body {
out = append(out, stmtText(s))
}
case *ast.Globl:
out = append(out, "globl "+d.Name.Raw)
case *ast.Data:
out = append(out, "data "+d.Name.Raw)
case *ast.Include:
out = append(out, "include "+d.Header.Text)
case *ast.Preproc:
out = append(out, "preproc "+d.Raw)
}
}
for _, s := range f.Orphans {
out = append(out, stmtText(s))
}
return out
}
// stmtText renders one statement for stmtSignature.
func stmtText(s ast.Stmt) string {
switch s := s.(type) {
case *ast.Label:
return "label " + s.Name.Text
case *ast.Instr:
parts := make([]string, 0, len(s.Operands)+1)
parts = append(parts, s.Mnemonic.Text)
for _, op := range s.Operands {
parts = append(parts, op.Raw)
}
return "instr " + strings.Join(parts, " ")
default:
return "stmt"
}
}
// lexOperands lexes a single operand string and drops the EOF token.
func lexOperands(s string) []token.Token {
toks := lexer.Tokenize(s)
+6
View File
@@ -31,6 +31,12 @@ func FuzzFormatIdempotency(f *testing.F) {
f.Add("TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ AX, BX\n\tRET\n")
f.Add("TEXT ·f(SB),NOSPLIT,$0\n\tMOVQ AX,BX\n\n\n\tRET\n")
f.Add("garbage ### ???\n")
// Line-ending whitespace at the edge of a comment: a CR followed by more
// trailing whitespace once survived the first pass and disappeared on
// re-lexing, so formatting was not idempotent.
f.Add("//\r ")
f.Add("// loop \r\t\nMOVQ AX, BX\n")
f.Add("TEXT ·f(SB), NOSPLIT, $0 // tail\r\n\tMOVQ AX, BX\r\n\tRET\r\n")
f.Fuzz(func(t *testing.T, src string) {
once := Source(src)
@@ -0,0 +1,2 @@
go test fuzz v1
string("//\r ")
+62 -7
View File
@@ -16,8 +16,16 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// middleDot is the Plan 9 symbol separator (U+00B7), used in ·funcName(SB).
const middleDot = '\u00B7'
const (
// middleDot is the Plan 9 symbol separator (U+00B7), used in
// ·funcName(SB): it stands for the period between package path and name.
middleDot = '\u00B7'
// divisionSlash is the Plan 9 path separator (U+2215), used inside the
// package path of a symbol: internal∕runtime∕atomic·Xchg. Like the
// middle dot it is an identifier character, so a package path containing
// it lexes as one name; the ordinary slash (U+002F) stays punctuation.
divisionSlash = '\u2215'
)
// Lexer scans a source string one token at a time.
type Lexer struct {
@@ -119,7 +127,9 @@ func (l *Lexer) Next() token.Token {
// is a C-preprocessor line continuation (used by #define macros in the
// runtime .s files): splice the lines together by consuming both, so
// the whole macro becomes one logical line that the parser treats as an
// opaque preprocessor directive.
// opaque preprocessor directive. The backslash may also reach its
// newline across whitespace and a trailing comment ("…; \ // note\n"),
// which the toolchain's scanner skips the same way.
for {
c := l.cur()
if c == ' ' || c == '\t' || c == '\r' {
@@ -136,6 +146,16 @@ func (l *Lexer) Next() token.Token {
}
continue
}
if c == '\\' && l.continuationAhead() {
l.advance() // backslash, then the runes the scan saw
for !l.atEnd() && l.cur() != '\n' {
l.advance()
}
if !l.atEnd() {
l.advance() // the newline that closes the continuation
}
continue
}
break
}
@@ -185,16 +205,42 @@ func (l *Lexer) Next() token.Token {
}
}
// continuationAhead reports, without consuming anything, whether the
// backslash at the current position closes onto a newline through nothing
// but horizontal whitespace and one line comment. Positions after the
// backslash are inspected directly on the rune slice so a non-match leaves
// the scanner state untouched.
func (l *Lexer) continuationAhead() bool {
i := l.i + 1
for i < len(l.src) {
switch r := l.src[i]; {
case r == ' ' || r == '\t' || r == '\r':
i++
case r == '/' && i+1 < len(l.src) && l.src[i+1] == '/':
for i < len(l.src) && l.src[i] != '\n' {
i++
}
default:
return r == '\n'
}
}
return false
}
// lineComment consumes a // comment up to, but not including, the newline. A
// trailing \r is part of a CRLF line ending rather than comment content:
// dropping it keeps the formatter's output uniformly LF-terminated.
// trailing run of \r, spaces and tabs is line-ending whitespace rather than
// comment content, so it never enters the token text. Trimming only a \r
// directly before the token's end would make the text depend on what follows
// the comment (a newline or the end of the input): "//x\r " would carry the
// "\r " while "//x\r\n" would not, and a formatter that terminates the line
// with \n would then re-lex its own output to a shorter comment.
func (l *Lexer) lineComment(start token.Position) token.Token {
var b strings.Builder
for !l.atEnd() && l.cur() != '\n' {
b.WriteRune(l.cur())
l.advance()
}
return l.make(token.Comment, start, strings.TrimSuffix(b.String(), "\r"))
return l.make(token.Comment, start, strings.TrimRight(b.String(), " \t\r"))
}
// blockComment consumes a /* ... */ comment, tolerating an unterminated one.
@@ -399,6 +445,15 @@ func (l *Lexer) punct(start token.Position) token.Token {
case '|':
l.advance()
return l.make(token.Pipe, start, "|")
case ';':
l.advance()
return l.make(token.Semicolon, start, ";")
case '&':
l.advance()
return l.make(token.Ampersand, start, "&")
case '~':
l.advance()
return l.make(token.Tilde, start, "~")
default:
// Unknown rune: emit it as Illegal and move on.
l.advance()
@@ -413,7 +468,7 @@ func isHexDigit(r rune) bool {
}
func isIdentStart(r rune) bool {
return r == '_' || r == middleDot || unicode.IsLetter(r)
return r == '_' || r == middleDot || r == divisionSlash || unicode.IsLetter(r)
}
func isIdentChar(r rune) bool {
+26
View File
@@ -85,6 +85,19 @@ func TestLabelAndComment(t *testing.T) {
[]token.Kind{token.Ident, token.Colon, token.Ident, token.Ident, token.Comment})
}
func TestLineCommentTrailingWhitespace(t *testing.T) {
// A trailing run of CR, spaces and tabs is line-ending whitespace, not
// comment content. The token text must not depend on what follows the
// comment: before the trim covered only a CR directly before the token's
// end, "// loop\r " kept the CR while "// loop\r\n" dropped it, and the
// formatter re-lexed its own output to a shorter comment.
eq(t, texts("// loop\r"), []string{"// loop"})
eq(t, texts("// loop\r "), []string{"// loop"})
eq(t, texts("// loop \r\t\nMOVQ AX, BX"), []string{"// loop", "MOVQ", "AX", ",", "BX"})
// A CR inside the comment is content and stays.
eq(t, texts("// loops\rall"), []string{"// loops\rall"})
}
func TestAVX512Mnemonics(t *testing.T) {
eq(t, texts("VFMADD231PD Z14, Z12, Z10"),
[]string{"VFMADD231PD", "Z14", ",", "Z12", ",", "Z10"})
@@ -169,6 +182,19 @@ func TestNulIsIllegal(t *testing.T) {
eq(t, texts("MOVQ \x00 AX"), []string{"MOVQ", "\x00", "AX"})
}
func TestDivisionSlashInIdentifiers(t *testing.T) {
// U+2215 DIVISION SLASH is an identifier character, the way the
// toolchain's tokenizer treats it: the package path of a symbol is
// written with it (internal∕runtime∕atomic·Xchg) and must lex as one
// name. The ordinary slash (U+002F) stays punctuation.
eq(t, texts("CALL internal∕runtime∕atomic·Xchg(SB)"),
[]string{"CALL", "internal∕runtime∕atomic·Xchg", "(", "SB", ")"})
eq(t, texts("MOVQ sync∕atomic·Align(SB), AX"),
[]string{"MOVQ", "sync∕atomic·Align", "(", "SB", ")", ",", "AX"})
// It may also begin a name, like any letter of the toolchain's rule.
eq(t, kinds("∕x"), []token.Kind{token.Ident})
}
// TestOffsetsAroundInvalidByte pins Position.Offset against the original
// bytes: an invalid UTF-8 byte decodes to RuneError but advances the offset
// table by exactly one byte, so every later position stays a true byte
+6
View File
@@ -21,6 +21,12 @@ import (
// The check requires a parseable signature; functions without one, and
// functions whose parameters are all covered by frame reads, stay silent.
func checkABI0Args(t *ast.Text) []Diagnostic {
// An explicit <ABIInternal> TEXT reads its arguments from the register
// file by declaration (runtime·memmove<ABIInternal> is the canonical
// example), so the ABI0 frame contract does not apply to it.
if t.Name != nil && t.Name.ABI != "" {
return nil
}
params, ok := abiParamNames(t.Doc)
if !ok || len(params) == 0 {
return nil
+17
View File
@@ -82,6 +82,23 @@ func TestABIArgSizeSkipsRegisterABI(t *testing.T) {
}
}
// TestABI0ArgsSkipsABIInternal verifies the frame-read check does not fire for
// a TEXT declared <ABIInternal>: runtime·memmove<ABIInternal> and friends read
// their arguments from the register file by declaration, which is the correct
// spelling there, not the register-args port bug the rule hunts.
func TestABI0ArgsSkipsABIInternal(t *testing.T) {
diags := lintSrc(t, "#include \"textflag.h\"\n"+
"// func memmove(to, from unsafe.Pointer, n uintptr)\n"+
"TEXT ·memmove<ABIInternal>(SB), NOSPLIT, $0-24\n"+
"\tMOVQ AX, DI\n"+
"\tMOVQ BX, SI\n"+
"\tMOVQ CX, BX\n"+
"\tRET\n")
if codes(diags)[CodeABI0RegisterArgs] != 0 {
t.Fatalf("ABIInternal TEXT must not be checked against the FP frame: %+v", diags)
}
}
// TestUnreachableCode exercises the dead-code detection and its guard rails.
func TestUnreachableCode(t *testing.T) {
// Code after a RET is unreachable.
+111 -8
View File
@@ -338,10 +338,8 @@ func lintText(t *ast.Text, tab *arch.Table, archKnown bool, cfg Config, macros m
}
if isJump(cfg.Arch, upper) {
for _, op := range st.Operands {
if name, pos, ok := localLabelRef(op); ok && !tab.IsRegister(name) && !arch.IsPseudoReg(name) {
referenced[name] = pos
}
if name, pos, ok := branchTargetRef(cfg.Arch, upper, st.Operands, tab); ok {
referenced[name] = pos
}
}
}
@@ -649,17 +647,65 @@ func localLabelRef(op *ast.Operand) (string, token.Position, bool) {
return sym.Name, op.Pos, true
}
// branchTargetRef returns the local label a branch transfers control to: the
// bare symbol in the destination position, the last operand, since that is
// where the Plan 9 branch target sits. A register-named target is a
// register-indirect branch (JMP AX, arm64 BR R5, riscv64 JALR X6, loong64
// JIRL R1) and yields no reference, unless the encoder reads the target
// positionally (positionalBranchTarget): there a label may legitimately
// collide with a register alias, riscv64 ZERO being the ABI name of X0, and
// a label named zero is ordinary code.
func branchTargetRef(a arch.Arch, upper string, ops []*ast.Operand, tab *arch.Table) (string, token.Position, bool) {
if len(ops) == 0 {
return "", token.Position{}, false
}
name, pos, ok := localLabelRef(ops[len(ops)-1])
if !ok {
return "", token.Position{}, false
}
if !positionalBranchTarget(a, upper) && (tab.IsRegister(name) || arch.IsPseudoReg(name)) {
return "", token.Position{}, false
}
return name, pos, true
}
// positionalBranchTarget reports whether the encoder reads a bare-symbol
// operand of the branch as its label target from a fixed position, without
// consulting the register file. The riscv64 branch, JMP and JAL encoders do
// (labelFromOperand in asm/riscv_assemble.go), as do the loong64 branch,
// BFPT/BFPF and jump encoders (l64Label in asm/loong64_assemble.go). amd64
// never does, because a bare register operand to JMP/CALL/Jcc is a
// register-indirect branch; nor do the register-indirect forms of the RISC
// families (arm64 BR/BLR, riscv64 JALR/JR, loong64 JIRL).
func positionalBranchTarget(a arch.Arch, upper string) bool {
switch a {
case arch.RISCV:
return riscvBranches[upper] || upper == "JMP" || upper == "JAL"
case arch.LOONG64:
return loong64Branches[upper] || upper == "JMP" || upper == "B" ||
upper == "JAL" || upper == "BL"
}
return false
}
// riscvBranches and loong64Branches are the conditional-branch mnemonics; they
// are listed explicitly rather than matched by a "B" prefix so that bit-manip
// instructions (BCLR, BSET, …) are never mistaken for branches.
// instructions (BCLR, BSET, …) are never mistaken for branches. The sets
// mirror the encoder's own branch cases: the B-type table entries
// (riscv_encode.go), the branch-zero pseudos and the reversed branches
// BGT/BGTU/BLE/BLEU (riscv_assemble.go), and for loong64 the 16-bit branch
// table plus the single-register forms of l64branch21Table (BEQZ/BNEZ and the
// floating-point branches BFPT/BFPF).
var riscvBranches = map[string]bool{
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
"BEQZ": true, "BNEZ": true, "BLEZ": true, "BGEZ": true, "BLTZ": true, "BGTZ": true,
"BGT": true, "BGTU": true, "BLE": true, "BLEU": true,
}
var loong64Branches = map[string]bool{
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
"BLEZ": true, "BLTZ": true, "BGEZ": true, "BGTZ": true,
"BEQZ": true, "BNEZ": true, "BFPT": true, "BFPF": true,
}
// isJump reports whether the mnemonic is any branch.
@@ -676,7 +722,8 @@ func isJump(a arch.Arch, upper string) bool {
upper == "JR" || upper == "BR"
case arch.LOONG64:
return upper == "CALL" || loong64Branches[upper] ||
upper == "JIRL" || upper == "JMP" || upper == "BR"
upper == "JIRL" || upper == "JMP" || upper == "BR" ||
upper == "B" || upper == "JAL" || upper == "BL"
default: // amd64
return upper == "CALL" || strings.HasPrefix(upper, "J")
}
@@ -692,7 +739,8 @@ func isUnconditionalJump(a arch.Arch, upper string) bool {
return upper == "JMP" || upper == "J" || upper == "JAL" ||
upper == "JALR" || upper == "JR" || upper == "BR"
case arch.LOONG64:
return upper == "JMP" || upper == "JIRL" || upper == "BR"
return upper == "JMP" || upper == "JIRL" || upper == "BR" || upper == "B" ||
upper == "JAL" || upper == "BL"
default:
return upper == "JMP"
}
@@ -792,6 +840,43 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
return false
}
// shiftRotateBases are the shift and rotate mnemonics without their width
// suffix. These are the instructions whose encoder path (encodeShift) reads
// the count from the first operand.
var shiftRotateBases = map[string]bool{
"SHL": true, "SHR": true, "SAR": true, "SAL": true,
"ROL": true, "ROR": true, "RCL": true, "RCR": true,
}
// isShiftCountOperand reports whether operand i of mnem is the shift count.
// The ISA fixes the shift/rotate count register at CL: the D2/D3 group (and
// C0/C1 for immediates) encode the count outside the ModRM register field,
// so the count operand is 8-bit by definition no matter how wide the data is.
// The count arrives as the first of the two operands; the one-operand form
// does not exist.
func isShiftCountOperand(mnem string, i, nops int) bool {
if nops != 2 || i != 0 {
return false
}
if shiftRotateBases[mnem] {
return true
}
if len(mnem) > 1 {
switch mnem[len(mnem)-1] {
case 'Q', 'L', 'W', 'B':
return shiftRotateBases[mnem[:len(mnem)-1]]
}
}
return false
}
// isSetcc reports whether the mnemonic is a SETcc: SET plus a condition code.
// The membership test is the encoder's own SET dispatch, which asm.Encodable
// mirrors.
func isSetcc(mnem string) bool {
return strings.HasPrefix(mnem, "SET") && asm.Encodable(mnem)
}
// checkRegisterWidth detects amd64 register-width mismatches. The naming
// truth of the Go assembler governs: AX, BX, CX, DX, SI, DI, BP, SP and
// R8-R15 ARE the 64-bit register names (there are no separate EAX/RAX
@@ -802,6 +887,14 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
// register (EAX under the gasm alias extension, or a byte form), and byte
// registers in L/W operations.
func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
// A SETcc stores one byte: the destination is an 8-bit register or an
// 8-bit memory location by definition (0F 90+cc), whichever condition it
// tests. The trailing letter of spellings like SETPL or SETEQ is part of
// the condition code, not an operand width, so the whole family is
// exempt from the suffix logic.
if isSetcc(mnem) {
return ""
}
// Determine expected width from mnemonic suffix.
var expected int // 0=unknown, 8/4/2/1=bytes
switch {
@@ -816,10 +909,20 @@ func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
default:
return "" // no suffix, can't determine width
}
for _, op := range ops {
for i, op := range ops {
if op.Kind != ast.OpAddr || op.Addr.Sym == nil {
continue
}
// Only a bare register carries a width to compare: frame and static
// symbol references (ch+8(FP), foo(SB)) and memory operands are not
// registers even when their name collides with one.
if op.Addr.Sym.Pseudo != "" || op.Addr.Base != "" || op.Addr.Index != "" {
continue
}
// The shift/rotate count is exempt: fixed at 8 bits by the ISA.
if isShiftCountOperand(mnem, i, len(ops)) {
continue
}
name := strings.ToLower(op.Addr.Sym.Name)
regWidth := amd64RegWidth(name)
if regWidth == 0 {
+214
View File
@@ -8,6 +8,7 @@ import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
@@ -21,6 +22,21 @@ func lintSrc(t *testing.T, src string) []Diagnostic {
return File(f, Config{Arch: arch.AMD64})
}
// lintArchFile parses and lints src under a, then hands the same file to
// assemble so the assertion is pinned against the encoder: a kernel the
// linter reasons about must also be one the encoder accepts.
func lintArchFile(t *testing.T, filename, src string, a arch.Arch, assemble func(*ast.File) (*asm.Image, error)) []Diagnostic {
t.Helper()
f, errs := parser.Parse(filename, src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := assemble(f); err != nil {
t.Fatalf("encoder rejects the kernel: %v", err)
}
return File(f, Config{Arch: a})
}
// lintSrcArch lints src under the architecture inferred from filename.
func lintSrcArch(t *testing.T, filename, src string) []Diagnostic {
t.Helper()
@@ -262,6 +278,138 @@ loop:
}
}
func TestRiscvBranchFamilyRegistersLabels(t *testing.T) {
// Every riscv64 pseudo-branch that references a label must register that
// reference: the reversed branches BGT/BGTU/BLE/BLEU (GOROOT's
// memmove_riscv64 branches with BGTU) and a label named like the ZERO
// register alias (GOROOT's memclr_riscv64 carries a label named zero;
// ZERO is the ABI name of X0) must not be reported unused.
diags := lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BGTU X10, X11, backward
BGT X10, X11, zero
BLE X10, X11, one
BLEU X10, X11, two
BEQZ X10, zero
BNEZ X10, one
JMP two
backward:
RET
zero:
RET
one:
RET
two:
RET
`)
if codes(diags)[CodeUnusedLabel] != 0 {
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
}
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("defined labels must resolve: %+v", diags)
}
// A branch to a truly undefined label still reports.
diags = lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BGT X10, X11, nowhere
RET
`)
if codes(diags)[CodeUndefinedLabel] != 1 {
t.Fatalf("undefined branch target must be flagged: %+v", diags)
}
// A register-indirect JALR is not a label reference.
diags = lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
JALR X1
RET
`)
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("register operand of JALR is not a label: %+v", diags)
}
}
func TestLoong64BranchFamilyRegistersLabels(t *testing.T) {
// The loong64 jumps and single-register branches (JAL, B, BL, BEQZ/BNEZ,
// BFPT/BFPF) all reference their label from the last operand; GOROOT's
// own basic kernels tail-call with JAL, so the reference must register.
diags := lintSrcArch(t, "f_loong64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BEQZ R4, fin
BNEZ R4, fin
BLTZ R4, fin
JAL fin
BL fin
B fin
RET
fin:
RET
`)
if codes(diags)[CodeUnusedLabel] != 0 {
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
}
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("defined labels must resolve: %+v", diags)
}
}
// TestBranchFamiliesAssemble pins the lint branch sets to the encoder: every
// mnemonic the linter classifies as a riscv64 or loong64 label branch must be
// a branch the encoder actually assembles, with the label in the last
// operand. If the encoder gains or renames a branch, this test fails and the
// set follows it.
func TestBranchFamiliesAssemble(t *testing.T) {
riscvForms := map[string]string{}
for m := range riscvBranches {
riscvForms[m] = m + " X10, X11, tgt"
}
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ"} {
riscvForms[m] = m + " X10, tgt"
}
riscvForms["JMP"] = "JMP tgt"
riscvForms["JAL"] = "JAL tgt"
loongForms := map[string]string{}
for _, m := range []string{"BEQ", "BNE", "BLT", "BGE", "BLTU", "BGEU"} {
loongForms[m] = m + " R4, R5, tgt"
}
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ", "BFPT", "BFPF"} {
loongForms[m] = m + " R4, tgt"
}
loongForms["JMP"] = "JMP tgt"
loongForms["B"] = "B tgt"
loongForms["JAL"] = "JAL tgt"
loongForms["BL"] = "BL tgt"
for m, form := range riscvForms {
src := "#include \"textflag.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\t" + form + "\n" +
"tgt:\n" +
"\tRET\n"
diags := lintArchFile(t, "f_riscv64.s", src, arch.RISCV, asm.AssembleFileRISCV)
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
t.Errorf("riscv64 %s: label reference not registered: %+v", m, diags)
}
}
for m, form := range loongForms {
src := "#include \"textflag.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\t" + form + "\n" +
"tgt:\n" +
"\tRET\n"
diags := lintArchFile(t, "f_loong64.s", src, arch.LOONG64, asm.AssembleFileLOONG64)
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
t.Errorf("loong64 %s: label reference not registered: %+v", m, diags)
}
}
}
func TestInvalidTextflag(t *testing.T) {
diags := lintSrc(t, `
#include "textflag.h"
@@ -413,6 +561,72 @@ TEXT ·f(SB), NOSPLIT, $0
}
}
func TestRegisterWidthShiftCount(t *testing.T) {
// The shift and rotate count lives in CL by ISA definition (the D2/D3
// group encodes the count outside the ModRM register field), so the count
// operand is 8-bit no matter how wide the data is: SHLQ CL, AX is the
// normal spelling of a 64-bit shift. The data operand keeps its check.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
SHLQ CL, AX
SHRL CL, BX
SARQ CL, CX
ROLL CL, DX
RORQ CL, R8
RCLL CL, R9
RCRQ CL, R10
MOVQ CL, R10
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 1 {
t.Fatalf("only the MOVQ CL data move must be flagged, got %+v", diags)
}
}
func TestRegisterWidthSetcc(t *testing.T) {
// A SETcc stores one byte whichever condition it tests (0F 90+cc), so
// SETNE AL is always right and the trailing letters of SETEQ, SETPL and
// SETLS are condition codes, not width suffixes.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
CMPQ AX, BX
SETNE AL
SETEQ AL
SETPL AL
SETLS AL
SETCC (BX)
SETGE (R8)
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
t.Fatalf("SETcc destinations are 8-bit by definition: %+v", diags)
}
if codes(diags)[CodeUnknownInstr] != 0 {
t.Fatalf("every SETcc spelling must be known: %+v", diags)
}
}
func TestRegisterWidthFrameNames(t *testing.T) {
// GOROOT's BSD syscall stubs carry frame parameters whose names collide
// with byte register names (kevent's ch and nch): MOVQ ch+8(FP), SI is a
// frame reference, not the CH register.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·kevent(SB), NOSPLIT, $0-36
MOVL kq+0(FP), DI
MOVQ ch+8(FP), SI
MOVL nch+16(FP), DX
MOVQ ev+24(FP), R10
MOVQ AX, ret+32(FP)
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
t.Fatalf("frame and static symbol names are not registers: %+v", diags)
}
}
func TestNonportableRegisterName(t *testing.T) {
diags := lintSrc(t, `
#include "textflag.h"
+146
View File
@@ -0,0 +1,146 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Constant-expression folding for operands. The toolchain's assembler
// evaluates arithmetic in every operand position, and macro-heavy GOROOT
// sources lean on it: parameterised bodies carry offsets like
// ((index*4)+0)(base), immediates like $(32-shift) and masks like
// $~63 or $(1<<0|1<<9). Substituting the parameters textually therefore
// leaves constant arithmetic behind, and the parser folds it here, keeping
// the operand AST identical to what the same literals written out would
// produce. Anything that is not a closed integer expression fails to fold
// and falls through to the ordinary operand paths.
package parser
import (
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// foldExpr evaluates the constant integer expression at the head of ts and
// returns its value together with the unconsumed tokens. ok is false when
// the tokens do not form an expression, which is the callers' signal to use
// the ordinary parsing paths.
func foldExpr(ts []token.Token) (val int64, rest []token.Token, ok bool) {
v, rest, ok := foldAdd(ts)
if !ok {
return 0, ts, false
}
return v, rest, true
}
// foldAdd parses addition-level expressions: +, - and | bind loosest, the
// Plan 9 convention that makes x<<1|3 read as (x<<1)|3.
func foldAdd(ts []token.Token) (int64, []token.Token, bool) {
v, rest, ok := foldMul(ts)
if !ok {
return 0, ts, false
}
for len(rest) > 0 {
kind := rest[0].Kind
if kind != token.Plus && kind != token.Minus && kind != token.Pipe {
return v, rest, true
}
w, r2, ok := foldMul(rest[1:])
if !ok {
return v, rest, true
}
switch kind {
case token.Plus:
v += w
case token.Minus:
v -= w
case token.Pipe:
v |= w
}
rest = r2
}
return v, rest, true
}
// foldMul parses multiplication-level expressions: *, / and the bit
// operators &, << and >>.
func foldMul(ts []token.Token) (int64, []token.Token, bool) {
v, rest, ok := foldFactor(ts)
if !ok {
return 0, ts, false
}
for len(rest) > 0 {
switch rest[0].Kind {
case token.Star:
w, r2, ok := foldFactor(rest[1:])
if !ok {
return v, rest, true
}
v *= w
rest = r2
case token.Slash:
w, r2, ok := foldFactor(rest[1:])
if !ok || w == 0 {
return v, rest, true
}
v /= w
rest = r2
case token.Ampersand:
w, r2, ok := foldFactor(rest[1:])
if !ok {
return v, rest, true
}
v &= w
rest = r2
case token.LShift:
w, r2, ok := foldFactor(rest[1:])
if !ok || w < 0 || w >= 64 {
return v, rest, true
}
v <<= uint(w)
rest = r2
case token.RShift:
w, r2, ok := foldFactor(rest[1:])
if !ok || w < 0 || w >= 64 {
return v, rest, true
}
v >>= uint(w)
rest = r2
default:
return v, rest, true
}
}
return v, rest, true
}
// foldFactor parses a number, a parenthesised expression, or a unary sign
// or complement.
func foldFactor(ts []token.Token) (int64, []token.Token, bool) {
if len(ts) == 0 {
return 0, ts, false
}
switch ts[0].Kind {
case token.Number:
v, ok := tryInt(ts[0].Text)
if !ok {
return 0, ts, false
}
return v, ts[1:], true
case token.LParen:
v, rest, ok := foldAdd(ts[1:])
if !ok || len(rest) == 0 || rest[0].Kind != token.RParen {
return 0, ts, false
}
return v, rest[1:], true
case token.Minus:
v, rest, ok := foldFactor(ts[1:])
if !ok {
return 0, ts, false
}
return -v, rest, true
case token.Plus:
return foldFactor(ts[1:])
case token.Tilde:
v, rest, ok := foldFactor(ts[1:])
if !ok {
return 0, ts, false
}
return ^v, rest, true
}
return 0, ts, false
}
+169 -8
View File
@@ -32,9 +32,8 @@ func (e Error) Error() string {
// returned file is usable even when errors is non-empty.
func Parse(path, src string) (*ast.File, []error) {
tokens := lexer.Tokenize(src)
lines := splitLines(tokens)
p := &state{path: path}
p.parse(lines)
p.parse(statementLines(tokens))
return p.file, p.errs
}
@@ -73,6 +72,46 @@ func splitLines(tokens []token.Token) [][]token.Token {
return lines
}
// statementLines turns the token stream into the logical lines the parser
// reads: physical lines split at the ';' statement separators, exactly the
// way the expansion path treats the expanded bodies. The runtime writes
// "ROLQ $3, DI; ROLQ $13, DI" and "REP; MOVSB" in plain files, and the
// separator carries no meaning beyond the break. Comments are statement
// text, not structure: the lexer delivers a whole comment as one token, so
// a ';' inside a comment is never a separator; a comment after a statement
// stays on that statement's line; and a comment that sits between
// statements (the runtime's "NO_LOCAL_POINTERS; /* … */" style) stands as
// its own logical line, like a whole-line comment.
func statementLines(tokens []token.Token) [][]token.Token {
var out [][]token.Token
var cur []token.Token
flush := func() {
if len(cur) > 0 {
out = append(out, cur)
cur = nil
}
}
for _, t := range tokens {
switch t.Kind {
case token.EOF:
// The stream's terminator is not statement content.
case token.Newline, token.Semicolon:
flush()
case token.Comment:
if len(cur) > 0 {
cur = append(cur, t)
} else {
out = append(out, []token.Token{t})
}
flush()
default:
cur = append(cur, t)
}
}
flush()
return out
}
func (p *state) parse(lines [][]token.Token) {
p.file = &ast.File{Path: p.path, Macros: map[string]bool{}}
for _, line := range lines {
@@ -265,7 +304,7 @@ func (p *state) parseGlobl(line []token.Token) *ast.Globl {
rest = rest[1:]
}
if len(rest) > 0 && rest[0].Kind == token.Dollar {
g.Size = parseOperand(rest)
g.Size = parseOperand(rest, false)
}
return g
}
@@ -282,7 +321,7 @@ func (p *state) parseData(line []token.Token) *ast.Data {
d.Name = sym
d.Width = width
if len(valuePart) > 0 {
d.Value = parseOperand(stripComment(valuePart))
d.Value = parseOperand(stripComment(valuePart), false)
}
return d
}
@@ -293,8 +332,13 @@ func (p *state) parseInstr(line []token.Token) {
return
}
instr := &ast.Instr{Mnemonic: body[0], Comment: comment}
for _, grp := range splitOperands(body[1:]) {
if op := parseOperand(grp); op != nil {
grps := splitOperands(body[1:])
for i, grp := range grps {
// Only the final operand slot may carry a bare constant: the
// toolchain reads the trailing 1 of CMPSD X1, X0, 1 as $1
// (math/floor_amd64.s), while an earlier bare number names an
// absolute address, a form this parser keeps out of the tree.
if op := parseOperand(grp, i == len(grps)-1); op != nil {
instr.Operands = append(instr.Operands, op)
}
}
@@ -378,8 +422,10 @@ func setName(raw string, sym *ast.Symbol) {
// --- operand parsing --------------------------------------------------------
// parseOperand parses one operand group into an Operand.
func parseOperand(g []token.Token) *ast.Operand {
// parseOperand parses one operand group into an Operand. allowBare marks
// the final operand slot of an instruction, where the toolchain reads a
// bare constant expression as an immediate.
func parseOperand(g []token.Token, allowBare bool) *ast.Operand {
g = stripComment(g)
if len(g) == 0 {
return nil
@@ -392,9 +438,25 @@ func parseOperand(g []token.Token) *ast.Operand {
}
op.Kind = ast.OpAddr
op.Addr = parseAddress(g)
// A trailing bare constant leaves every address field empty: the
// grammar sees no register, memory reference or symbol, and the closed
// constant expression is the whole group. Read it as the immediate it
// names, exactly what the $ spelling would produce.
if allowBare && isEmptyAddress(op.Addr) {
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
op.Kind = ast.OpImmediate
op.Imm = ast.Immediate{Val: v, HasVal: true}
}
}
return op
}
// isEmptyAddress reports whether parseAddress populated nothing, its sign
// that the group is no register, memory reference, symbol or register range.
func isEmptyAddress(a ast.Address) bool {
return a.Sym == nil && a.Base == "" && a.Index == "" && a.Range == nil && a.Shift == ""
}
// parseImmediate parses the tokens following a '$'.
func parseImmediate(g []token.Token) ast.Immediate {
var imm ast.Immediate
@@ -408,6 +470,18 @@ func parseImmediate(g []token.Token) ast.Immediate {
return imm
}
}
// A constant expression introduced by '(' or '~'. Textual macro
// substitution leaves arithmetic such as $(32-shift) and $~63 behind,
// and the toolchain evaluates it in place; only shapes the ordinary
// paths below cannot read reach the folder, so every existing form
// keeps its exact parse.
if g[0].Kind == token.LParen || g[0].Kind == token.Tilde {
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
imm.Val = v
imm.HasVal = true
return imm
}
}
i := 0
if g[i].Kind == token.Minus {
imm.Neg = true
@@ -415,6 +489,17 @@ func parseImmediate(g []token.Token) ast.Immediate {
} else if g[i].Kind == token.Plus {
i++
}
// A constant expression after the sign: $-(R - 8), $+(32-shift). The
// toolchain folds the negated value in place (the cgo ABI macros write
// ADJSP $-(REGS_HOST_TO_ABI0_STACK - 8)), so the sign applies to the
// folded value exactly as it does to a bare literal.
if i < len(g) && (g[i].Kind == token.LParen || g[i].Kind == token.Tilde) {
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) == 0 {
imm.Val = v
imm.HasVal = true
return imm
}
}
if i < len(g) && g[i].Kind == token.Number {
text := g[i].Text
if v, ok := tryInt(text); ok {
@@ -446,6 +531,14 @@ func parseAddress(g []token.Token) ast.Address {
if len(g) == 0 {
return addr
}
// A bracketed register range, [Z0-Z3]: the amd64 4FMAPS/4VNNIW
// multi-source operand. The bracket runes arrive as Illegal tokens
// (the lexer has no bracket kind), so the shape matches on their text.
if isBracket(g[0], "[") && len(g) == 5 && g[1].Kind == token.Ident &&
g[2].Kind == token.Minus && g[3].Kind == token.Ident && isBracket(g[4], "]") {
addr.Range = &ast.RegRange{Lo: g[1].Text, Hi: g[3].Text, Pos: g[0].Pos}
return addr
}
// Symbol-with-pseudo form: name[<>][+off](PSEUDO).
// When the prefix is not a valid symbol name (e.g. a bare number like
// 0(SP) in RISC-V), sym is nil, and we fall through to regular memory
@@ -459,6 +552,31 @@ func parseAddress(g []token.Token) ast.Address {
}
i := 0
// A parenthesised constant expression as the displacement: substituted
// macro bodies carry ((index*4)+0)(base) shapes. As with the signed
// number path below, the value is committed only when a base group
// follows.
if i < len(g) && g[i].Kind == token.LParen {
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
addr.Offset = v
addr.HasOff = true
i = len(g) - len(rest)
}
}
// The same expression under a leading sign: -(24+8)(X6) puts the sign
// outside the fold. The base group must follow for the value to
// commit, exactly as in the unsigned branch above.
if i < len(g) && (g[i].Kind == token.Minus || g[i].Kind == token.Plus) &&
i+1 < len(g) && g[i+1].Kind == token.LParen {
if v, rest, ok := foldExpr(g[i+1:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
if g[i].Kind == token.Minus {
v = -v
}
addr.Offset = v
addr.HasOff = true
i = len(g) - len(rest)
}
}
// Optional leading displacement before a '(' base group. A sign pushes
// the parenthesis one token further out: -4(DX) has it at i+2.
if isSignedNumber(g, i) {
@@ -515,6 +633,16 @@ func parseAddress(g []token.Token) ast.Address {
}
}
}
// A lone (index*scale) group is the VSIB index-only form: the
// gather/scatter families address memory through a scaled vector index
// with no base register, 8(X4*1). The two-group grammar below reads
// (base)(index*scale), so a first group whose member carries a scale
// factor can only be an index.
if isIndexGroup(g[i:]) {
addr.Index = g[i+1].Text
addr.Scale = int(parseInt(g[i+3].Text))
i += 5
}
// First parenthesised group: the base register.
if i < len(g) && g[i].Kind == token.LParen {
i++
@@ -556,6 +684,26 @@ func parseAddress(g []token.Token) ast.Address {
if i > 0 && i < len(g) {
addr.Shift = joinRaw(g[i:])
}
// A lone (possibly signed) number is an absolute address: MOVL $0xf1,
// 0xf1 stores through the bare displacement with no base at all. In
// operand position a number without $ is an address, never a value.
if addr.Sym == nil && addr.Base == "" && addr.Index == "" && !addr.HasOff {
neg := false
j := 0
if j < len(g) && (g[j].Kind == token.Minus || g[j].Kind == token.Plus) {
neg = g[j].Kind == token.Minus
j++
}
if j == len(g)-1 && g[j].Kind == token.Number {
v := parseInt(g[j].Text)
if neg {
v = -v
}
addr.Offset = v
addr.HasOff = true
return addr
}
}
return addr
}
@@ -571,6 +719,19 @@ func findPseudoParen(g []token.Token) int {
return -1
}
// isBracket reports whether t is a square bracket. The lexer has no bracket
// kind, so '[' and ']' arrive as Illegal tokens.
func isBracket(t token.Token, text string) bool {
return t.Kind == token.Illegal && t.Text == text
}
// isIndexGroup reports whether g begins with a complete (index*scale) group:
// one identifier followed by a scale factor, all inside a single parenthesis.
func isIndexGroup(g []token.Token) bool {
return len(g) >= 5 && g[0].Kind == token.LParen && g[1].Kind == token.Ident &&
g[2].Kind == token.Star && g[3].Kind == token.Number && g[4].Kind == token.RParen
}
// --- token helpers ----------------------------------------------------------
// splitOperands splits a token slice on top-level commas (commas outside any
+257
View File
@@ -401,3 +401,260 @@ func TestInt64MinimumImmediate(t *testing.T) {
t.Errorf("imm.Float = %q, want empty", imm.Float)
}
}
// TestDivisionSlashPackagePath covers the runtime's package-path spelling:
// U+2215 DIVISION SLASH separates the elements of an import path inside a
// symbol (internal∕runtime∕atomic·Xchg), and the middle dot still separates
// the package from the name. The whole spelling must reach the symbol, not
// stop at the first slash.
func TestDivisionSlashPackagePath(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tCALL internal∕runtime∕atomic·Xchg(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
instr := txt.Body[0].(*ast.Instr)
sym := instr.Operands[0].Addr.Sym
if sym == nil {
t.Fatal("operand carries no symbol")
}
if sym.Pkg != "internal∕runtime∕atomic" {
t.Errorf("pkg = %q, want internal∕runtime∕atomic", sym.Pkg)
}
if sym.Name != "Xchg" {
t.Errorf("name = %q, want Xchg", sym.Name)
}
if sym.Raw != "internal∕runtime∕atomic·Xchg(SB)" {
t.Errorf("raw = %q", sym.Raw)
}
}
// TestSemicolonStatements covers the plain parse path: ';' separates
// statements on one line exactly as it does inside macro expansion, and a
// ';' inside a comment is comment text.
func TestSemicolonStatements(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tROLQ $3, DI; ROLQ $13, DI\n\tMOVQ AX, BX // note; still comment\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4", len(txt.Body))
}
first := txt.Body[0].(*ast.Instr)
if first.Mnemonic.Text != "ROLQ" || len(first.Operands) != 2 {
t.Errorf("first statement = %+v, want ROLQ with two operands", first.Mnemonic)
}
second := txt.Body[1].(*ast.Instr)
if second.Mnemonic.Text != "ROLQ" || len(second.Operands) != 2 {
t.Errorf("second statement = %s, want ROLQ with two operands", second.Mnemonic.Text)
}
// The trailing comment belongs to the second MOVQ, semicolon included.
third := txt.Body[2].(*ast.Instr)
if third.Mnemonic.Text != "MOVQ" || third.Comment != "note; still comment" {
t.Errorf("third = %s, comment %q", third.Mnemonic.Text, third.Comment)
}
}
// TestSemicolonAfterLabel covers a label sharing its line with two
// statements.
func TestSemicolonAfterLabel(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\nloop: NOP; NOP\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4 (label, two instructions, RET)", len(txt.Body))
}
if _, ok := txt.Body[0].(*ast.Label); !ok {
t.Errorf("first statement = %T, want *ast.Label", txt.Body[0])
}
for i, want := range []string{"NOP", "NOP", "RET"} {
in, ok := txt.Body[i+1].(*ast.Instr)
if !ok || in.Mnemonic.Text != want {
t.Errorf("statement %d = %v, want %s", i+1, txt.Body[i+1], want)
}
}
}
// TestParseEqualsZeroOptions pins the contract that ParseWithOptions with
// the zero Options reproduces Parse, here for the semicolon split.
func TestParseEqualsZeroOptions(t *testing.T) {
src := "TEXT \u00b7f(SB), $0\n\tNOP; NOP\n\tRET\n"
a, errsA := Parse("t.s", src)
b, errsB := ParseWithOptions("t.s", src, Options{})
if len(errsA) > 0 || len(errsB) > 0 {
t.Fatalf("errors: %v / %v", errsA, errsB)
}
ta, tb := texts(a), texts(b)
if len(ta) != len(tb) {
t.Fatalf("decl counts differ: %d vs %d", len(ta), len(tb))
}
for i := range ta {
if len(ta[i].Body) != len(tb[i].Body) {
t.Fatalf("TEXT %d: body lengths differ: %d vs %d", i, len(ta[i].Body), len(tb[i].Body))
}
}
}
// TestBracketRegisterRange pins the amd64 multi-source operand of the
// 4FMAPS/4VNNIW families: the bracket group [Z0-Z3] names four consecutive
// source registers and must reach the AST as a register range instead of an
// empty address.
func TestBracketRegisterRange(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), [Z0-Z3], K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
if len(in.Operands) != 4 {
t.Fatalf("operands = %d, want 4", len(in.Operands))
}
rng := in.Operands[1]
if rng.Kind != ast.OpAddr {
t.Errorf("range operand kind = %v, want OpAddr", rng.Kind)
}
if rng.Addr.Range == nil {
t.Fatalf("range operand = %+v, want a register range", rng.Addr)
}
if rng.Addr.Range.Lo != "Z0" || rng.Addr.Range.Hi != "Z3" {
t.Errorf("range = %s-%s, want Z0-Z3", rng.Addr.Range.Lo, rng.Addr.Range.Hi)
}
if rng.Addr.Sym != nil || rng.Addr.Base != "" || rng.Addr.Index != "" || rng.Addr.Shift != "" {
t.Errorf("range operand carries stray address fields: %+v", rng.Addr)
}
if rng.Raw != "[ Z0 - Z3 ]" {
t.Errorf("range raw = %q, want the verbatim spelling", rng.Raw)
}
}
// TestBracketRegisterRangeNotList pins that arm64-style register lists, whose
// members carry arrangements, stay out of the simple range shape: they remain
// plain bracketed groups the arm64 encoder reads from Raw. A comma inside
// brackets is a top-level comma, so a multi-member list spans several
// operands, exactly the shape the arm64 encoder's list scan stitches back.
func TestBracketRegisterRangeNotList(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []string{"[ V21.B16 ]", "V3.B16 ]"} {
in := fn.Body[i].(*ast.Instr)
op := in.Operands[len(in.Operands)-1]
if op.Addr.Range != nil {
t.Errorf("%s: range = %v, want nil", in.Mnemonic.Text, op.Addr.Range)
}
if op.Raw != want {
t.Errorf("operand %d raw = %q, want %q", i, op.Raw, want)
}
}
}
// TestVSIBIndexOnly pins the gather/scatter memory operand with a scaled
// vector index and no base register: 8(X4*1) must carry index and scale and
// leave the base empty, not strand the scale in the shift suffix.
func TestVSIBIndexOnly(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVPGATHERDQ Y0, 8(X4*1), Y6\n\tVPGATHERDQ Y0, (X4*2), Y6\n\tVPGATHERDQ Y0, -8(X4*1), Y6\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
want := []ast.Address{
{Index: "X4", Scale: 1, Offset: 8, HasOff: true},
{Index: "X4", Scale: 2},
{Index: "X4", Scale: 1, Offset: -8, HasOff: true},
}
for i, w := range want {
in := fn.Body[i].(*ast.Instr)
a := in.Operands[1].Addr
if a.Base != "" || a.Index != w.Index || a.Scale != w.Scale || a.Offset != w.Offset || a.HasOff != w.HasOff || a.Shift != "" {
t.Errorf("operand %d = %+v, want %+v", i, a, w)
}
}
}
// TestVSIBTwoGroupKeepsBase pins that the ordinary (base)(index*scale)
// grammar is untouched by the index-only recognition.
func TestVSIBTwoGroupKeepsBase(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVP4DPWSSD 7(SI)(DI*1), [Z2-Z5], K4, Z17\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
a := in.Operands[0].Addr
if a.Base != "SI" || a.Index != "DI" || a.Scale != 1 || a.Offset != 7 || !a.HasOff {
t.Errorf("address = %+v, want base SI index DI scale 1 offset 7", a)
}
if in.Operands[1].Addr.Range == nil || in.Operands[1].Addr.Range.Lo != "Z2" || in.Operands[1].Addr.Range.Hi != "Z5" {
t.Errorf("second operand = %+v, want range Z2-Z5", in.Operands[1].Addr)
}
}
// TestBareTrailingImmediate pins the toolchain's bare constant spelling in
// the final operand slot: CMPSD X1, X0, 1 reads as $1 (math/floor_amd64.s).
// Earlier slots keep the strict grammar, so a bare number there stays an
// address rather than becoming an immediate.
func TestBareTrailingImmediate(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tCMPSD X1, X0, 1\n\tCMPSD X1, X0, -1\n\tADDQ AX, 1+2\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []int64{1, -1, 3} {
in := fn.Body[i].(*ast.Instr)
last := in.Operands[len(in.Operands)-1]
if last.Kind != ast.OpImmediate || !last.Imm.HasVal || last.Imm.Val != want {
t.Errorf("operand %d = %+v, want immediate %d", i, last, want)
}
}
// A bare number outside the final slot is not an immediate.
file2, errs2 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tADDQ 1, AX\n\tRET\n")
if len(errs2) > 0 {
t.Fatalf("parse errors: %v", errs2)
}
fn2 := file2.Decls[0].(*ast.Text)
first := fn2.Body[0].(*ast.Instr).Operands[0]
if first.Kind != ast.OpAddr {
t.Errorf("non-final bare number kind = %v, want OpAddr", first.Kind)
}
// A bare name in the final slot stays a symbol: labels are names, not
// constants, and jump targets depend on the distinction.
file3, errs3 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tJMP loop\nloop: NOP\n\tRET\n")
if len(errs3) > 0 {
t.Fatalf("parse errors: %v", errs3)
}
fn3 := file3.Decls[0].(*ast.Text)
jmp := fn3.Body[0].(*ast.Instr)
if jmp.Operands[0].Kind != ast.OpAddr || jmp.Operands[0].Addr.Sym == nil || jmp.Operands[0].Addr.Sym.Name != "loop" {
t.Errorf("jump target = %+v, want label loop", jmp.Operands[0])
}
}
// TestSignedParenDisplacement pins a sign before a parenthesised
// displacement expression: -(24+8)(X6) negates the folded value and keeps
// the base group, the shape GOROOT's riscv64 and loong64 files use.
func TestSignedParenDisplacement(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0-0\n\tMOV X7, -(24+8)(X6)\n\tMOV X7, +(16)(X6)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
text := file.Decls[0].(*ast.Text)
ins := text.Body[0].(*ast.Instr)
op := ins.Operands[1] // Plan 9 order: the destination address is last
if !op.Addr.HasOff || op.Addr.Offset != -32 {
t.Errorf("-(24+8): offset = %v hasOff=%v, want -32 true", op.Addr.Offset, op.Addr.HasOff)
}
if op.Addr.Base != "X6" {
t.Errorf("-(24+8): base = %q, want X6", op.Addr.Base)
}
ins = text.Body[1].(*ast.Instr)
op = ins.Operands[1]
if !op.Addr.HasOff || op.Addr.Offset != 16 || op.Addr.Base != "X6" {
t.Errorf("+(16): offset = %v hasOff=%v base=%q, want 16 true X6", op.Addr.Offset, op.Addr.HasOff, op.Addr.Base)
}
}
+571
View File
@@ -0,0 +1,571 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// The preprocessor turns #define and #include directives into the token
// stream the parser really sees, the way the Go toolchain's assembler does:
// object and parameterised macros expand at the point of use, and an
// #include splices the named file's lines in place of the directive. The
// pass runs only on the assembly path (gasm asm, diff, the corpus audit),
// where the result is machine code; parsing for the linter, formatter and
// language server keeps the raw file so their view of #define lines, and
// therefore their macro-aware behaviour, is unchanged.
package parser
import (
"fmt"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"unicode/utf8"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// Options controls the optional preprocessing applied before a file is
// parsed. The zero value reproduces Parse exactly.
type Options struct {
// IncludeDirs lists the -I directories searched for #include files,
// in order, after the including file's own directory.
IncludeDirs []string
// Expand enables macro expansion, include splicing and the
// statement-separator reading of ';' that the expanded bodies rely on.
Expand bool
// Predefines names the macros defined before the file is read. The
// go command drives go tool asm with -D GOOS_<goos> -D GOARCH_<arch>,
// and GOROOT's own headers (go_tls.h, asm_riscv64.h) select their
// platform blocks with #ifdef on exactly those names, so an assembler
// without them cannot see the platform definitions at all.
Predefines map[string]string
}
// ParseWithOptions parses src like Parse, optionally preprocessing it first.
// The returned file is usable even when errors is non-empty.
func ParseWithOptions(path, src string, opts Options) (*ast.File, []error) {
tokens := lexer.Tokenize(src)
var lines [][]token.Token
var errs []error
if opts.Expand {
pp := &preproc{opts: opts, macros: map[string]*macroDef{}}
for name, value := range opts.Predefines {
pp.macros[name] = &macroDef{name: name, body: lexer.Tokenize(value)}
}
lines = pp.fileLines(path, tokens, token.Position{})
errs = pp.errs
} else {
lines = statementLines(tokens)
}
p := &state{path: path}
p.parse(lines)
return p.file, append(errs, p.errs...)
}
// maxExpansionDepth bounds recursive macro expansion; the toolchain's
// assembler gives up after 100 nested invocations without producing a token.
const maxExpansionDepth = 100
// textflagHeader names the one header gasm does not splice: its flag macros
// (NOSPLIT, RODATA, …) are consumed by name throughout gasm's parser,
// encoders and linter, and expanding them to their numeric constants would
// leave every consumer blind to them.
const textflagHeader = "textflag.h"
// macroDef is one #define. A nil args slice is an object macro; a non-nil
// (possibly empty) one is parameterised, the C distinction between
// "#define A(x)" and "#define A (x)".
type macroDef struct {
name string
args []string
body []token.Token
}
// preproc carries the state of one expansion pass: the live macro table, the
// chain of files currently being read, for cycle detection, and the
// conditional-inclusion stack of #ifdef regions.
type preproc struct {
opts Options
macros map[string]*macroDef
errs []error
stack []string // absolute paths of files being read, innermost last
ifdefStack []bool // one entry per open #ifdef/#ifndef, its truth
}
// enabled reports whether the position being read is inside a live
// conditional branch. Directives inside a disabled branch contribute
// nothing, and its content lines are dropped, exactly as the toolchain's
// input stack does.
func (pp *preproc) enabled() bool {
return len(pp.ifdefStack) == 0 || pp.ifdefStack[len(pp.ifdefStack)-1]
}
func (pp *preproc) errorf(pos token.Position, format string, args ...any) {
pp.errs = append(pp.errs, Error{Pos: pos, Msg: fmt.Sprintf(format, args...)})
}
// fileLines tokenizes and preprocesses one file into logical lines.
// Directive lines are kept (the parser records them for the tooling);
// #include lines are replaced by the included file's lines. includePos is
// the position of the #include that pulled this file in, zero for the
// top-level file, and only serves cycle diagnostics.
func (pp *preproc) fileLines(path string, tokens []token.Token, includePos token.Position) [][]token.Token {
abs, err := filepath.Abs(path)
if err != nil {
abs = filepath.Clean(path)
}
if slices.Contains(pp.stack, abs) {
if includePos.IsValid() {
pp.errorf(includePos, "#include %q: include cycle (%s is already being read)", path, filepath.Base(path))
}
return nil
}
pp.stack = append(pp.stack, abs)
var out [][]token.Token
for _, line := range splitLines(tokens) {
if len(line) == 0 {
out = append(out, line)
continue
}
if line[0].Kind == token.Hash {
out = append(out, pp.directive(line, filepath.Dir(path))...)
continue
}
if !pp.enabled() {
continue
}
out = append(out, splitOnSemicolons(pp.expandTokens(line))...)
}
pp.stack = pp.stack[:len(pp.stack)-1]
if len(pp.stack) == 0 && len(pp.ifdefStack) > 0 {
// The stack is per-input, shared across includes, so only the
// top-level file's end can decide the input was left unclosed.
pp.errorf(token.Position{Line: 1, Column: 1}, "unclosed #ifdef or #ifndef")
}
return out
}
// directive processes one '#' line and returns the lines to keep in the
// stream: every directive line is kept as-is for the parser (which records
// it), except #include, which is replaced by the spliced content.
// Conditionals are tracked on every line; every other directive is inert
// inside a disabled branch.
func (pp *preproc) directive(line []token.Token, dir string) [][]token.Token {
if len(line) < 2 || line[1].Kind != token.Ident {
return [][]token.Token{line}
}
switch line[1].Text {
case "ifdef", "ifndef":
pp.ifdef(line, line[1].Text == "ifndef")
case "else":
pp.elseBranch(line)
case "endif":
pp.endif(line)
case "define":
if pp.enabled() {
pp.define(line)
}
case "undef":
if pp.enabled() {
pp.undef(line)
}
case "include":
if pp.enabled() {
return pp.include(line, dir)
}
default:
// #line and unknown directives are recorded but not interpreted:
// conservative support keeps the parser's view intact and files
// using them fail on their content, not silently.
}
return [][]token.Token{line}
}
// ifdef handles "#ifdef NAME" and "#ifndef NAME", pushing the branch's truth
// onto the conditional stack. A branch opened inside a disabled region is
// itself disabled, however the name resolves.
func (pp *preproc) ifdef(line []token.Token, inverted bool) {
truth := false
if len(line) >= 3 && line[2].Kind == token.Ident {
_, defined := pp.macros[line[2].Text]
truth = defined != inverted
} else {
pp.errorf(line[0].Pos, "expected identifier after #%s", line[1].Text)
}
if !pp.enabled() {
truth = false
}
pp.ifdefStack = append(pp.ifdefStack, truth)
}
// elseBranch flips the innermost conditional's truth, but only when the
// region enclosing it is itself live: the toolchain keeps outer overrides.
func (pp *preproc) elseBranch(line []token.Token) {
if len(pp.ifdefStack) == 0 {
pp.errorf(line[0].Pos, "unmatched #else")
return
}
if len(pp.ifdefStack) == 1 || pp.ifdefStack[len(pp.ifdefStack)-2] {
pp.ifdefStack[len(pp.ifdefStack)-1] = !pp.ifdefStack[len(pp.ifdefStack)-1]
}
}
// endif closes the innermost conditional.
func (pp *preproc) endif(line []token.Token) {
if len(pp.ifdefStack) == 0 {
pp.errorf(line[0].Pos, "unmatched #endif")
return
}
pp.ifdefStack = pp.ifdefStack[:len(pp.ifdefStack)-1]
}
// define parses "#define NAME[(formals)] body" into the macro table. The
// body runs to the end of the logical line (the lexer has already spliced
// backslash continuations) and stops at a comment, which never expands.
func (pp *preproc) define(line []token.Token) {
if len(line) < 3 || line[2].Kind != token.Ident {
return
}
name := line[2]
args := []string(nil)
body := line[3:]
// The definition is parameterised only when '(' follows the name
// directly; the toolchain separates "#define A(x)" from
// "#define A (x)" by adjacency, and so does the column check here.
if len(body) > 0 && body[0].Kind == token.LParen &&
body[0].Pos.Column == name.Pos.Column+utf8.RuneCountInString(name.Text) {
args = []string{}
i := 1
for i < len(body) && body[i].Kind != token.RParen {
if body[i].Kind == token.Ident {
args = append(args, body[i].Text)
}
i++
}
if i < len(body) {
body = body[i+1:]
} else {
body = nil
}
}
if i := slices.IndexFunc(body, func(t token.Token) bool { return t.Kind == token.Comment }); i >= 0 {
body = body[:i]
}
if _, exists := pp.macros[name.Text]; exists {
// The toolchain refuses redefinition, so a file the oracle accepts
// never redefines; failing here keeps that contract visible.
pp.errorf(name.Pos, "redefinition of macro %s", name.Text)
}
pp.macros[name.Text] = &macroDef{name: name.Text, args: args, body: pp.bodyWithBreaks(body)}
}
// bodyWithBreaks records the statement boundaries the continuations carry.
// The lexer splices backslash-continued lines into one logical line, but the
// toolchain keeps the newline as a token in the stored body, which is how a
// multi-instruction body without semicolons (the arm64 style) still splits
// into statements on expansion. A line change inside the logical line is
// exactly a continuation, so the boundary is restored from the positions.
func (pp *preproc) bodyWithBreaks(body []token.Token) []token.Token {
out := make([]token.Token, 0, len(body))
for i, t := range body {
if i > 0 && t.Pos.Line != body[i-1].Pos.Line {
out = append(out, token.Token{Kind: token.Newline, Text: "\n", Pos: t.Pos, End: t.Pos})
}
out = append(out, t)
}
return out
}
// undef handles "#undef NAME", which the toolchain honours and requires to
// name a defined macro.
func (pp *preproc) undef(line []token.Token) {
if len(line) < 3 || line[2].Kind != token.Ident {
return
}
if _, ok := pp.macros[line[2].Text]; !ok {
pp.errorf(line[2].Pos, "#undef for undefined macro %s", line[2].Text)
return
}
delete(pp.macros, line[2].Text)
}
// include resolves and splices "#include \"file\"". A header that cannot be
// read keeps the directive line in the stream, with a diagnostic.
func (pp *preproc) include(line []token.Token, dir string) [][]token.Token {
if len(line) < 3 || line[2].Kind != token.String {
return [][]token.Token{line}
}
header := line[2]
name, err := strconv.Unquote(header.Text)
if err != nil {
pp.errorf(header.Pos, "unquoting include file name: %v", err)
return [][]token.Token{line}
}
if filepath.Base(name) == textflagHeader {
// Flag macros are handled natively (see textflagHeader); the
// directive stays so tools still see the include.
return [][]token.Token{line}
}
resolved, ok := pp.resolve(name, dir)
if !ok {
searched := append([]string{dir}, pp.opts.IncludeDirs...)
pp.errorf(header.Pos, "#include %q: file not found (searched %s)", name, strings.Join(searched, ", "))
return [][]token.Token{line}
}
src, err := os.ReadFile(resolved)
if err != nil {
pp.errorf(header.Pos, "#include %q: %v", name, err)
return [][]token.Token{line}
}
return pp.fileLines(resolved, lexer.Tokenize(string(src)), header.Pos)
}
// resolve looks an include name up the way the toolchain does: as written
// (relative to the working directory), then relative to the including
// file's directory, then in each -I directory in order.
func (pp *preproc) resolve(name, dir string) (string, bool) {
candidates := []string{name}
if !filepath.IsAbs(name) {
candidates = append(candidates, filepath.Join(dir, name))
for _, d := range pp.opts.IncludeDirs {
candidates = append(candidates, filepath.Join(d, name))
}
}
for _, c := range candidates {
if st, err := os.Stat(c); err == nil && !st.IsDir() {
return c, true
}
}
return "", false
}
// expandTokens expands every macro invocation in a token sequence,
// recursively, with a depth guard. A body is spliced into the sequence in
// place and rescanned, the way the toolchain's input stack re-reads pushed
// tokens: an object macro may name a parameterised one, and the argument
// list of the expansion may then come from the tokens that follow.
func (pp *preproc) expandTokens(in []token.Token) []token.Token {
s := in
i := 0
consecutive := 0
for i < len(s) {
t := s[i]
if t.Kind != token.Ident {
i++
consecutive = 0
continue
}
def, suffix := pp.macroFor(t.Text)
if def == nil {
i++
consecutive = 0
continue
}
// The guard mirrors the toolchain's: 100 nested invocations in a
// row without a plain token between them means recursion.
consecutive++
if consecutive > maxExpansionDepth {
pp.errorf(t.Pos, "recursive macro invocation (deeper than %d levels)", maxExpansionDepth)
return nil
}
if def.args == nil {
body := restamp(def.body, t.Pos)
if suffix != "" {
// The macro was reached only through a compound spelling
// (ACC0.B16 over "#define ACC0 V8"), so the selector has
// to travel with the expansion.
body = appendSelector(body, suffix, t.Pos)
}
s = append(s[:i], append(body, s[i+1:]...)...)
continue
}
// A parameterised macro invoked without its parentheses stands
// unexpanded, naming itself, as in the toolchain.
if i+1 >= len(s) || s[i+1].Kind != token.LParen {
i++
consecutive = 0
continue
}
args, next := pp.collectArgs(s, i+1, t)
if args == nil {
return nil
}
// A zero-argument macro may be invoked as NAME().
if len(def.args) == 0 && len(args) == 1 && len(args[0]) == 0 {
args = nil
}
if len(args) != len(def.args) {
pp.errorf(t.Pos, "wrong arg count for macro %s: got %d, want %d", t.Text, len(args), len(def.args))
i = next
consecutive = 0
continue
}
sub := make([]token.Token, 0, len(def.body))
for _, bt := range def.body {
if bt.Kind == token.Ident {
if k := slices.Index(def.args, bt.Text); k >= 0 {
sub = append(sub, restamp(args[k], t.Pos)...)
continue
}
// A parameter used with an element or lane selector: the
// lexer folds A.S4 into one identifier, so the whole-token
// match above cannot see the parameter. The toolchain
// lexes the period separately and substitutes the name
// alone; splitting at the FIRST period and pasting the
// argument back in front of the selector is the equivalent
// for this lexer.
if k, sel := parameterSelector(bt.Text, def.args); k >= 0 {
sub = append(sub, restamp(pasteSelector(args[k], sel), t.Pos)...)
continue
}
}
sub = append(sub, bt)
}
s = append(s[:i], append(sub, s[next:]...)...)
}
return s
}
// macroFor finds the macro a use names. The lexer folds NAME.selector into
// one identifier token, so a macro written behind a selector suffix
// (ACC0.B16 over "#define ACC0 V8") never matches a whole-token table
// lookup; the toolchain splits on the period and reads the two halves, so
// the prefix before the FIRST period is tried here as well and the caller
// re-attaches the suffix to whatever the macro expands to. Only a whole
// name counts: AB.S4 does not reach a macro named A, and a parameterised
// macro is not hidden behind a selector, because its invocation would need
// the parentheses to follow the bare name.
func (pp *preproc) macroFor(text string) (*macroDef, string) {
if def := pp.macros[text]; def != nil {
return def, ""
}
if j := strings.IndexByte(text, '.'); j > 0 {
if def := pp.macros[text[:j]]; def != nil && def.args == nil {
return def, text[j:]
}
}
return nil, ""
}
// appendSelector glues a selector suffix onto an object macro's expansion:
// the selector binds to the identifier the expansion ends with, the way the
// toolchain's operand parser reads V0 and .B16 back as one register
// spelling. An expansion that does not end in an identifier carries the
// selector as its own token, which the parser then reports where it cannot
// parse it.
func appendSelector(body []token.Token, suffix string, pos token.Position) []token.Token {
if n := len(body); n > 0 && body[n-1].Kind == token.Ident {
body[n-1].Text += suffix
return body
}
return append(body, token.Token{Kind: token.Ident, Text: suffix, Pos: pos, End: pos})
}
// parameterSelector reports the argument a compound body token names: the
// parameter whose whole name occupies the text before the token's FIRST
// period, with the selector that follows. k is negative when no parameter
// matches, which leaves tokens like AB.S4 untouched even though a parameter
// A is bound.
func parameterSelector(text string, args []string) (int, string) {
j := strings.IndexByte(text, '.')
if j <= 0 {
return -1, ""
}
if k := slices.Index(args, text[:j]); k >= 0 {
return k, text[j:]
}
return -1, ""
}
// pasteSelector joins an argument with the selector a compound body token
// carries, textually: the selector binds to the identifier the argument
// ends with, so A.S4 over the argument V0.B16 spells V0.B16.S4, exactly the
// operand the toolchain's split-then-substitute leaves behind. An argument
// with no trailing identifier carries the selector as a separate token,
// which the parser then reports where it cannot parse it.
func pasteSelector(val []token.Token, suffix string) []token.Token {
if len(val) == 0 {
return []token.Token{{Kind: token.Ident, Text: suffix}}
}
out := slices.Clone(val)
if n := len(out); out[n-1].Kind == token.Ident {
out[n-1].Text += suffix
return out
}
return append(out, token.Token{Kind: token.Ident, Text: suffix})
}
// collectArgs reads the actual argument tokens of an invocation; the opening
// parenthesis is at start. Commas separate arguments except inside nested
// parentheses. A nil result means the list was unterminated, which is a
// diagnostic.
func (pp *preproc) collectArgs(in []token.Token, start int, name token.Token) ([][]token.Token, int) {
var args [][]token.Token
var cur []token.Token
nesting := 0
for i := start + 1; i < len(in); i++ {
t := in[i]
switch t.Kind {
case token.LParen:
nesting++
cur = append(cur, t)
case token.RParen:
if nesting == 0 {
return append(args, cur), i + 1
}
nesting--
cur = append(cur, t)
case token.Comma:
if nesting == 0 {
args = append(args, cur)
cur = nil
continue
}
cur = append(cur, t)
case token.Comment:
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
return nil, i
default:
cur = append(cur, t)
}
}
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
return nil, len(in)
}
// restamp copies body tokens to the invocation's position, so diagnostics
// and the line table point where the macro was used, as the toolchain's
// input stack does.
func restamp(body []token.Token, pos token.Position) []token.Token {
out := make([]token.Token, len(body))
for i, t := range body {
t.Pos, t.End = pos, pos
out[i] = t
}
return out
}
// splitOnSemicolons breaks a token sequence at ';' statement separators and
// at the Newline markers that record continuation boundaries inside macro
// bodies, producing the logical lines the parser expects. The separators
// carry no meaning beyond the break, so the pieces are exactly what the same
// statements on separate lines would produce.
func splitOnSemicolons(ts []token.Token) [][]token.Token {
var out [][]token.Token
start := 0
for i, t := range ts {
if t.Kind == token.Semicolon || t.Kind == token.Newline {
if i > start {
out = append(out, ts[start:i])
}
start = i + 1
}
}
if start < len(ts) {
out = append(out, ts[start:])
}
return out
}
+675
View File
@@ -0,0 +1,675 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package parser
import (
"os"
"path/filepath"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
)
// expand parses src with preprocessing enabled and returns the first TEXT's
// body instructions as "MNEMONIC operand|operand" strings, the shape the
// expansion assertions below compare against. Runs of spaces are
// collapsed: Raw renders a token group as its tokens joined with single
// spaces, so "$(32-7)" arrives as "$ ( 32 - 7 )" and the comparison must
// not depend on that spelling.
func expand(t *testing.T, src string) (*ast.File, []string) {
t.Helper()
f, errs := ParseWithOptions("t_amd64.s", src, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
ts := texts(f)
if len(ts) == 0 {
t.Fatalf("no TEXT in:\n%s", src)
}
var got []string
for _, s := range ts[0].Body {
in, ok := s.(*ast.Instr)
if !ok {
continue
}
var ops []string
for _, op := range in.Operands {
ops = append(ops, op.Raw)
}
line := in.Mnemonic.Text + " " + strings.Join(ops, ", ")
got = append(got, strings.ReplaceAll(line, " ", ""))
}
return f, got
}
func wantLines(t *testing.T, got []string, want ...string) {
t.Helper()
strip := func(lines []string) string {
var out []string
for _, l := range lines {
out = append(out, strings.ReplaceAll(l, " ", ""))
}
return strings.Join(out, "\n")
}
if strip(got) != strip(want) {
t.Errorf("expanded body:\n %s\nwant:\n %s", strings.Join(got, "\n "), strings.Join(want, "\n "))
}
}
func TestObjectMacroExpandsAtUse(t *testing.T) {
_, got := expand(t, `
#define REGTMP CX
#define TWICE ADDQ CX, AX; ADDQ CX, AX
TEXT ·f(SB), NOSPLIT, $0
MOVQ 8(SP), REGTMP
TWICE
RET
`)
wantLines(t, got,
"MOVQ 8(SP), CX",
"ADDQ CX, AX",
"ADDQ CX, AX",
"RET",
)
}
func TestParameterisedMacroSubstitutesArguments(t *testing.T) {
f, errs := ParseWithOptions("t_amd64.s", `
#define ROUND1(a, index, const, shift) \
ADDQ $const, a; \
MOVW (index*4)(SP), a; \
RORQ $(32-shift), a
TEXT ·f(SB), NOSPLIT, $0
ROUND1(AX, 3, 0xd76aa478, 7)
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
body := texts(f)[0].Body
add := body[0].(*ast.Instr)
if add.Mnemonic.Text != "ADDQ" || !add.Operands[0].Imm.HasVal ||
add.Operands[0].Imm.Val != 0xd76aa478 || add.Operands[1].Addr.Sym == nil ||
add.Operands[1].Addr.Sym.Name != "AX" {
t.Errorf("ADDQ operands substituted wrong: %+v %+v", add.Operands[0].Imm, add.Operands[1].Addr)
}
mov := body[1].(*ast.Instr)
if addr := mov.Operands[0].Addr; !addr.HasOff || addr.Offset != 12 {
t.Errorf("MOVW offset = %+v, want 12 from 3*4", addr)
}
ror := body[2].(*ast.Instr)
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
t.Errorf("RORQ immediate = %+v, want 25 from (32-7)", ror.Operands[0].Imm)
}
}
func TestMacroArgumentsKeepCommasInParens(t *testing.T) {
// An argument may itself be an unparenthesised expression: the tokens
// substitute verbatim and the parser folds the result, as the
// toolchain's parser does.
f, errs := ParseWithOptions("t_amd64.s", `
#define LOAD(dst, off) MOVQ off(SP), dst
TEXT ·f(SB), NOSPLIT, $0
LOAD(AX, 1*8)
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
addr := in.Operands[0].Addr
if !addr.HasOff || addr.Offset != 8 {
t.Errorf("offset = %+v, want 8", addr)
}
if sym := in.Operands[1].Addr.Sym; sym == nil || sym.Name != "AX" {
t.Errorf("destination = %+v, want AX", in.Operands[1].Addr)
}
}
func TestNestedMacroInvocations(t *testing.T) {
// An object macro naming a parameterised one, and a parameterised body
// invoking another parameterised macro: the toolchain's input stack
// rescans substituted tokens, and so does expansion here.
_, got := expand(t, `
#define DOUBLE(x) ADDQ x, x
#define TWICE2 DOUBLE
#define FOUR(a, b) DOUBLE(a); DOUBLE(b)
TEXT ·f(SB), NOSPLIT, $0
TWICE2(AX)
FOUR(AX, CX)
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"ADDQ AX, AX",
"ADDQ CX, CX",
"RET",
)
}
func TestMultiLineBodySplitsWithoutSemicolons(t *testing.T) {
// The arm64 style: backslash-continued lines with no semicolons. The
// continuation newline is a statement boundary, as in the toolchain.
_, got := expand(t, `
#define PAIR \
ADDQ AX, AX \
MOVQ AX, CX
TEXT ·f(SB), NOSPLIT, $0
PAIR
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"MOVQ AX, CX",
"RET",
)
}
func TestZeroArgumentMacro(t *testing.T) {
_, got := expand(t, `
#define BARRIER()
TEXT ·f(SB), NOSPLIT, $0
BARRIER()
RET
`)
wantLines(t, got, "RET")
}
func TestParameterisedWithoutParensStandsAsName(t *testing.T) {
// A parameterised macro invoked without its parentheses names itself,
// which the parser then reports as an unknown instruction rather than
// silently expanding nothing.
f, errs := ParseWithOptions("t_amd64.s", `
#define M(x) ADDQ x, x
TEXT ·f(SB), NOSPLIT, $0
M
RET
`, Options{Expand: true})
if len(errs) != 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
if len(fn.Body) == 0 {
t.Fatal("body empty")
}
in, ok := fn.Body[0].(*ast.Instr)
if !ok || in.Mnemonic.Text != "M" {
t.Fatalf("bare parameterised macro did not stand as its name: %+v", fn.Body[0])
}
}
func TestDefinitionScoping(t *testing.T) {
// A definition applies from its point onward: the use before the
// #define stays untouched.
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
SPECIAL
#define SPECIAL ADDQ AX, AX
SPECIAL
RET
`)
wantLines(t, got,
"SPECIAL",
"ADDQ AX, AX",
"RET",
)
}
func TestUndefRemovesMacro(t *testing.T) {
_, got := expand(t, `
#define TEMP AX
TEXT ·f(SB), NOSPLIT, $0
TEMP
#undef TEMP
TEMP
RET
`)
wantLines(t, got,
"AX",
"TEMP",
"RET",
)
}
func TestUndefUndefinedMacroIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#undef NOSUCH\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "undefined macro NOSUCH") {
t.Fatalf("#undef of an undefined macro: got %v, want an error naming it", errs)
}
}
func TestRedefinitionIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define A X\n#define A Y\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "redefinition of macro A") {
t.Fatalf("redefinition: got %v, want an error", errs)
}
}
func TestRecursiveMacroIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define A B\n#define B A\nTEXT ·f(SB), NOSPLIT, $0\n\tA\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "recursive macro invocation") {
t.Fatalf("recursion: got %v, want a recursive-macro error, not a hang", errs)
}
}
func TestWrongArgumentCountIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define M(a, b) ADDQ a, b\nTEXT ·f(SB), NOSPLIT, $0\n\tM(AX)\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "wrong arg count for macro M") {
t.Fatalf("arg count: got %v, want an error", errs)
}
}
func TestConditionalsSelectOneBranch(t *testing.T) {
_, got := expand(t, `
#define MODE2
TEXT ·f(SB), NOSPLIT, $0
#ifdef MODE2
ADDQ AX, AX
#else
SUBQ AX, AX
#endif
#ifndef MODE2
SUBQ CX, CX
#else
ADDQ CX, CX
#endif
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"ADDQ CX, CX",
"RET",
)
}
func TestConditionalsHideDefinitionsAndIncludes(t *testing.T) {
// A definition inside a disabled branch must not exist, and an
// unresolvable include there must not be followed.
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
#ifdef NOTDEFINED
#define HIDEN ADDQ AX, AX
#include "nowhere.h"
#endif
HIDEN
RET
`)
wantLines(t, got, "HIDEN", "RET")
}
func TestUnclosedConditionalIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#ifdef X\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unclosed #ifdef") {
t.Fatalf("unclosed conditional: got %v, want an error", errs)
}
}
func TestUnmatchedConditionalDelimitersAreErrors(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#endif\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #endif") {
t.Fatalf("unmatched #endif: got %v, want an error", errs)
}
_, errs = ParseWithOptions("t_amd64.s", "#else\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #else") {
t.Fatalf("unmatched #else: got %v, want an error", errs)
}
}
// includeTree writes a directory of include files and returns its path.
func includeTree(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, content := range files {
path := filepath.Join(dir, name)
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
func TestIncludeSplicesAndDefinesAreShared(t *testing.T) {
dir := includeTree(t, map[string]string{
"consts.h": "#define KONST $42\n",
})
f, errs := ParseWithOptions("t_amd64.s", `
#include "consts.h"
TEXT ·f(SB), NOSPLIT, $0
MOVQ KONST, AX
RET
`, Options{Expand: true, IncludeDirs: []string{dir}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
if in.Mnemonic.Text != "MOVQ" || strings.ReplaceAll(in.Operands[0].Raw, " ", "") != "$42" {
t.Fatalf("include splicing failed: %+v", in)
}
}
func TestIncludeResolutionOrder(t *testing.T) {
// The including file's directory wins over the -I list, and the -I list
// is searched in order.
src := includeTree(t, map[string]string{
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"inc/which.h": "#define WHO ONE\n",
"first/which.h": "#define WHO TWO\n",
"second/which.h": "#define WHO THREE\n",
})
main := filepath.Join(src, "inc", "main.s")
body, err := os.ReadFile(main)
if err != nil {
t.Fatal(err)
}
// The header exists in the including file's directory and in two -I
// directories; the source-directory copy must win.
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
filepath.Join(src, "first"), filepath.Join(src, "second"),
}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
found := false
for _, d := range f.Decls {
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO ONE") {
found = true
}
}
if !found {
t.Error("the including file's directory did not win include resolution")
}
}
func TestIncludeSearchesIncludeDirsInOrder(t *testing.T) {
src := includeTree(t, map[string]string{
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"first/which.h": "#define WHO TWO\n",
"second/which.h": "#define WHO THREE\n",
})
main := filepath.Join(src, "inc", "main.s")
body, err := os.ReadFile(main)
if err != nil {
t.Fatal(err)
}
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
filepath.Join(src, "first"), filepath.Join(src, "second"),
}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
for _, d := range f.Decls {
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO THREE") {
t.Error("the second -I directory was searched before the first")
}
}
}
func TestIncludeCycleIsDetected(t *testing.T) {
src := includeTree(t, map[string]string{
"a.s": "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"b.s": "#include \"a.s\"\n",
})
_, errs := ParseWithOptions(filepath.Join(src, "a.s"), "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "include cycle") {
t.Fatalf("include cycle: got %v, want a cycle diagnostic, not a hang", errs)
}
}
func TestUnresolvableIncludeIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
Options{Expand: true, IncludeDirs: []string{t.TempDir()}})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), `#include "nothere.h"`) {
t.Fatalf("missing include: got %v, want a clear diagnostic", errs)
}
}
func TestTextflagHeaderIsNeverSpliced(t *testing.T) {
// textflag.h resolves nowhere here, yet the file must parse: the flag
// names are consumed natively and the include stays in the tree.
f, errs := ParseWithOptions("t_amd64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
hasInclude := false
for _, d := range f.Decls {
if _, ok := d.(*ast.Include); ok {
hasInclude = true
}
}
if !hasInclude {
t.Error("textflag.h include was dropped from the tree")
}
}
func TestSemicolonSplitsRawLinesToo(t *testing.T) {
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
BYTE $0x0f; BYTE $0x1f
RET
`)
wantLines(t, got, "BYTE $0x0f", "BYTE $0x1f", "RET")
}
func TestParseUnchangedWithoutExpand(t *testing.T) {
// Without Expand the preprocessor must not exist: a macro invocation
// stays an unexpanded instruction line. The ';' statement separator is
// not part of the preprocessor: the plain parse path splits on it the
// same way the expansion path does, so both spellings agree.
f, errs := Parse("t_amd64.s", `
#define TWICE ADDQ AX, AX
TEXT ·f(SB), NOSPLIT, $0
TWICE
BYTE $0x0f; BYTE $0x1f
RET
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
var mnemonics []string
for _, s := range fn.Body {
if in, ok := s.(*ast.Instr); ok {
mnemonics = append(mnemonics, in.Mnemonic.Text)
}
}
if strings.Join(mnemonics, " ") != "TWICE BYTE BYTE RET" {
t.Errorf("non-expanding parse changed: %v", mnemonics)
}
}
func TestConstantExpressionFolding(t *testing.T) {
// The shapes substituted macro bodies leave behind: parenthesised
// arithmetic in immediates and displacements, tilde complements. The
// assertions read the semantic fields; Raw keeps the operand's tokens
// in the canonicalised rendering, not the folded values.
f, errs := ParseWithOptions("t_amd64.s", `
TEXT ·f(SB), NOSPLIT, $0
RORQ $(32-7), AX
ANDQ $~63, AX
MOVQ ((2*4)+0)(SP), AX
MOVQ $((1<<3)|(1<<1)), AX
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
body := texts(f)[0].Body
ror := body[0].(*ast.Instr)
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
t.Errorf("RORQ immediate = %+v, want 25", ror.Operands[0].Imm)
}
and := body[1].(*ast.Instr)
if !and.Operands[0].Imm.HasVal || and.Operands[0].Imm.Val != -64 {
t.Errorf("ANDQ immediate = %+v, want -64", and.Operands[0].Imm)
}
mov := body[2].(*ast.Instr)
addr := mov.Operands[0].Addr
if !addr.HasOff || addr.Offset != 8 || addr.Base != "SP" {
t.Errorf("MOVQ address = %+v, want 8(SP)", addr)
}
mov2 := body[3].(*ast.Instr)
if !mov2.Operands[0].Imm.HasVal || mov2.Operands[0].Imm.Val != 10 {
t.Errorf("MOVQ immediate = %+v, want 10", mov2.Operands[0].Imm)
}
}
func TestConstantExpressionFoldsWithoutExpand(t *testing.T) {
// Folding is a parser capability, not a preprocessing one: a
// hand-written $(32-7) folds the same way with expansion off.
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tRORQ $(32-7), AX\n\tRET\n", Options{})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
if !in.Operands[0].Imm.HasVal || in.Operands[0].Imm.Val != 25 {
t.Errorf("Imm = %+v, want 25", in.Operands[0].Imm)
}
}
func TestParameterWithSelectorSubstitutes(t *testing.T) {
// The lexer folds A.S4 into one identifier token, so a parameter used
// with an element or lane selector never matched the whole-token
// substitution; the toolchain's lexer splits on the period and its
// substitution sees the name alone. Several parameters carry selectors
// in one body here, which is the chacha8_arm64.s QR shape in miniature.
_, got := expand(t, `
#define QR(A, B, C, D) VADD A.S4, B.S4, C.S4; VEOR D.B16, A.B16, D.B16
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V2.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"RET",
)
}
func TestSelectorWithCompoundArgumentPastesTextually(t *testing.T) {
// An argument that is itself one compound identifier pastes verbatim:
// A.S4 over V0.B16 spells V0.B16.S4, the operand the toolchain's
// split-then-substitute leaves behind.
_, got := expand(t, `
#define M(A) VADD A.S4, A.S4, A.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0.B16)
RET
`)
wantLines(t, got, "VADD V0.B16.S4, V0.B16.S4, V0.B16.S4", "RET")
}
func TestSelectorAlongsideBareParameter(t *testing.T) {
// A body may use the parameter bare and suffixed, and the argument may
// itself end in a selector; neither disturbs the other.
_, got := expand(t, `
#define M(A) VADD A, A.S4, A
TEXT ·f(SB), NOSPLIT, $0
M(V0)
M(V1.B16)
RET
`)
wantLines(t, got,
"VADD V0, V0.S4, V0",
"VADD V1.B16, V1.B16.S4, V1.B16",
"RET",
)
}
func TestSelectorKeepsNonParameterPrefixes(t *testing.T) {
// The prefix before the period must be the whole parameter name:
// AB.S4 never reaches a parameter A.
_, got := expand(t, `
#define M(A) VADD AB.S4, A.S4, AB.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0)
RET
`)
wantLines(t, got, "VADD AB.S4, V0.S4, AB.S4", "RET")
}
func TestSelectorExpandsMacroValuedArgument(t *testing.T) {
// gcm_arm64.s invokes mulRound(B1) where B1 is itself an object macro:
// the paste stays rescannable, so B1.D1 still expands to V1.D1 the way
// the toolchain's rescan of substituted tokens does.
_, got := expand(t, `
#define B1 V1
#define mulRound(X) VPMULL X.D1, T1.D1, T3.Q1
TEXT ·f(SB), NOSPLIT, $0
mulRound(B1)
RET
`)
wantLines(t, got, "VPMULL V1.D1, T1.D1, T3.Q1", "RET")
}
func TestObjectMacroBehindSelectorExpands(t *testing.T) {
// Ordinary code writes ACC0.B16 where ACC0 is an object macro; the
// toolchain expands the alias because its lexer reads the selector as
// its own token, and the lookup here must reach the macro through the
// compound spelling the same way.
_, got := expand(t, `
#define ACC0 V8
TEXT ·f(SB), NOSPLIT, $0
VEOR ACC0.B16, ACC0.B16, ACC0.B16
RET
`)
wantLines(t, got, "VEOR V8.B16, V8.B16, V8.B16", "RET")
}
func TestChacha8QRMacroExpands(t *testing.T) {
// The real QR round of chacha8_arm64.s end to end: every parameter
// carries a selector somewhere, and the round is sixteen instructions.
_, got := expand(t, `
#define QR(A, B, C, D) \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VREV32 D.H8, D.H8; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $12, V30.S4, B.S4; VSRI $20, V30.S4, B.S4; \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VTBL V31.B16, [D.B16], D.B16; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $7, V30.S4, B.S4; VSRI $25, V30.S4, B.S4
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VREV32 V3.H8, V3.H8",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $12, V30.S4, V1.S4",
"VSRI $20, V30.S4, V1.S4",
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VTBL V31.B16, [V3.B16], V3.B16",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $7, V30.S4, V1.S4",
"VSRI $25, V30.S4, V1.S4",
"RET",
)
}
func TestNotAnExpressionFallsBack(t *testing.T) {
// Symbol immediates and floats must keep their ordinary parse.
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ $1.5, AX\n\tMOVQ $·sym(SB), AX\n\tRET\n", Options{})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
mov1 := fn.Body[0].(*ast.Instr)
if mov1.Operands[0].Imm.HasVal || mov1.Operands[0].Imm.Float != "1.5" {
t.Errorf("float immediate parsed as %+v", mov1.Operands[0].Imm)
}
mov2 := fn.Body[1].(*ast.Instr)
if mov2.Operands[0].Imm.Sym == nil {
t.Errorf("symbol immediate parsed as %+v", mov2.Operands[0].Imm)
}
}
+44
View File
@@ -0,0 +1,44 @@
// Differential kernel: the _dbar (acquire/release) atomic exchange
// variants against the Go toolchain's loong64enc1.s rows.
#include "textflag.h"
TEXT ·AMXORDBW(SB), NOSPLIT, $0
AMXORDBW R14, (R13), R12
RET
TEXT ·AMXORDBV(SB), NOSPLIT, $0
AMXORDBV R14, (R13), R12
RET
TEXT ·AMMAXDBW(SB), NOSPLIT, $0
AMMAXDBW R14, (R13), R12
RET
TEXT ·AMMAXDBV(SB), NOSPLIT, $0
AMMAXDBV R14, (R13), R12
RET
TEXT ·AMMINDBW(SB), NOSPLIT, $0
AMMINDBW R14, (R13), R12
RET
TEXT ·AMMINDBV(SB), NOSPLIT, $0
AMMINDBV R14, (R13), R12
RET
TEXT ·AMMAXDBWU(SB), NOSPLIT, $0
AMMAXDBWU R14, (R13), R12
RET
TEXT ·AMMAXDBVU(SB), NOSPLIT, $0
AMMAXDBVU R14, (R13), R12
RET
TEXT ·AMMINDBWU(SB), NOSPLIT, $0
AMMINDBWU R14, (R13), R12
RET
TEXT ·AMMINDBVU(SB), NOSPLIT, $0
AMMINDBVU R14, (R13), R12
RET
+191
View File
@@ -0,0 +1,191 @@
// The AVX-512 families behind the avx512enc gap: AES round ops, integer
// VNNI and bit algorithms, word shifts and permutes with an immediate or a
// register count, lane broadcasts and extracts, gather and scatter prefetch
// hints, opmask broadcasts, the high/low half moves and the non-temporal
// stores. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func avx512int(p *byte, n int) uint64
TEXT ·avx512int(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
MOVQ n+16(FP), CX
// AES rounds through the EVEX spellings, masks included.
VAESENC Z20, Z21, Z22
VAESENCLAST Z23, Z24, Z25
VAESDEC (SI), Z26, Z27
VAESDECLAST Z28, Z29, Z30
// Integer VNNI and the bit algorithm group.
VPDPBUSD Z1, Z2, K2, Z3
VPDPBUSDS Z4, Z5, K2, Z6
VPDPWSSD Z7, Z8, Z9
VPDPWSSDS Z10, Z11, K2, Z12
VPOPCNTW Z12, K3, Z13
VPOPCNTB Z14, Z15
VGF2P8MULB Z16, Z17, K4, Z18
VGF2P8AFFINEQB $7, Z18, Z19, K5, Z20
// Byte/word arithmetic with saturation and masks.
VPADDSB Z1, Z2, K1, Z3
VPADDUSW Z3, Z4, K1, Z5
VPSUBSW Z5, Z6, K1, Z7
VPSUBUSB Z7, Z8, K1, Z9
VPSADBW Z9, Z10, Z11
VPMULHRSW Z11, Z12, Z13
VPMULHW Z13, Z14, Z15
VPUNPCKLBW Z15, Z16, K2, Z17
VPUNPCKHBW Z17, Z18, K2, Z19
VPUNPCKLWD Z19, Z20, K2, Z21
VPUNPCKHWD Z21, Z22, K2, Z23
VPCMPEQB Z23, Z24, K2, K3
VPCMPGTW Z25, Z26, K2, K3
VPCMPEQQ Z27, Z28, K2
VPMULTISHIFTQB Z29, Z30, K3, Z31
VDBPSADBW $3, Z1, Z2, K3, Z3
MOVQ CX, ret+16(FP)
RET
// func avx512perm(p *byte) uint64
TEXT ·avx512perm(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Permutations: immediate and register counts, ternary logic.
VALIGNQ $3, Z1, Z2, K1, Z3
VPERMT2B Z3, Z4, K1, Z5
VPERMT2W Z5, Z6, K1, Z7
VPERMT2PS Z7, Z8, K1, Z9
VPERMI2W Z9, Z10, K1, Z11
VPERMI2PS Z11, Z12, K1, Z13
VPERMI2PD Z13, Z14, K1, Z15
VPERMB Z15, Z16, K1, Z17
VPERMW Z17, Z18, K1, Z19
VPERMPS Z19, Z20, Z21
VPERMD Z20, Z21, Z22
VPERMQ $1, Z1, K2, Z2
VPERMQ Z3, Z4, K2, Z5
VPERMPD $1, Z5, K2, Z6
VPERMPD Z7, Z8, K2, Z9
VPERMILPS $5, Z9, K2, Z10
VPERMILPS Z11, Z12, K2, Z13
VPERMILPD $1, Z13, K2, Z14
VPERMILPD Z15, Z16, K2, Z17
VPTERNLOGD $6, Z17, Z18, K2, Z19
VPTERNLOGQ $9, Z19, Z20, K2, Z21
// Lane shuffle and blend families.
VSHUFPD $1, Z1, Z2, K1, Z3
VSHUFPS $2, Z4, Z5, K1, Z6
VBLENDMPD Z7, Z8, K1, Z9
VBLENDMPS Z9, Z10, K1, Z11
VPBLENDMB Z11, Z12, K1, Z13
VPBLENDMW Z13, Z14, K1, Z15
VPBLENDMD Z15, Z16, K1, Z17
VPBLENDMQ Z17, Z18, K1, Z19
// Conflicts and leading zero counts.
VPCONFLICTD Z1, K1, Z2
VPCONFLICTQ Z3, K1, Z4
VPLZCNTD Z5, K1, Z6
VPLZCNTQ Z7, K1, Z8
// Compress and expand, byte and word widths.
VPCOMPRESSB Z1, K1, (SI)
VPCOMPRESSW Z2, K1, (SI)
VPEXPANDB (SI), K1, Z3
VPEXPANDW (SI), K1, Z4
MOVQ SI, ret+8(FP)
RET
// func avx512shift(p *byte) uint64
TEXT ·avx512shift(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Variable shifts and shuffles with masks.
VPSLLVW Z1, Z2, K1, Z3
VPSRLVW Z3, Z4, K1, Z5
VPSRAVW Z5, Z6, K1, Z7
VPSHLDVW Z7, Z8, K1, Z9
VPSHRDVW Z9, Z10, K1, Z11
VPSHLDVD Z11, Z12, K1, Z13
VPSHLDVQ Z13, Z14, K1, Z15
VPSHRDVD Z15, Z16, K1, Z17
VPSHRDVQ Z17, Z18, K1, Z19
// Immediate shifts, the word/byte-quad widths and masks.
VPSLLW $3, Z1, K2, Z2
VPSRLW $5, Z3, K2, Z4
VPSRAW $7, Z5, K2, Z6
VPSLLDQ $9, Z7, Z8
VPSRLDQ $11, Z9, Z10
// Register-count shifts and their memory-count forms.
VPSLLD X1, Z2, K1, Z3
VPSRLD 16(SI), Z4, K1, Z5
VPSLLQ X6, Z7, K1, Z8
VPSRLQ X9, Z10, K1, Z11
VPSLLW X12, Z13, K1, Z14
VPSRAW X15, Z16, K1, Z17
VPSRAQ $13, Z12, K1, Z13
VPSRAD X14, Z15, K1, Z16
// Lane shuffles in and out.
VPSHLDW $2, Z1, Z2, K1, Z3
VPSHLDQ $4, Z3, Z4, K1, Z5
VPSHRDW $6, Z5, Z6, K1, Z7
VPSHRDQ $8, Z7, Z8, K1, Z9
VPSHUFBITQMB Z9, Z10, K3
VPTESTMB Z11, Z12, K4
VPTESTNMQ Z13, Z14, K5
MOVQ SI, ret+8(FP)
RET
// func avx512float(x float64) float64
TEXT ·avx512float(SB), NOSPLIT, $0-16
// Square roots, compares and the EXP2/RCP28 helpers.
MOVQ x+0(FP), AX
VSQRTPD Z1, K1, Z2
VSQRTPS Z3, K1, Z4
VSQRTSD X1, X2, K1, X3
VSQRTSS X3, X4, X5
VCOMISD X5, X6
VUCOMISS X7, X8
VEXP2PD Z5, K1, Z6
VRCP28PD Z7, K1, Z8
VRCP28SD X9, X8, K1, X10
VRSQRT28PS Z11, K1, Z12
VRSQRT28SS X11, X10, K1, X12
VCVTSD2SS X1, X2, X3
VCVTSS2SD X3, X2, K1, X4
VFMADD132PD Z1, Z2, K1, Z3
VFMADD231SD X1, X2, K1, X3
VFMSUBADD213PS Z3, Z4, K1, Z5
VFNMSUB231PD Z5, Z6, K1, Z7
// Broadcasts and masked moves.
VBROADCASTF32X2 X1, K1, Z2
VBROADCASTI64X2 (SI), K1, Z3
VMOVUPS Z1, K2, Z3
VMOVSD X14, X5, K3, X22
VMOVSS X18, X3, K2, X25
VMOVHPS (SI), X18, X19
VMOVHPS X20, 8(SI)
VMOVLHPS X16, X5, X17
VMOVNTDQ Z7, (SI)
VMOVNTDQA 64(SI), Z8
VMOVNTPD Z9, (SI)
MOVQ SI, ret+8(FP)
RET
// func avx512mask(p *byte) uint64
TEXT ·avx512mask(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Omask broadcasts and the K register logic.
VPBROADCASTMB2Q K1, Z2
VPBROADCASTMW2D K3, Z4
KUNPCKWD K6, K4, K1
KADDB K2, K3, K5
KORW K1, K2, K7
// Gather and scatter prefetch hints.
VGATHERPF0DPD K5, (SI)(Y29*8)
VSCATTERPF1DPS K2, (SI)(Z28*4)
// Masked gathers ride the EVEX spelling; the data length wins L'L.
VGATHERDPD (SI)(X10*4), K7, Y22
VPSCATTERDQ Y6, K2, (SI)(X4*1)
// Lane extracts to general registers.
VPEXTRB $3, X1, AX
VPEXTRD $1, X2, DI
VPINSRQ $1, SI, X3, X4
VEXTRACTI32X4 $1, Z1, X5
VINSERTI64X2 $1, X6, Z7, K2, Z8
MOVQ SI, ret+8(FP)
RET
+32
View File
@@ -0,0 +1,32 @@
// The runtime bookkeeping statements: FUNCDATA and PCDATA contribute no
// text bytes on any architecture, and amd64 now matches. They sit between
// real instructions here, with plain, static and offset symbol references
// on the FUNCDATA lines, so the byte counts prove the zero contribution.
#include "textflag.h"
// func bookkeep(x int64) int64
TEXT ·bookkeep(SB), NOSPLIT, $0-16
PCDATA $0, $-1
MOVQ x+0(FP), AX
PCDATA $1, $-2
FUNCDATA $0, args_stackmap(SB)
ADDQ $1, AX
FUNCDATA $5, arginfo0(SB)
PCDATA $1, $3
MOVQ AX, ret+8(FP)
FUNCDATA $1, externalfuncdata(SB)
PCDATA $0, $0
RET
// func bookkeepstatic() int64
TEXT ·bookkeepstatic(SB), NOSPLIT, $0-8
// A static symbol and a defined data symbol as the funcdata target.
// (A symbol+offset target the toolchain itself refuses.)
FUNCDATA $2, fdtable<>(SB)
FUNCDATA $3, undefsym(SB)
MOVQ $7, AX
MOVQ AX, ret+0(FP)
RET
GLOBL fdtable<>(SB), NOPTR, $16
+31
View File
@@ -0,0 +1,31 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 bookkeeping statements: the funcdata.h
// pseudo-directives (GO_ARGS, NO_LOCAL_POINTERS, FUNCDATA, PCDATA) contribute
// no instruction bytes, and every function is byte-compared against
// go tool asm.
#include "textflag.h"
#include "funcdata.h"
// func bookkeep()
TEXT ·bookkeep(SB), NOSPLIT, $8-0
GO_ARGS
FUNCDATA $3, inline_tree(SB)
PCDATA $1, $2
MOVD R1, 0(RSP)
RET
// func bookkeepNoLocals()
TEXT ·bookkeepNoLocals(SB), NOSPLIT, $16-0
NO_LOCAL_POINTERS
PCDATA $0, $0
PCDATA $1, $1
MOVD R2, 8(RSP)
RET
// func bookkeepPlain()
TEXT ·bookkeepPlain(SB), NOSPLIT, $0-0
MOVD R3, R4
RET
+40
View File
@@ -0,0 +1,40 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the riscv64 bookkeeping statements and the
// slot-relative branches: FUNCDATA and PCDATA (the expanded forms of the
// funcdata.h macros, contributing no bytes), UNDEF (the toolchain's ebreak),
// and the JMP N(PC) slot jumps including the self-loop and the backward form.
#include "textflag.h"
TEXT ·bookkeep(SB), NOSPLIT, $8-8
FUNCDATA $1, marks<>(SB)
PCDATA $1, $-1
MOV ZERO, ret+0(FP)
PCDATA $1, $1
UNDEF
MOV $1, X10
RET
TEXT ·slots(SB), NOSPLIT, $0-0
MOV $1, X10
JMP 2(PC)
MOV $64, X11
MOV $128, X12
MOV $2, X11
MOV $3, X12
BEQ X10, X11, skip
JMP -2(PC)
skip:
JMP 0(PC)
TEXT ·marksreader(SB), NOSPLIT, $0-8
MOV $marks<>(SB), X10
MOV (X10), X11
MOV X11, ret+0(FP)
RET
GLOBL marks<>(SB), RODATA, $8
DATA marks<>+0(SB)/8, $1234605616436508552
+21
View File
@@ -0,0 +1,21 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the loong64 two-operand BEQ/BNE spellings the
// msan trampolines use: BEQ Rj, target compares against R0 (the beqz form).
#include "textflag.h"
TEXT ·branch2(SB), NOSPLIT, $0-8
MOVV arg+0(FP), R4
BEQ R4, zero
ADDV $1, R4, R4
zero:
MOVV $16, R5
BNE R4, done
ADDV $2, R4, R4
done:
MOVV R4, ret+0(FP)
RET
File diff suppressed because it is too large Load Diff
+48
View File
@@ -0,0 +1,48 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Carry arithmetic, logical shifts, register aliases with element selectors
// and the ADC/SBC immediate spellings: the shapes nat_arm64.s, p256 and
// gcm_arm64.s exercise. Byte-for-byte against go tool asm.
#include "textflag.h"
#define acc0 V8
#define acc1 V9
#define const0 R15
#define POLY V15
// carry pins the ADC/SBC family: the $0 spellings in two and three
// operands, and the register-carry forms.
TEXT ·carry(SB), NOSPLIT, $0-0
ADC $0, R20
ADC $0, R20, R4
SBCS $0, R4
SBCS $0, R4, R12
SBCS R15, R4, R12
SBC $0, R1
ADCSW $0, R2, R3
RET
// shift pins the shifted-register forms including ROR, which only the
// logical family accepts.
TEXT ·shift(SB), NOSPLIT, $0-0
ANDW R9@>7, R19, R26
AND R1@>33, R2, R3
ADD R1<<11, R2, R3
SUB R1->33, R2
ORR R5<<2, R6, R7
RET
// vecalias pins the vector aliases with element selectors and the
// structure loads with aliased members.
TEXT ·vecalias(SB), NOSPLIT, $0-0
MOVD $0xC2, R1
VMOV R1, POLY.D[0]
VMOV R0, POLY.D[1]
VEOR POLY.B16, POLY.B16, POLY.B16
VLD1 (R0), [acc0.B16]
VLD1.P (R0), [acc0.B16, acc1.B16]
VST1 [acc0.B16, acc1.B16], (R1)
VST1.P [acc0.B16, acc1.B16], 32(R1)
RET
+27
View File
@@ -0,0 +1,27 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the loong64 DATA value forms the runtime's exp and
// asm files use: floating-point initialisers stored as IEEE-754 bits and
// string initialisers zero-padded within their declared width.
#include "textflag.h"
TEXT ·floatbits(SB), NOSPLIT, $0-8
MOVV $floats<>(SB), R12
MOVD 8(R12), F0
MOVD F0, ret+0(FP)
RET
TEXT ·stringhead(SB), NOSPLIT, $0-8
MOVV $msg<>(SB), R12
MOVV (R12), R13
MOVV R13, ret+0(FP)
RET
GLOBL floats<>(SB), RODATA, $16
DATA floats<>+0(SB)/8, $0.0
DATA floats<>+8(SB)/8, $0.5
GLOBL msg<>(SB), RODATA, $20
DATA msg<>+0(SB)/20, $"call frame too large"
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVQ handlers+8(SB), AX
MOVQ AX, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVQ $42, AX
MOVQ AX, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $·handler+5(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVD handlers+8(SB), R4
MOVD R4, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVZ $42, R4
MOVD R4, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $extentry(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVQ a+0(FP), DI
MOVQ b+8(FP), SI
CALL internal∕runtime∕atomic·Xchg(SB)
MOVQ AX, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVQ AX, ret+0(FP)
RET
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVD a+0(FP), R4
MOVD b+8(FP), R5
CALL internal∕runtime∕atomic·Xchg(SB)
MOVD R4, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVD R0, ret+0(FP)
RET
+33
View File
@@ -0,0 +1,33 @@
// The three-operand SHL/SHR forms, which go tool asm encodes as SHLD/SHRD:
// immediate and CL (or its CX spelling) counts at the Q and W widths, next
// to the two-operand CX-count spelling GOROOT's bignum kernels use. Every
// result is folded back so no instruction is dead.
#include "textflag.h"
// func dblshift(x, y uint64) uint64
TEXT ·dblshift(SB), NOSPLIT, $0-24
MOVQ x+0(FP), SI
MOVQ y+8(FP), DI
MOVQ $12, CX
SHLQ $13, SI, DI
SHRQ $7, DI, SI
SHLQ CX, SI, DI
SHRQ CX, DI, SI
SHLQ CX, SI
SHLQ $9, DI
SHLW $1, SI, DI
SHRW $3, DI, SI
XORQ DI, SI
MOVQ SI, ret+16(FP)
RET
// func dblshift32(a, b uint32) uint32
TEXT ·dblshift32(SB), NOSPLIT, $0-12
MOVL a+0(FP), SI
MOVL b+4(FP), DI
SHLL $5, SI, DI
SHRL $2, DI, SI
XORL SI, DI
MOVL DI, ret+8(FP)
RET

Some files were not shown because too many files have changed in this diff Show More