Compare commits

...
72 Commits
Author SHA1 Message Date
petrbalvin e836d6150d docs: changelog entry for the loong64 JIT enablement
Test / test (push) Successful in 2m7s
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin 8a51b060da feat(cmd): enable loong64 JIT execution, all trampolines qemu-validated
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin d3d47db727 test(verify): seed the arm64 ABI kernel arguments
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin 0758556b7d docs: changelog entries for the parity round and corpus number
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin ddb8440340 fix(cmd): padding-aware ground-truth comparison
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin f15ff66fb1 fix(riscv64): accept the g spelling of the goroutine register
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin 187e4856d3 feat(amd64): encode the mixed-width extend family and PMOVMSKB
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin d315a998ce fix(arm64): store-exclusive operand order and large-frame parity
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin a6f3828c02 docs: changelog entries for the review fixes
Test / test (push) Successful in 2m4s
Assisted-by: GLM 5.3
2026-09-19 23:49:27 +02:00
petrbalvin e3b35bb817 style(testdata): canonical gasm formatting for the verify kernels
Assisted-by: GLM 5.3
2026-09-19 23:49:27 +02:00
petrbalvin eb0a89e58d ci(release): state the version contract inline
Assisted-by: GLM 5.3
2026-09-19 23:49:27 +02:00
petrbalvin dd32d9e66e chore(justfile): one-line install-man comment and long flag forms
Assisted-by: GLM 5.3
2026-09-19 23:49:27 +02:00
petrbalvin 3a73acb20a docs: drop process labels and refresh the architecture and manual pages
Assisted-by: GLM 5.3
2026-09-19 23:49:27 +02:00
petrbalvin 7604a9443f fix(cmd): usage exit codes, asm output file and cross-arch ground truth
Assisted-by: GLM 5.3
2026-09-19 23:49:19 +02:00
petrbalvin b3908fc43d fix(lsp): parse-error survival, symbol ranges and UTF-16 positions
Assisted-by: GLM 5.3
2026-09-19 23:49:19 +02:00
petrbalvin eb8b0cd316 fix(lint): trailing-label CFG guard and the goroutine alias
Assisted-by: GLM 5.3
2026-09-19 23:49:19 +02:00
petrbalvin a8bfd54ed2 fix(debug): hardware watchpoints, signal stops and breakpoint restore
Assisted-by: GLM 5.3
2026-09-19 23:49:19 +02:00
petrbalvin 375182ef1f fix(verify): arm64 stack save, adaptive canary and host gating
Assisted-by: GLM 5.3
2026-09-19 23:49:19 +02:00
petrbalvin 87b1081c53 fix(goobj): external package and symbol indices and arm64 pair relocations
Assisted-by: GLM 5.3
2026-09-19 23:49:13 +02:00
petrbalvin f3c8510a58 fix(elf): relocation records, DWARF tables and per-architecture frame data
Assisted-by: GLM 5.3
2026-09-19 23:49:13 +02:00
petrbalvin ebdf14939f fix(loong64): FP immediates through R30 and unsigned branch forms
Assisted-by: GLM 5.3
2026-09-19 23:49:13 +02:00
petrbalvin 79a2c16bac fix(riscv64): compressed store offsets, FENCE and branch range checks
Assisted-by: GLM 5.3
2026-09-19 23:49:13 +02:00
petrbalvin 401386956c fix(arm64): encode shifts, divides and multiplies and align sizes with emission
Assisted-by: GLM 5.3
2026-09-19 23:49:07 +02:00
petrbalvin 4258131a3a fix(amd64): correct guard displacements, frameless FP offsets and immediate ranges
Assisted-by: GLM 5.3
2026-09-19 23:49:07 +02:00
petrbalvin 94e09e8070 fix(format): preserve flag separators and normalise CRLF input
Assisted-by: GLM 5.3
2026-09-19 23:48:47 +02:00
petrbalvin 7aefe6a42d fix(parser): parse ABI markers and keep TEXT decls usable on errors
Assisted-by: GLM 5.3
2026-09-19 23:48:47 +02:00
petrbalvin ac1c05c793 fix(lexer): tokenise the flag separator and handle NUL and invalid UTF-8
Assisted-by: GLM 5.3
2026-09-19 23:48:47 +02:00
petrbalvin 93c47a312a feat(docs): man pages for gasm and every command, guarded against CLI drift
Test / test (push) Successful in 2m4s
Assisted-by: GLM 5.3 Flash
2026-09-19 21:18:43 +02:00
petrbalvin 708d0a0a5e docs: trim the changelog entries to user-visible deltas
Assisted-by: GLM 5.3 Flash
2026-09-19 20:54:56 +02:00
petrbalvin 3c8f7cb411 test(format): pin the fuzz-found crashers as regression seeds
Assisted-by: GLM 5.3 Flash
2026-09-19 20:48:51 +02:00
petrbalvin 7c5b7a1419 docs: add the changelog entries and the corpus number to the readme
Assisted-by: GLM 5.3 Flash
2026-09-19 20:48:51 +02:00
petrbalvin bc3f448738 feat(format): fuzz targets for the parser and formatter
Assisted-by: GLM 5.3 Flash
2026-09-19 20:41:43 +02:00
petrbalvin f37f183577 feat(riscv64): GOROOT instruction shapes, DATA order and offset expressions
Assisted-by: GLM 5.3 Flash
2026-09-19 19:58:43 +02:00
petrbalvin 1e77e58250 feat(gasm): audit a .s corpus with audit-instructions --corpus
Assisted-by: GLM 5.3 Flash
2026-09-19 19:27:30 +02:00
petrbalvin 1d0969ed64 feat(gasm): select the asm and diff architecture with -GOARCH
Assisted-by: GLM 5.3 Flash
2026-09-19 19:20:47 +02:00
petrbalvin 23c001be51 feat(asm): encode indirect JMP and CALL on all four architectures
Assisted-by: GLM 5.3 Flash
2026-09-19 19:17:07 +02:00
petrbalvin 96e81cc98d docs: add the Plan 9 assembly case and real-use note to the README
Test / test (push) Successful in 2m6s
2026-09-19 18:06:18 +02:00
petrbalvin c834d98210 docs: bring the document set into the standard shape
Test / test (push) Successful in 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 03d6d4da54 style: put the repository assembly in gasm fmt canonical form
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 0b42ce7952 style: use one spelling for colour across the CLI
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 288a64ccd2 ci: align the pipelines with the hand-written templates
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 5fddfa704b build: declare the exact toolchain and the canonical recipes
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin a2bb5eeb4e chore: drop the stale comment from the ignore list
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:14 +02:00
petrbalvin 48449b7a7f build: declare the go1.27.1 toolchain
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 3de043c494 docs: add SECURITY.md and record the round in the CHANGELOG
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 0078f7be5c style: purge em dashes from the produced text
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 6a7317d141 chore: trim the ignore list to the convention
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin d08523caa5 docs: move the recipe and version descriptions with the behaviour
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 20e4b8d9c4 ci: align the pipelines with the hand-written templates
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 61f4247cef refactor(gasm): report the toolchain-recorded version
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 049872ddff build: restore the canonical justfile recipe set
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 3669f64ff6 build: install the gasm binary into the user-local bin directory
Test / vet (push) Successful in 46s
Test / test (push) Successful in 2m44s
Test / build (push) Successful in 42s
2026-09-14 23:41:21 +02:00
petrbalvin fff9f75595 chore: prepare release v0.33.0
Release / build (amd64, linux) (push) Successful in 49s
Release / build (arm64, linux) (push) Successful in 43s
Release / build (loong64, linux) (push) Successful in 46s
Release / build (riscv64, linux) (push) Successful in 45s
Test / vet (push) Successful in 47s
Release / release (push) Successful in 18s
Test / test (push) Successful in 2m39s
Test / build (push) Successful in 43s
2026-09-14 23:36:19 +02:00
petrbalvin 40476546df fix(asm): close the oracle parity gaps in frame addressing and calls 2026-09-14 23:25:14 +02:00
petrbalvin 70218e84ba feat(asm): emit the loong64 stack-split guard for big frames 2026-09-14 22:38:58 +02:00
petrbalvin db50b98179 feat(asm): emit the loong64 stack-split guard for small and medium frames 2026-09-14 21:21:58 +02:00
petrbalvin 2e2c0b82a0 feat(asm): emit the riscv64 stack-split guard and fix large-frame addressing 2026-09-14 21:09:13 +02:00
petrbalvin 8dc1e98ca1 feat(asm): emit the arm64 stack-split guard and morestack block 2026-09-14 20:49:03 +02:00
petrbalvin 1d8e68c574 feat(asm): emit the amd64 stack-split guard and morestack block 2026-09-14 20:35:55 +02:00
petrbalvin 89fa6ea15e feat(lsp): resolve definition and references across open documents 2026-09-14 18:50:55 +02:00
petrbalvin edc20ffa97 feat(cmd): add gasm dis and share the decoder with the debugger 2026-09-14 18:47:08 +02:00
petrbalvin 50db6615b2 feat(cmd): add gofmt-style -l and -d modes to gasm fmt 2026-09-14 18:47:08 +02:00
petrbalvin 95f1d6f083 style: replace em dashes in the scaffold comments 2026-09-14 18:22:25 +02:00
petrbalvin f43e791e5a chore: add .qwen to the gitignore metadata block 2026-09-14 18:22:18 +02:00
petrbalvin 1691c81095 style: replace em and en dashes across sources 2026-09-14 18:22:18 +02:00
petrbalvin 2db563be07 refactor(cmd): consolidate cross-arch verify and drop dead code 2026-09-14 18:22:00 +02:00
petrbalvin 4f190ee1a2 refactor(debug): move watchpoint slot state into the session 2026-09-14 18:22:00 +02:00
petrbalvin 909f874797 fix(lsp): recover from handler panics and decode client uris 2026-09-14 18:22:00 +02:00
petrbalvin e307bf830f fix(lint): guard unnamed TEXT and refresh the textflag table 2026-09-14 18:22:00 +02:00
petrbalvin 953c258d6a fix(asm): make arm64 and loong64 relocations match the toolchain 2026-09-14 18:22:00 +02:00
petrbalvin c6f0286732 fix(asm): encode amd64 frame adjustments above 127 bytes with imm32 2026-09-14 18:22:00 +02:00
petrbalvin f5fc22d390 fix(parser): reject malformed TEXT frames and parse signed frame sizes 2026-09-14 18:22:00 +02:00
235 changed files with 15641 additions and 3669 deletions
+37
View File
@@ -0,0 +1,37 @@
# Race, Go. Dispatched by hand, and never a gate on a push or a tag: the release tag is
# cut only after `just gates` has already raced the tree, so this workflow is the
# explicit second opinion, not a step of the release.
#
# The race detector roughly doubles both time and memory, which the shared runner box
# cannot afford on every push. Locally it belongs to `just gates`, which runs it once per
# task; here it is a decision rather than a routine.
#
# Every step is one command, so the step that fails is the gate that failed.
name: Race
on:
workflow_dispatch:
env:
# One core: parallelism buys no speed here and costs memory the box does not have.
GOFLAGS: -p=1
GOMAXPROCS: "2"
jobs:
race:
runs-on: fedora
timeout-minutes: 20
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version-file: go.mod
cache: true
- name: Install gcc
# The race detector needs cgo and the runner image carries no C compiler.
run: dnf install -y gcc
- name: Race
run: go test -race -count=1 -timeout 10m ./...
+255 -86
View File
@@ -1,73 +1,198 @@
# Release — gasm binaries. Runs on version tags (v0.28.0) pushed to main. # Release, Go binaries. Runs on version tags (v1.2.3) pushed to main.
#
# The module sits at the repository root: the toolchain records a version only for a root
# module, measured on go1.27.1, so a build of a module in a subdirectory reports (devel)
# even at its own <module>/vX.Y.Z tag and this workflow's smoke test can never pass for
# it. A Go repository is one module at the root.
#
# The version contract these steps implement: nothing is injected. The toolchain records
# the tag into the binary's build information, so the build simply has to happen at the
# tag, which the trigger guarantees.
#
# The gates run in their own job, once, before the matrix, minus the race detector: race
# never runs on a push path or a tag, and the local gate raced this tree before the tag
# was cut. Putting the gates inside the matrix would run the whole suite once per target
# on the box that also hosts the forge. Each job validates the tag for itself rather than
# passing a value between jobs, so no workflow feature has to be trusted for the version
# to reach the file name.
name: Release name: Release
on: on:
push: push:
tags: ["v*"] tags: ["v*"]
env:
# The box is shared with the forge, so parallelism is bounded on purpose. The gates job
# needs it most; the build jobs inherit it for their parallel compilation.
GOFLAGS: -p=1
GOMAXPROCS: "2"
jobs: jobs:
gates:
runs-on: fedora
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version-file: go.mod
cache: true
- name: Install Perl
# Perl for the steps below. The install is a no-op where the package
# is already present.
run: dnf install -y perl
- name: Validate the tag
env:
VERSION: ${{ gitea.ref_name }}
run: |
perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ m{^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$}
or die qq{ERROR: expected a semver tag like v1.2.3, got: $v\n};
print qq{tag $v\n};
'
- name: Build
run: go build ./...
- name: Format
run: |
perl -e '
open(my $g, q{-|}, q{gofmt}, q{-l}, q{.}) or die qq{gofmt: $!};
my @bad = <$g>;
close($g);
print @bad;
exit(@bad ? 1 : 0);
'
- name: Vet
run: go vet ./...
- name: Modernise
run: go fix -diff ./...
- name: Tests
# The same command as in test.yml, so the floor is the same number everywhere.
run: go test -count=1 -timeout 10m -coverprofile=coverage.out ./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... ./lint/... ./lsp/... ./parser/... ./token/... ./verify/...
- name: Coverage floor
run: |
perl -e '
open(my $c, q{-|}, q{go}, q{tool}, q{cover}, q{-func=coverage.out}) or die qq{cover: $!};
my $total;
while (my $l = <$c>) { $total = $1 if $l =~ m{^total:\s+\S+\s+([0-9.]+)%} }
close($c);
die qq{no total line in coverage.out\n} unless defined $total;
printf qq{Total coverage: %s%%\n}, $total;
exit($total < 80 ? 1 : 0);
'
build: build:
runs-on: fedora runs-on: fedora
timeout-minutes: 25
needs: gates
strategy: strategy:
fail-fast: false fail-fast: false
matrix: matrix:
# Portable targets: amd64, arm64, loong64 and riscv64 on Linux, at the toolchain
# default level. No 32-bit, no wasm, no macOS, no Windows. FreeBSD stays out until
# verify/jit.go ports off syscall.Mprotect: the Go syscall package defines no
# Mprotect for freebsd, and verify/jit.go:50 calls it to drop the write bit from
# the JIT mapping, so every freebsd target fails to build with "undefined:
# syscall.Mprotect" (verified for amd64, arm64 and riscv64 on go1.27.1).
include: include:
- goos: linux - goos: linux
goarch: amd64 goarch: amd64
- goos: linux - goos: linux
goarch: arm64 goarch: arm64
- goos: linux
goarch: riscv64
- goos: linux - goos: linux
goarch: loong64 goarch: loong64
- goos: linux
goarch: riscv64
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
- uses: actions/setup-go@v6 - uses: actions/setup-go@v6
with: with:
go-version: "1.27" go-version-file: go.mod
cache: true
- name: Download dependencies - name: Install Perl
run: go mod download run: dnf install -y perl
- name: Validate tag and build - name: Validate the tag
id: build id: version
env: env:
VERSION: ${{ gitea.ref_name }} VERSION: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ m{^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$}
or die qq{ERROR: expected a semver tag like v1.2.3, got: $v\n};
(my $nv = $v) =~ s{^v}{};
open(my $o, q{>>}, $ENV{GITEA_OUTPUT}) or die qq{GITEA_OUTPUT: $!};
print $o qq{version_no_v=$nv\n};
close($o);
print qq{version $nv\n};
'
if ! echo "$VERSION" | grep -qE '^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$'; then - name: Build
echo "ERROR: expected a semver tag like v1.2.3, got: '$VERSION'" env:
exit 1 VERSION_NO_V: ${{ steps.version.outputs.version_no_v }}
fi GOOS: ${{ matrix.goos }}
GOARCH: ${{ matrix.goarch }}
VERSION_NO_V="${VERSION#v}" CGO_ENABLED: "0"
echo "version_no_v=${VERSION_NO_V}" >> "$GITEA_OUTPUT" run: |
# Nothing is injected. The toolchain records the tag into the binary's build
mkdir -p bin # information, so the version is right because this build happens at the tag, and
GOOS=${{ matrix.goos }} GOARCH=${{ matrix.goarch }} CGO_ENABLED=0 \ # there is no path for anyone to get wrong. -s -w only strips symbols.
go build -ldflags "-s -w -X main.version=${VERSION_NO_V}" \ go build -ldflags "-s -w" -o "bin/gasm-${VERSION_NO_V}-${GOOS}-${GOARCH}" ./cmd/gasm
-o "bin/gasm-${VERSION_NO_V}-${{ matrix.goos }}-${{ matrix.goarch }}" \
./cmd/gasm
# Artifacts stay on v3: v4 and later detect Gitea as GHES and abort.
- name: Upload artifact - name: Upload artifact
uses: actions/upload-artifact@v3 uses: actions/upload-artifact@v3
with: with:
name: gasm-${{ matrix.goos }}-${{ matrix.goarch }} name: gasm-${{ matrix.goos }}-${{ matrix.goarch }}
path: bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} path: bin/gasm-${{ steps.version.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }}
if-no-files-found: error if-no-files-found: error
- name: Smoke test - name: Smoke test
# Only a binary matching the runner can be run here. The check is not that --version
# exits cleanly but that it reports the tag and nothing more: a build outside version
# control reports (devel), and a build whose tree was dirty reports +dirty, and both
# would otherwise be published.
if: matrix.goos == 'linux' && matrix.goarch == 'amd64' if: matrix.goos == 'linux' && matrix.goarch == 'amd64'
env:
TAG: ${{ gitea.ref_name }}
BIN: bin/gasm-${{ steps.version.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }}
run: | run: |
chmod +x bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} perl -e '
./bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} --version my $want = $ENV{TAG} // die qq{ERROR: no tag\n};
open(my $bin, q{-|}, $ENV{BIN}, q{--version}) or die qq{$ENV{BIN}: $!};
my $got = <$bin>;
close($bin);
$got = defined $got ? $got : q{};
chomp $got;
index($got, $want) >= 0
or die qq{ERROR: the binary printed "$got", which does not contain $want. Version control was disabled, so there is no recorded version.\n};
index($got, q{+dirty}) < 0
or die qq{ERROR: the binary printed "$got". The tree was dirty at build time, which means the checkout was not the tag, or the build artefacts are not ignored.\n};
print qq{$ENV{BIN} reports $got\n};
'
release: release:
runs-on: fedora runs-on: fedora
timeout-minutes: 15
needs: build needs: build
permissions: permissions:
# contents: read is required for the checkout: a job that declares any
# permissions gets a token scoped to exactly those, and releases: write
# alone leaves the fetch with no read access, which Gitea answers with
# a 404 "Repository not found". Verified on the instance 2026-09-16.
contents: read
releases: write releases: write
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
@@ -77,81 +202,125 @@ jobs:
with: with:
path: dist path: dist
- name: Extract CHANGELOG section - name: Install Perl
run: dnf install -y perl
- name: Extract the CHANGELOG section
env: env:
VERSION: ${{ gitea.ref_name }} VERSION: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail # Each step derives what it needs from the tag, so no value has to travel between
VERSION_NO_V="${VERSION#v}" # jobs.
perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ s{^v}{};
open(my $vout, q{>}, q{version-no-v.txt}) or die qq{version-no-v.txt: $!};
print $vout $v;
close($vout);
open(my $in, q{<}, q{CHANGELOG.md}) or die qq{CHANGELOG.md: $!};
my @lines = <$in>;
close($in);
my ($start, $end) = (-1, scalar @lines);
for my $i (0 .. $#lines) {
if ($start < 0) { $start = $i if $lines[$i] =~ m{^##\s+\[\Q$v\E\]} }
elsif ($lines[$i] =~ m{^##\s+\[}) { $end = $i; last }
}
$start >= 0 or die qq{ERROR: no CHANGELOG section for $v, expected a heading like: ## [$v] - YYYY-MM-DD\n};
my @body = grep { m{\S} } @lines[$start + 1 .. $end - 1];
@body or die qq{ERROR: the CHANGELOG section for $v is empty\n};
open(my $out, q{>}, q{release-body.md}) or die qq{release-body.md: $!};
print $out @body;
close($out);
printf qq{notes for %s: %d lines\n}, $v, scalar @body;
'
sed -n "/^## \[${VERSION_NO_V}\] /,/^## \[/p" CHANGELOG.md \ - name: Build the release request
| sed '$d' \ run: |
| tail -n +2 \ perl -e '
> release-body.md open(my $vin, q{<}, q{version-no-v.txt}) or die qq{version-no-v.txt: $!};
my $v = <$vin>;
close($vin);
chomp $v;
open(my $in, q{<:raw}, q{release-body.md}) or die qq{release-body.md: $!};
my $body = do { local $/; <$in> };
close($in);
# Byte-oriented escaping: JSON is UTF-8, so non-ASCII passes through and only the
# characters JSON forbids are rewritten.
$body =~ s/([\\"])/\\$1/g;
$body =~ s/\t/\\t/g;
$body =~ s/\r//g;
$body =~ s/\n/\\n/g;
$body =~ s/([\x00-\x08\x0b\x0c\x0e-\x1f])/sprintf(q{\u%04x}, ord($1))/ge;
my $json = sprintf(qq{{"tag_name":"v%s","name":"v%s","body":"%s","draft":false,"prerelease":false}}, $v, $v, $body);
open(my $out, q{>}, q{release.json}) or die qq{release.json: $!};
print $out $json;
close($out);
print qq{release.json written for v$v\n};
'
if [ ! -s release-body.md ]; then - name: Create the release
echo "ERROR: no CHANGELOG section found for ${VERSION_NO_V}"
echo "Expected a heading like: ## [${VERSION_NO_V}] — YYYY-MM-DD"
exit 1
fi
- name: Create release
env: env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }} GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
GITEA_SERVER_URL: ${{ gitea.server_url }} GITEA_SERVER_URL: ${{ gitea.server_url }}
GITEA_REPOSITORY: ${{ gitea.repository }} GITEA_REPOSITORY: ${{ gitea.repository }}
GITEA_REF_NAME: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
my @cmd = (q{curl}, q{-sS}, q{-o}, q{response.json}, q{-w}, q{%{http_code}},
BODY=$(sed -e 's/\\/\\\\/g' -e 's/"/\\"/g' -e 's/\t/\\t/g' -e 's/\r//g' release-body.md | sed ':a;N;$!ba;s/\n/\\n/g') q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
BODY="\"${BODY}\"" q{-H}, q{Content-Type: application/json},
q{-X}, q{POST},
response=$(curl -sS -w '\n%{http_code}' \ qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases},
-H "Authorization: token ${GITEA_TOKEN}" \ q{--data-binary}, q{@release.json});
-H "Content-Type: application/json" \ open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
-X POST \ my $code = <$curl>;
"${GITEA_SERVER_URL}/api/v1/repos/${GITEA_REPOSITORY}/releases" \ my $ok = close($curl);
-d "{\"tag_name\":\"${GITEA_REF_NAME}\",\"name\":\"${GITEA_REF_NAME}\",\"body\":${BODY},\"draft\":false,\"prerelease\":false}") my $exit = $? >> 8;
$code = defined $code ? $code : q{};
http_code=$(echo "$response" | tail -1) $ok or die qq{ERROR: curl failed (exit $exit) calling $ENV{GITEA_SERVER_URL}\n};
payload=$(echo "$response" | sed '$d') open(my $r, q{<:raw}, q{response.json}) or die qq{response.json: $!};
my $body = do { local $/; <$r> };
echo "HTTP ${http_code}" close($r);
if [ "$http_code" != "201" ]; then $code eq q{201} or die qq{ERROR: the release was not created, HTTP $code: $body\n};
echo "Failed to create release: ${payload}" $body =~ m{"id"\s*:\s*([0-9]+)} or die qq{ERROR: no release id in the response: $body\n};
exit 1 open(my $o, q{>}, q{release-id.txt}) or die qq{release-id.txt: $!};
fi print $o $1;
close($o);
RELEASE_ID=$(echo "$payload" | grep -oE '"id"[[:space:]]*:[[:space:]]*[0-9]+' | head -1 | grep -oE '[0-9]+') print qq{release id $1\n};
echo "Created release ID=${RELEASE_ID}" '
printf '%s' "${RELEASE_ID}" > release-id.txt
- name: Upload assets - name: Upload assets
env: env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }} GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
GITEA_SERVER_URL: ${{ gitea.server_url }} GITEA_SERVER_URL: ${{ gitea.server_url }}
GITEA_REPOSITORY: ${{ gitea.repository }} GITEA_REPOSITORY: ${{ gitea.repository }}
GITEA_REF_NAME: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
RELEASE_ID=$(cat release-id.txt) open(my $f, q{<}, q{release-id.txt}) or die qq{release-id.txt: $!};
my $id = <$f>;
for binary in dist/gasm-*/gasm-*; do close($f);
[ -f "$binary" ] || continue chomp $id;
fname=$(basename "$binary") my @files = grep { -f $_ } glob(q{dist/*/*});
echo "Uploading ${fname}..." @files or die qq{ERROR: no assets under dist/\n};
http_code=$(curl -sS -o /dev/null -w '%{http_code}' \ my $bad = 0;
-H "Authorization: token ${GITEA_TOKEN}" \ for my $path (@files) {
-H "Content-Type: application/octet-stream" \ (my $name = $path) =~ s{.*/}{};
-X POST \ my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
--data-binary "@${binary}" \ q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
"${GITEA_SERVER_URL}/api/v1/repos/${GITEA_REPOSITORY}/releases/${RELEASE_ID}/assets?name=${fname}") q{-H}, q{Content-Type: application/octet-stream},
echo " HTTP ${http_code}" q{-X}, q{POST}, q{--data-binary}, qq{@$path},
if [ "$http_code" != "201" ]; then qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
echo "Failed to upload ${fname}" open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
exit 1 my $code = <$curl>;
fi my $ok = close($curl);
done my $exit = $? >> 8;
$code = defined $code ? $code : q{};
echo "Release ${GITEA_REF_NAME} is live." unless ($ok) {
printf qq{%s: curl failed (exit %d)\n}, $name, $exit;
$bad = 1;
next;
}
printf qq{%s: HTTP %s\n}, $name, $code;
$bad = 1 if $code ne q{201};
}
exit($bad ? 1 : 0);
'
+82 -76
View File
@@ -1,4 +1,17 @@
# Test — gasm-devkit. Runs on push and pull request to development. # Test, Go. Push and pull request to development. Never on main.
#
# The gates are the ones the justfile's `gates` recipe runs, minus race: the shared
# runner box cannot afford the race detector on every push, so it lives in race.yml.
# The box is one core and 2 GB beside Gitea, so parallelism is bounded on purpose and
# everything runs in one job. Extra jobs would duplicate the checkout, the Go setup and
# the dependency download three times without buying any parallelism.
#
# Every step is one command, so the step that fails is the gate that failed, and no shell
# option has to be trusted for the run to stop. The scripted steps are Perl, not shell and
# not Python: Perl behaves the same on both runner images, there is no bashism to trip over
# on ash, and it is one language instead of two. The Perl uses builtins only, because
# Fedora packages the Perl modules separately and nothing beyond `perl` itself may be
# assumed present.
name: Test name: Test
on: on:
@@ -7,90 +20,83 @@ on:
pull_request: pull_request:
branches: [development] branches: [development]
env:
# One core: parallelism buys no speed here and costs memory the box does not have.
GOFLAGS: -p=1
GOMAXPROCS: "2"
# A superseded run of the same ref is cancelled instead of queueing behind one that
# no longer matters. Verified on Gitea 1.27.1 on 2026-09-17: a queued run whose ref
# moved on is cancelled before it ever reaches the runner, while a run already
# dispatched there runs to completion.
concurrency:
group: ${{ gitea.workflow }}-${{ gitea.ref }}
cancel-in-progress: true
jobs: jobs:
vet:
runs-on: fedora
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version: "1.27"
- name: Download dependencies
run: go mod download
- name: gofmt
run: |
set -euo pipefail
unformatted=$(gofmt -l .)
if [ -n "$unformatted" ]; then
echo "These files need gofmt:"
echo "$unformatted"
exit 1
fi
- name: go vet
run: go vet ./...
test: test:
runs-on: fedora runs-on: fedora
needs: vet timeout-minutes: 10
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
- uses: actions/setup-go@v6 - uses: actions/setup-go@v6
with: with:
go-version: "1.27" # The module is the source of truth for the version, so it cannot drift.
go-version-file: go.mod
cache: true
- name: Download dependencies - name: Install Perl
run: go mod download # The runner images are minimal and Perl is not guaranteed. The install is a
# no-op where it is already present; drop this step once verified on the box.
- name: Install gcc run: dnf install -y perl
run: dnf install -y gcc
- name: go test -race
run: go test -race -count=1 ./...
- name: Coverage gate — 80 % minimum
run: |
set -euo pipefail
# Exclude packages inherently untestable without hardware:
# debug — interactive ptrace, requires a live process
# cmd/gasm — CLI glue, covered by integration tests
go test -coverprofile=coverage.out \
sourcedock.dev/petrbalvin/gasm-devkit/arch \
sourcedock.dev/petrbalvin/gasm-devkit/asm \
sourcedock.dev/petrbalvin/gasm-devkit/ast \
sourcedock.dev/petrbalvin/gasm-devkit/format \
sourcedock.dev/petrbalvin/gasm-devkit/lexer \
sourcedock.dev/petrbalvin/gasm-devkit/lint \
sourcedock.dev/petrbalvin/gasm-devkit/lsp \
sourcedock.dev/petrbalvin/gasm-devkit/parser \
sourcedock.dev/petrbalvin/gasm-devkit/token \
sourcedock.dev/petrbalvin/gasm-devkit/verify
coverage=$(go tool cover -func=coverage.out | awk '/^total:/ { gsub("%", "", $3); print $3 }')
echo "Total coverage: ${coverage}%"
if awk -v c="$coverage" 'BEGIN { exit !(c+0 < 80) }'; then
echo "ERROR: coverage ${coverage}% is below the 80% threshold"
exit 1
fi
build:
runs-on: fedora
needs: test
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version: "1.27"
- name: Download dependencies
run: go mod download
# The steps follow the `gates` order of the justfile contract: build, format,
# vet, test. The vet gate is go vet and go fix -diff, two steps here.
- name: Build - name: Build
run: go build -ldflags="-s -w" -o bin/gasm ./cmd/gasm run: go build ./...
- name: Smoke test - name: Format
run: ./bin/gasm --version run: |
perl -e '
open(my $g, q{-|}, q{gofmt}, q{-l}, q{.}) or die qq{gofmt: $!};
my @bad = <$g>;
close($g);
print @bad;
exit(@bad ? 1 : 0);
'
- name: Vet
run: go vet ./...
- name: Modernise
# Exits non-zero when it has something to rewrite, so it needs no output capture.
run: go fix -diff ./...
- name: Tests
# The suite must be fast: a push pipeline that cannot finish in a few minutes moves
# its heavy part behind a dispatch. The inner timeout matches the job's, so a
# hanging test reports its own goroutine dump rather than a silent job kill.
# The pattern is `packages` in the project's justfile: the logic packages, since a
# thin cmd/ would drag the total under the floor. release.yml runs the same
# command, so the floor is the same number everywhere. ./verify/... carries the
# live oracle-parity comparison against `go tool asm` (the TestGroundTruth
# suites); the runner's Go setup provides both the tool and GOROOT.
run: go test -count=1 -timeout 10m -coverprofile=coverage.out ./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... ./lint/... ./lsp/... ./parser/... ./token/... ./verify/...
- name: Oracle parity
# Re-run the live go-tool-asm comparison as its own step so that a parity
# regression names the gate that failed instead of hiding inside the suite.
run: go test -count=1 -timeout 10m -run 'TestGroundTruth' ./verify/...
- name: Coverage floor
run: |
perl -e '
open(my $c, q{-|}, q{go}, q{tool}, q{cover}, q{-func=coverage.out}) or die qq{cover: $!};
my $total;
while (my $l = <$c>) { $total = $1 if $l =~ m{^total:\s+\S+\s+([0-9.]+)%} }
close($c);
die qq{no total line in coverage.out\n} unless defined $total;
printf qq{Total coverage: %s%%\n}, $total;
exit($total < 80 ? 1 : 0);
'
+3 -14
View File
@@ -1,24 +1,13 @@
# Metadata (always first, per repo convention)
.idea/ .idea/
.zcode/ .zcode/
.mimocode/
# Binaries # Build output
/gasm
/bin/ /bin/
*.exe /gasm
# Test and coverage artefacts
coverage.out coverage.out
*.test *.test
# Crash dumps # Crash dumps from the emulator runs
core core
core.* core.*
*.core *.core
# Scratch / temporary work
_scratch/
# ZCode workspace
.zcode
+496 -148
View File
File diff suppressed because it is too large Load Diff
+99 -78
View File
@@ -1,107 +1,128 @@
# Contributing to gasm-devkit # Contributing
Thanks for contributing to gasm-devkit. Contributions to **gasm-devkit** are governed by the Contributor terms
below; submitting one means you accept them.
## Contributor terms
1. This project belongs to its owner alone. The owner decides what is
accepted, in what form and when; the decision is final and needs no
justification.
2. By submitting a contribution you assign to Petr Balvín
<opensource@petrbalvin.org> all present and future copyright and
related rights in it, worldwide, for the full term of the rights,
with the right to relicense and sublicense without restriction,
including under proprietary terms.
3. Where that assignment is not effective, it counts as a perpetual,
irrevocable, royalty-free licence with the same scope.
4. To the fullest extent permitted by law, you waive any right of
attribution and integrity in the contribution. The project names no
contributors and keeps no credits list.
5. By submitting you represent that the work is yours and that you
hold the rights to assign it as above.
## Development setup ## Development setup
Requirements: Go 1.27 or later, the [just](https://github.com/casey/just) Requirements: Go 1.27.1, the exact version the `go` directive in `go.mod`
command runner, and a Linux host on amd64, arm64, riscv64 or loong64. declares, and [just](https://github.com/casey/just) for the recipes.
```sh ```sh
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
cd gasm-devkit cd gasm-devkit
just install # download module dependencies just build
just build # go vet + gofmt check just gates
just test # full suite, race detector, 80 % coverage gate
``` ```
## Workflow ## Workflow
1. Branch from `development`; never commit directly to `main` (`main` is 1. Branch from `development`. Never commit directly to `main`, which is release-only.
release-only: merge from `development`, then tag). 2. Commit in [Conventional Commits](https://www.conventionalcommits.org/) form:
2. Commit with [Conventional Commits](https://www.conventionalcommits.org/): `type(scope): description`, subject line only, imperative mood, lowercase after the
`type(scope): description`: subject line only, imperative mood, colon, no trailing full stop. Allowed types: `feat`, `fix`, `docs`, `style`,
lowercase after the colon, no trailing dot. Allowed types: `feat`, `refactor`, `perf`, `test`, `chore`, `ci`, `build`, `revert`.
`fix`, `docs`, `style`, `refactor`, `perf`, `test`, `chore`, `ci`, 3. One logical change per commit. A refactor, a behaviour change and a formatting pass
`build`, `revert`. The only line after the subject is the trailer: are three commits, never one.
`Assisted-by: <model-name>`. No `Co-Authored-By`, no `Signed-off-by`, 4. Record every user-visible change in `CHANGELOG.md` under `## [development]`.
no other trailers. 5. Add or update tests. Coverage stays at 80 percent or more; it is a hard gate.
3. Record every user-visible change in `CHANGELOG.md` under 6. Update the documentation when the public API, the configuration or the behaviour
`## [development]` (categories: Added, Changed, Fixed, Removed, changes.
Security). 7. Open a pull request against `development`.
4. Add or update tests; coverage must stay **at or above 80 %** (hard
gate, enforced by CI).
5. Update the documentation when behaviour, flags or the public surface
change.
6. Open a pull request against `development`.
Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`; Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`. The release
CI builds and publishes the binaries for all four architectures. workflow builds the assets and publishes the release and its notes.
## Code style ## Code style
`gofmt` and `go vet` via `just fmt` / `just build`; both must pass with `gofmt` and `go vet` run through `just fmt` and `just vet`, with zero diff and zero
zero output; `go fix -diff ./...` must report nothing on touched packages. warnings tolerated. `just gates` is the definition of done in one command, and the recipe
file names what it contains. Errors are checked explicitly, wrapped as
`fmt.Errorf("context: %w", err)`, and nothing panics outside `main`. The recipe file holds
the commands, and the language and standard-library surface is the one the `go` directive
in `go.mod` pins.
- Standard library only in production code; `golang.org/x/arch` is used - `golang.org/x/arch` is the one module dependency, and it is linked into the binary:
in tests only (round-trip decoding) and is never linked into the `gasm` `gasm dis` and the debugger's listings decode through it. Everything else is the
binary. standard library.
- No cgo, no C, no external toolchains at runtime. - No cgo, no C, no external toolchain at runtime.
- Explicit `if err != nil`; errors wrapped with - The parser, lexer and formatter are hand-written; the `arch` instruction tables are
`fmt.Errorf("context: %w", err)`; no panics outside `main`. generated only by `_gen/gen.go` (`just gen`) and never edited by hand.
- The parser, lexer and formatter are hand-written; the `arch` instruction - Assembly committed to the repository goes through `gasm fmt` and `gasm lint`, so a
tables are generated only via `_gen/gen.go` (`just gen`), never edited. `.s` file that `gasm fmt -l .` lists is unfinished.
## Running a single test New source files open with the project's two-line licence header, whose SPDX
identifier matches `LICENSE`. Configuration files, workflows and dotfiles do not carry
it.
```sh ## AI contribution policy
go test -run TestVexGroundTruth ./asm/
go test -run TestGroundTruthBasic ./verify/
go test -run TestGOObjectLinkAndRun ./asm/
go test -run TestFuzzWideCopy ./verify/
```
The interactive debugger (`gasm debug`) requires a compiled binary on AI tools are welcome as productivity aids and are a normal part of modern software
`$PATH`; `go run` does not work for the traced child process. Install development. What matters is that the contribution stays understandable, reviewable and
first with `just install-bin`. genuinely useful.
## CI (Gitea Actions) - **Disclose the assistance.** If AI helped draft any part of a commit, issue, pull
request or review, say so.
- **Commit messages carry exactly one trailer**, as a git trailer on the line after a
blank line that closes the subject:
Workflows live in `.gitea/workflows/` and run on self-hosted runners: ```
Assisted-by: MODEL
```
Name the model that did the work, spelled the way its maker spells it, for example
`GLM 5.3`, `DeepSeek V4.1 Flash` or `Qwen 3.8 Flash`. No `Co-Authored-By`, no `Signed-off-by`,
no other trailers, and no prose: the trailer is the disclosure.
- **Issues and pull requests** attribute the assistance in a comment, for example
`_Assisted-by: GLM 5.3_`. It does not belong in the pull request description.
- **Take responsibility.** You are accountable for the accuracy, completeness and
intent of everything you submit, whether or not AI produced it.
- **Review before marking ready.** Read the diff carefully, run it locally, and add the
tests it needs. Do not mark a pull request ready until you can defend every change in
it.
- **Quality over quantity.** Contributions that look like un-reviewed output, or whose
author cannot engage substantively during review, may be closed.
- **Preferred models.** Prefer open-weight models with transparent training data and
minimal output filtering.
AI assists. It does not replace judgement.
## Continuous integration
Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | What it does | | Workflow | Trigger | What it does |
|----------|---------|--------------| |---|---|---|
| Test | push / PR to `development` | gofmt check, `go vet`, `go test -race`, 80 % coverage gate | | Test | push or pull request to `development` | build, format check, vet, modernisation, the test suite with the coverage floor |
| Release | tag `v*` | cross-compiles binaries for linux/{amd64,arm64,riscv64,loong64} and publishes the Gitea release | | Release | a `v*` tag | the same gates as Test, then the matrix build, the proven version and the release itself; the race detector runs locally in `just gates` before the tag is cut |
The Definition of Done (`just build` + `just test` + `just fmt`) must The local equivalent is `just gates`, which is the same set plus the race detector. The
still pass locally before pushing. race detector also has its own workflow, dispatched by hand; it never runs on a push or a
tag, where it would double the time and the memory a shared runner cannot spare.
## AI Contribution Policy
AI tools are welcome as productivity aids. What matters is that
contributions remain understandable, reviewable, and genuinely useful.
- **Disclose AI use.** If you used AI to draft or generate any part of a
commit, issue, pull request, or code review, say so clearly.
- **Commit messages:** end every commit with exactly one trailer:
`Assisted-by: <model-name>` (e.g. `Assisted-by: GLM 5.3`).
- **Pull requests and issues:** attribute AI assistance in one trailing
line, e.g. `_Assisted-by: GLM 5.3_`. Do not paste it into the PR
description as a section.
- **Take responsibility.** You remain accountable for the accuracy,
completeness, and intent of everything you submit.
- **Review before marking ready.** Read AI-generated diffs carefully, run
them locally, and add or update tests where appropriate.
- **Preferred models.** Prefer open-weight models with transparent
training data: **GLM**, **DeepSeek**, and **MiMo**.
## Reporting bugs ## Reporting bugs
Open an issue at Open an issue at `https://sourcedock.dev/petrbalvin/gasm-devkit/issues` with the
[sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit/issues) version, the operating system and architecture, the exact command, the full output,
with the version (`gasm --version`), OS and architecture, the exact and the expected against the actual behaviour.
command, the full output, and the expected versus actual behaviour.
**Security issues:** email **opensource@petrbalvin.org** instead of opening **Security issues do not go in the issue tracker.** Report them as
a public issue. [SECURITY.md](SECURITY.md) describes, to **opensource@petrbalvin.org**.
+124 -26
View File
@@ -1,13 +1,63 @@
# gasm-devkit # Plan 9 assembly tooling, inside and outside Go
Developer tooling for **GAsm**, Go's built-in Plan 9 assembler. > **Warning: this is an experiment.** gasm-devkit is under active
> development and is not stable. The version is 0.x.x: commands, flags,
> output formats and behaviour can change without warning at any time.
> A 1.0.0 release is light years away. Nothing in this document is a
> stability promise. For all of that, this is not a paper project: gasm
> is already in active use and is tested on real assembly work.
Go ships an assembler but no tooling for it: there is no syntax highlighting, **GAsm** is Go's Plan 9 assembler, and Go ships it without tooling:
no autocomplete, no linter, no static analyser, no formatter, no standalone there is no formatter, no linter, no static analyser, no standalone
assembler and no debugger for `.s` files. Developers write assembly blind, assembler and no debugger for `.s` files. Developers write assembly
validate it by benchmark, and debug it by print statement. gasm-devkit is the blind, validate it by benchmark, and debug it by print statement.
missing toolkit: a single, self-contained binary, `gasm`, that brings proper gasm-devkit is the missing toolkit: a single, self-contained binary,
developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64. `gasm`, that serves both purposes.
- **Help develop Plan 9 assembly.** Formatting, linting, disassembly,
dynamic verification, a source-level debugger and a language server,
for `.s` files in Go programs.
- **Use Plan 9 assembly outside the Go toolchain.** `gasm asm` encodes
on its own, with no Go installation in the loop, and writes raw
images, linkable ELF objects with DWARF5 debug sections, or the Go
toolchain's own GOOBJ format, which `go build` consumes in place of
the toolchain's output.
## Why Plan 9 assembly
Plan 9 assembly is the quiet triumph of the field. One syntax across
every architecture Go builds for: the same source-first operand order,
the same four pseudo-registers, the same frame convention, whether the
target is x86, ARM, RISC-V or LoongArch. Learn it once and you can
read a kernel on any of them.
Compare the alternatives. Intel syntax and AT&T syntax disagree on the
one question every instruction answers, which operand is the source
and which is the destination, so half the world writes it one way,
half the other, and every assembly programmer carries both in their
head forever. GNU as settles the argument with directives that switch
dialects mid-file (`.intel_syntax noprefix`), a percent sign on every
register and a dollar on every immediate: punctuation that carries
nothing the operand order did not already say. And the x86 family
fragments again underneath: NASM is not MASM is not GAS, each with its
own directive zoo and macro language, so every project picks a dialect
and every reader learns a different one by accident.
Plan 9 assembly has none of it. Registers are bare names. Memory is
one notation, `offset(base)`, extended by an index and a scale when
the instruction needs it. Arguments arrive named and offset-checked:
`x+0(FP)` is the argument x, on every architecture, and `go vet`
polices the offsets against the Go prototype.
```text
AT&T (GNU as): movq %rax, -16(%rbp)
Plan 9 (Go): MOVQ AX, total-16(SP)
```
The same lines, but only one of them tells you what the number is for.
The syntax is uppercase, regular and boring, which is the highest
compliment a language for machine code can earn. gasm-devkit exists
to give that syntax the tooling it deserves.
## Features ## Features
@@ -16,7 +66,8 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
directly. directly.
- **Formatter.** `gasm fmt` canonicalises indentation, operand spacing, - **Formatter.** `gasm fmt` canonicalises indentation, operand spacing,
per-function mnemonic alignment and blank-line layout: `gofmt` for assembly, per-function mnemonic alignment and blank-line layout: `gofmt` for assembly,
operating recursively on directories the way `go fmt` does. operating recursively on directories the way `go fmt` does. `-l` lists
files whose formatting differs and `-d` prints a unified diff.
- **Linter.** `gasm lint` runs 18 conservative static checks, among them - **Linter.** `gasm lint` runs 18 conservative static checks, among them
`undefined-label`, `abi-argsize` (declared frame vs the `// func` signature), `undefined-label`, `abi-argsize` (declared frame vs the `// func` signature),
`register-clobber` (Go ABI register liveness over the control-flow graph), `register-clobber` (Go ABI register liveness over the control-flow graph),
@@ -24,7 +75,11 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
- **Standalone assembler.** `gasm asm` encodes all four architectures without - **Standalone assembler.** `gasm asm` encodes all four architectures without
the Go toolchain and writes raw images, linkable ELF objects (with DWARF5 the Go toolchain and writes raw images, linkable ELF objects (with DWARF5
debug sections) or the Go toolchain's own GOOBJ format, which `go build` debug sections) or the Go toolchain's own GOOBJ format, which `go build`
consumes in place of the toolchain's output. consumes in place of the toolchain's output. Framed functions get the
stack-split guard and the morestack block, byte-identical to the
toolchain's, so split functions link too.
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
offsets after assembling, or disassembles raw bytes from a file or stdin.
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into - **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
executable memory: smoke calls, ABI checks (sentinel registers, red-zone executable memory: smoke calls, ABI checks (sentinel registers, red-zone
canary), differential fuzzing against the `go tool asm` build, and canary), differential fuzzing against the `go tool asm` build, and
@@ -36,17 +91,17 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
push and pull diagnostics, semantic-token highlighting, go-to-definition, push and pull diagnostics, semantic-token highlighting, go-to-definition,
find references, rename, formatting, inlay hints, code actions, signature find references, rename, formatting, inlay hints, code actions, signature
help, document highlights, workspace symbol search, #include document help, document highlights, workspace symbol search, #include document
links and folding ranges over stdio. links and folding ranges over stdio; definition, references and rename
work across every open document.
- **Comparators and audits.** `gasm diff` compares the machine code of two - **Comparators and audits.** `gasm diff` compares the machine code of two
assembly files byte-for-byte, `gasm profile` shows basic-block structure, assembly files byte-for-byte, `gasm profile` shows basic-block structure,
`gasm audit-instructions` diffs the encoder against the installed toolchain, `gasm audit-instructions` diffs the encoder against the installed toolchain,
and `gasm scaffold` generates a differential test skeleton for a kernel. and `gasm scaffold` generates a differential test skeleton for a kernel.
- **Complete instruction coverage.** The instruction tables are generated
from the Go toolchain's own assembler source, so the toolkit recognises
every mnemonic the real assembler accepts; `just gen` refreshes them.
### Architecture support ### Architecture support
Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised | | Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------| |--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases | | AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
@@ -57,27 +112,64 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
"Common opcodes" are the instructions shared by every architecture (`RET`, "Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally `JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
carries the traditional conditional-jump spellings (`JZ`, `JNZ`, `JA`, `JC`, carries the traditional conditional-jump spellings (`JZ`, `JNZ`, `JA`, `JC`,
...) that the assembler accepts as aliases. Regenerating the tables is one ...) that the assembler accepts as aliases. The tables are generated from
command (`just gen`) and requires only a Go installation; the committed output the Go toolchain's own assembler source (`just gen` refreshes them), so
has no runtime dependency on the toolchain. every mnemonic the real assembler accepts is recognised; what the encoder
can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 127 of 627 files (20.3 %)
assembling for every target architecture today, with the top failure
reasons per architecture; the number moves with every release.
## Direction
The plan, in the order it is being worked:
- **Extended instruction support.** Two layers. First, encoding
coverage for every mnemonic the Go toolchain itself accepts, closed in
order of how often real code needs each instruction;
`gasm audit-instructions` measures the gap. Second, the larger work:
an extended instruction set the toolchain does not know at all. The
toolchain-derived tables stay generated and untouched; only the
extended instructions are hand-maintained, with their own spellings
and encoders, verified by execution on real hardware because the
toolchain offers no ground truth to compare against. The gaps exist
on every architecture, amd64 included.
- **Full GOOBJ and ELF compilation.** The destination is a complete,
standalone compilation path: linkable ELF objects for consumers outside
Go, and GOOBJ objects that `go build` links directly. Through GOOBJ, a
Go program will be able to use machine instructions that the Go
toolchain itself does not support; through ELF, Plan 9 assembly becomes
usable outside Go entirely.
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
architectures and is where the binary builds. FreeBSD follows: the
JIT's executable-memory mapping and the ptrace debugger layer are the
two pieces of porting work. Other unix systems may follow those two.
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
No others are planned.
## Install ## Install
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
linux/loong64 are on the linux/loong64 are on the
[releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases). [releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases).
From source (Go 1.27 or later): From source (Go 1.27.1):
```sh ```sh
go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest
``` ```
Or from a repository checkout, with the development version stamped: Or from a repository checkout:
```sh ```sh
just install-bin just install
``` ```
The installed binary reports the version the toolchain recorded: the tag
on a tagged checkout, a pseudo-version naming the commit below one.
## Quick start ## Quick start
```sh ```sh
@@ -102,9 +194,13 @@ gasm verify --call add --args a=2,b=3 hello_amd64.s # JIT-call it with argumen
```sh ```sh
gasm fmt # reformat every .s below here, like go fmt gasm fmt # reformat every .s below here, like go fmt
gasm fmt -w kernel_amd64.s # canonicalise one file in place gasm fmt -w kernel_amd64.s # canonicalise one file in place
gasm fmt -l *.s # list files whose formatting differs
gasm fmt -d kernel_amd64.s # print a unified diff instead
gasm lint *.s # static checks gasm lint *.s # static checks
gasm asm --format elf -o k.o k.s # assemble to a linkable ELF object gasm asm --format elf -o k.o k.s # assemble to a linkable ELF object
gasm asm --format goobj -p pkg/path -o k.o k.s # Go object, consumed by go build gasm asm --format goobj -p pkg/path -o k.o k.s # Go object, consumed by go build
gasm dis k.s # assemble, then list each function
gasm dis -a amd64 - < dump.bin # disassemble raw bytes from stdin
gasm verify --ground-truth k.s # byte-for-byte vs go tool asm gasm verify --ground-truth k.s # byte-for-byte vs go tool asm
gasm verify --fuzz k.s # differential fuzz vs the go tool asm build gasm verify --fuzz k.s # differential fuzz vs the go tool asm build
gasm debug --func name k.s # interactive debugger gasm debug --func name k.s # interactive debugger
@@ -132,9 +228,9 @@ infers the target architecture from the file-name suffix
## Development ## Development
```sh ```sh
just install # download module dependencies just build # compile, zero errors and zero warnings
just build # go vet + gofmt check, zero errors and zero warnings just test # the suite, no cache, the 80 % coverage floor
just test # full suite, race detector, 80 % coverage gate just gates # build, fmt-check, vet, test, race: the definition of done
just fmt # gofmt the tree just fmt # gofmt the tree
just gen # regenerate the instruction tables from the Go toolchain just gen # regenerate the instruction tables from the Go toolchain
``` ```
@@ -145,14 +241,16 @@ recipe.
## Documentation ## Documentation
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/CLI.md](docs/CLI.md): full command reference - [docs/CLI.md](docs/CLI.md): full command reference
- man pages: `just install-man` installs gasm(1) and one page per command
into ~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes - [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [docs/DECISIONS.md](docs/DECISIONS.md): deferred design decisions
- [CHANGELOG.md](CHANGELOG.md): release history - [CHANGELOG.md](CHANGELOG.md): release history
## Licence ## Licence
BSD-3-Clause — see [LICENSE](LICENSE). BSD-3-Clause; see [LICENSE](LICENSE).
Copyright © 2026 [Petr Balvín](https://petrbalvin.org) Copyright © 2026 [Petr Balvín](https://petrbalvin.org)
+40
View File
@@ -0,0 +1,40 @@
# Security policy
## Supported versions
Security fixes go to the newest release and to the `development` branch. Older
releases do not receive them.
| Version | Supported |
|---|---|
| 0.33.0 | yes |
| older releases | no |
## Reporting a vulnerability
**Do not open a public issue for a security problem.** A public report tells everyone
about the flaw before there is a fix. Report it privately to
**opensource@petrbalvin.org**.
Include:
- the version or commit you tested, and the platform
- what the problem is, and what an attacker gains from it
- the smallest reproducer you have, ideally a test or a single command
- a suggested fix, if you have one
## What to expect
- A human reads the report, and you get an acknowledgement.
- You are kept informed while the fix is being made, and told when it ships.
- The fix is released before the details are published, and the timing is agreed with
you.
- The reporter is credited in the release notes unless they ask otherwise.
## Out of scope
- Findings that require the attacker to already run code as the user, or to have local
access.
- Missing hardening with no demonstrated impact.
- Flaws in a third-party dependency: report them to that project, and to this one only
when this project's use of it makes them reachable.
+3
View File
@@ -55,6 +55,7 @@ const (
Mask // AVX-512 mask register (K) Mask // AVX-512 mask register (K)
Float // arm64 floating-point register (F) Float // arm64 floating-point register (F)
VecARM // arm64 SIMD/vector register (V) VecARM // arm64 SIMD/vector register (V)
VecSIMD // architecture-neutral SIMD/vector register (LoongArch LSX/LASX)
Special // architecture-special register Special // architecture-special register
) )
@@ -73,6 +74,8 @@ func (c RegClass) String() string {
return "float" return "float"
case VecARM: case VecARM:
return "vector (arm64)" return "vector (arm64)"
case VecSIMD:
return "vector"
case Special: case Special:
return "special" return "special"
default: default:
+2 -2
View File
@@ -29,7 +29,7 @@ func arm64Registers() []Register {
regs = append(regs, Register{Name: name, Class: class, Desc: desc}) regs = append(regs, Register{Name: name, Class: class, Desc: desc})
} }
// General-purpose integer registers R0–R30. // General-purpose integer registers R0-R30.
for i := 0; i <= 30; i++ { for i := 0; i <= 30; i++ {
add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register") add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register")
} }
@@ -145,7 +145,7 @@ func arm64Curated() []Instr {
for _, op := range []string{ for _, op := range []string{
"LDAXR", "LDAXRB", "LDAXRH", "LDAXRW", "STXR", "STXRB", "STXRH", "STXRW", "LDAXR", "LDAXRB", "LDAXRH", "LDAXRW", "STXR", "STXRB", "STXRH", "STXRW",
"LDAR", "LDARB", "LDARH", "LDARW", "STLR", "STLRB", "STLRH", "STLRW", "LDAR", "LDARB", "LDARH", "LDARW", "STLR", "STLRB", "STLRH", "STLRW",
"LDADD", "LDCLR", "LDEOR", "LDSET", "SWP", "CAS", "CASAL", "CASL", "CASAL", "LDADD", "LDCLR", "LDEOR", "LDSET", "SWP", "CAS", "CASAL", "CASL",
} { } {
t = append(t, i(op, "Atomic memory operation")) t = append(t, i(op, "Atomic memory operation"))
} }
+2 -2
View File
@@ -31,10 +31,10 @@ func loong64Registers() []Register {
add(fmt.Sprintf("F%d", i), Float, "floating-point register") add(fmt.Sprintf("F%d", i), Float, "floating-point register")
} }
for i := 0; i <= 31; i++ { for i := 0; i <= 31; i++ {
add(fmt.Sprintf("V%d", i), VecARM, "LSX 128-bit vector register") add(fmt.Sprintf("V%d", i), VecSIMD, "LSX 128-bit vector register")
} }
for i := 0; i <= 31; i++ { for i := 0; i <= 31; i++ {
add(fmt.Sprintf("X%d", i), VecARM, "LASX 256-bit vector register") add(fmt.Sprintf("X%d", i), VecSIMD, "LASX 256-bit vector register")
} }
return regs return regs
} }
+69
View File
@@ -4,6 +4,7 @@
package asm package asm
import ( import (
"encoding/binary"
"os" "os"
"os/exec" "os/exec"
"path/filepath" "path/filepath"
@@ -61,6 +62,62 @@ TEXT ·add(SB), NOSPLIT, $0-24
} }
} }
// TestGOObjectAARCH64PairReloc pins the ADRP-pair relocation shape against
// the toolchain's own object for the same source: exactly one R_ADDRARM64
// of Siz 8 at the ADRP word (cmd/internal/obj/arm64/asm7.go adds a single
// Siz-8 relocation per pair and the linker patches both instructions from
// it). gasm's assembler records the ADRP+ADD form as two word relocs; the
// emitter must coalesce them, not emit two Siz-4 records.
func TestGOObjectAARCH64PairReloc(t *testing.T) {
f, errs := parser.Parse("gv_arm64.s", `
#include "textflag.h"
TEXT ·getv(SB), NOSPLIT, $0-8
MOVD $v<>(SB), R4
MOVD R4, ret+0(FP)
RET
GLOBL v<>(SB), RODATA, $8
DATA v<>+0(SB)/8, $7
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
obj, err := img.GOObjectAARCH64("main", "gv_arm64.s")
if err != nil {
t.Fatalf("GOObjectAARCH64: %v", err)
}
v := openGoobj(t, obj)
relocs := v.blk(blkReloc)
le := binary.LittleEndian
// Two DWARF relocs on the lines/DIE symbols, then the code's one pair
// relocation.
if len(relocs) != 3*23 {
t.Fatalf("relocs = %d bytes, want three entries", len(relocs))
}
cr := relocs[2*23:]
if off := int32(le.Uint32(cr[0:])); off != 0 {
t.Errorf("pair reloc off = %d, want 0 (the ADRP word)", off)
}
if siz := cr[4]; siz != 8 {
t.Errorf("pair reloc siz = %d, want 8", siz)
}
if typ := le.Uint16(cr[5:]); typ != relocArm64Addr {
t.Errorf("pair reloc type = %d, want %d (R_ADDRARM64)", typ, relocArm64Addr)
}
if pkg := le.Uint32(cr[15:]); pkg != pkgIdxSelf {
t.Errorf("pair reloc PkgIdx = %#x, want pkgIdxSelf", pkg)
}
// The GLOBL is the first package definition.
if sym := le.Uint32(cr[19:]); sym != 0 {
t.Errorf("pair reloc SymIdx = %d, want 0 (the GLOBL definition)", sym)
}
}
// TestGOObjectAARCH64Link does an end-to-end link test: it cross-compiles a // TestGOObjectAARCH64Link does an end-to-end link test: it cross-compiles a
// Go program for arm64, substitutes the gasm-produced object into the package // Go program for arm64, substitutes the gasm-produced object into the package
// archive, re-links with cmd/link, and verifies the symbol appears in the // archive, re-links with cmd/link, and verifies the symbol appears in the
@@ -79,6 +136,14 @@ TEXT ·add(SB), NOSPLIT, $0-24
ADD R5, R4, R4 ADD R5, R4, R4
MOVD R4, ret+16(FP) MOVD R4, ret+16(FP)
RET RET
TEXT ·getv(SB), NOSPLIT, $0-8
MOVD $v<>(SB), R4
MOVD R4, ret+0(FP)
RET
GLOBL v<>(SB), RODATA, $8
DATA v<>+0(SB)/8, $7
` `
if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil { if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
@@ -86,11 +151,15 @@ TEXT ·add(SB), NOSPLIT, $0-24
mainSrc := `package main mainSrc := `package main
func add(a, b int64) int64 func add(a, b int64) int64
func getv() *int64
func main() { func main() {
if add(20, 22) != 42 { if add(20, 22) != 42 {
panic("bad add") panic("bad add")
} }
if getv() == nil {
panic("bad getv")
}
} }
` `
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil { if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
+404 -142
View File
@@ -12,17 +12,21 @@ import (
// assembleARM64 assembles an AArch64 (arm64) TEXT function body into machine // assembleARM64 assembles an AArch64 (arm64) TEXT function body into machine
// code. Every instruction is 4 bytes; the MOV pseudo-instruction and the // code. Every instruction is 4 bytes; the MOV pseudo-instruction and the
// immediate-arithmetic forms expand to 2–4 instructions when the immediate // immediate-arithmetic forms expand to 2-4 instructions when the immediate
// does not fit, so the layout is computed in two passes (sizes, then encoding // does not fit, so the layout is computed in two passes (sizes, then encoding
// with resolved branch targets). // with resolved branch targets).
// //
// The emitted bytes match the Go toolchain's arm64 assembler, which is the // The emitted bytes match the Go toolchain's arm64 assembler, which is the
// ground-truth oracle: prologue/epilogue, FP/SP frame mapping, branch // ground-truth oracle: prologue/epilogue, FP/SP frame mapping, branch
// encodings and the MOV immediate expansions all follow cmd/internal/obj/ // encodings and the MOV immediate expansions all follow cmd/internal/obj/
// arm64's asmout cases. // arm64's asmout cases. One deliberate difference: the stack-growth guard
// (the morestack check in the prologue and the call back into the runtime in
// the epilogue) is not emitted, so the bytes match only for NOSPLIT functions
// or zero-frame leaves, where the toolchain emits no guard either.
func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) { func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
fi := arm64ComputeFrame(t) fi := arm64ComputeFrame(t)
prologue := arm64Prologue(fi) prologue := arm64Prologue(fi)
guardLen := arm64GuardLen(fi)
chain := arm64JumpChain(t) chain := arm64JumpChain(t)
resolve := func(name string) string { resolve := func(name string) string {
if r, ok := chain[name]; ok { if r, ok := chain[name]; ok {
@@ -35,14 +39,14 @@ func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
var spadj []SpadjStep var spadj []SpadjStep
// The prologue (3 instructions when a small frame, 4 for large) // The prologue (3 instructions when a small frame, 4 for large)
// raises the SP delta by autosize. // raises the SP delta by autosize. The guard prefix shifts its PC.
if fi.autosize != 0 { if fi.autosize != 0 {
spadj = append(spadj, SpadjStep{PC: arm64PrologueSpadjPC(fi), Value: fi.autosize}) spadj = append(spadj, SpadjStep{PC: guardLen + arm64PrologueSpadjPC(fi), Value: fi.autosize})
} }
// Pass 1: label offsets from the instruction sizes. // Pass 1: label offsets from the instruction sizes.
offsets := map[string]int{} offsets := map[string]int{}
pos := len(prologue) pos := guardLen + len(prologue)
for _, stmt := range t.Body { for _, stmt := range t.Body {
switch s := stmt.(type) { switch s := stmt.(type) {
case *ast.Label: case *ast.Label:
@@ -52,9 +56,25 @@ func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
} }
} }
// Pass 2: encode. Relocation offsets are recorded function-relative. // Pass 2: encode. The guard prefix precedes the prologue; its branches
out := append([]byte(nil), prologue...) // target the morestack block at the end of the function, whose position
pc := len(prologue) // the first pass has settled.
bodyLen := 0
{
p := guardLen + len(prologue)
for _, stmt := range t.Body {
if in, ok := stmt.(*ast.Instr); ok {
p += arm64InstrSize(in, fi)
}
}
bodyLen = p - (guardLen + len(prologue))
}
var out []byte
if fi.needSplit {
out = append(out, arm64GuardBytes(fi, guardLen+len(prologue)+bodyLen)...)
}
out = append(out, prologue...)
pc := guardLen + len(prologue)
preCount := len(relocs) preCount := len(relocs)
var lines []LineEntry var lines []LineEntry
for _, stmt := range t.Body { for _, stmt := range t.Body {
@@ -67,7 +87,12 @@ func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err) return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err)
} }
for j := preCount; j < len(relocs); j++ { for j := preCount; j < len(relocs); j++ {
relocs[j].Off += pc - len(prologue) // Make the relocation offsets function-relative: each instruction
// records its reloc offset relative to its own start, and pc is
// that instruction's offset from the function start (prologue
// included). After shifts by the same amount.
relocs[j].Off += pc
relocs[j].After += pc
} }
preCount = len(relocs) preCount = len(relocs)
lines = append(lines, LineEntry{Offset: pc, Line: in.Pos().Line}) lines = append(lines, LineEntry{Offset: pc, Line: in.Pos().Line})
@@ -79,6 +104,12 @@ func assembleARM64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
out = append(out, code...) out = append(out, code...)
pc += len(code) pc += len(code)
} }
if fi.needSplit {
block, blReloc := arm64MoreStackBlock(pc)
out = append(out, block...)
relocs = append(relocs, blReloc)
pc += len(block)
}
return out, offsets, relocs, lines, spadj, nil return out, offsets, relocs, lines, spadj, nil
} }
@@ -158,7 +189,7 @@ func arm64InstrSize(instr *ast.Instr, fi arm64FrameInfo) int {
return arm64MovSize(mnem, ops, fi) return arm64MovSize(mnem, ops, fi)
case "ADD", "ADDW", "SUB", "SUBW", "AND", "ANDW", "ORR", "ORRW", "EOR", "EORW": case "ADD", "ADDW", "SUB", "SUBW", "AND", "ANDW", "ORR", "ORRW", "EOR", "EORW":
if len(ops) >= 2 && isImmOperand(ops[0]) { if len(ops) >= 2 && isImmOperand(ops[0]) {
v := immFromOperand(ops[0]) v := arm64Imm64(ops[0])
// Small immediate (0..4095 or -2048..-1) fits in one instruction. // Small immediate (0..4095 or -2048..-1) fits in one instruction.
if v >= 0 && v <= 0xFFF { if v >= 0 && v <= 0xFFF {
return 4 return 4
@@ -190,8 +221,12 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
if len(ops) != 1 { if len(ops) != 1 {
return nil, fmt.Errorf("WORD expects 1 operand, got %d", len(ops)) return nil, fmt.Errorf("WORD expects 1 operand, got %d", len(ops))
} }
return a64wordLE(uint32(immFromOperand(ops[0]))), nil w := arm64Imm64(ops[0])
case "B": if w < 0 || w > 0xFFFFFFFF {
return nil, fmt.Errorf("WORD: immediate %d does not fit a 32-bit word", w)
}
return a64wordLE(uint32(w)), nil
case "B", "JMP":
return encodeARM64Branch(mnem, ops, pc, offsets, false, relocs, resolve) return encodeARM64Branch(mnem, ops, pc, offsets, false, relocs, resolve)
case "BL", "CALL": case "BL", "CALL":
return encodeARM64Branch(mnem, ops, pc, offsets, true, relocs, resolve) return encodeARM64Branch(mnem, ops, pc, offsets, true, relocs, resolve)
@@ -205,6 +240,11 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
return encodeARM64BranchCond(mnem, enc.op, ops, pc, offsets, resolve) return encodeARM64BranchCond(mnem, enc.op, ops, pc, offsets, resolve)
} }
// Unconditional register branches (BR, BLR).
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FUncondBranch {
return encodeARM64RegBranch(mnem, enc.op, ops)
}
// ADD/SUB immediate. // ADD/SUB immediate.
if mnem == "ADD" || mnem == "ADDW" || mnem == "SUB" || mnem == "SUBW" || if mnem == "ADD" || mnem == "ADDW" || mnem == "SUB" || mnem == "SUBW" ||
mnem == "CMP" || mnem == "CMPW" || mnem == "CMN" || mnem == "CMNW" { mnem == "CMP" || mnem == "CMPW" || mnem == "CMN" || mnem == "CMNW" {
@@ -213,14 +253,19 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
} }
} }
// Shifts: immediate forms alias SBFM/UBFM/EXTR, register forms are the
// two-source LSLV/LSRV/ASRV/RORV.
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FShift {
return encodeARM64Shift(mnem, enc.op, ops)
}
// Multiply-accumulate: MADD/MSUB Rm, Ra, Rn, Rd.
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FDPR4 {
return encodeARM64MAddSub(mnem, enc.op, ops)
}
// Register-register data processing. // Register-register data processing.
// ASR/LSL/LSR/ROR with immediate operands use bitfield encoding (SBFM/UBFM).
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FDPSR { if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FDPSR {
isShift := mnem == "ASR" || mnem == "ASRW" || mnem == "LSL" || mnem == "LSLW" ||
mnem == "LSR" || mnem == "LSRW" || mnem == "ROR" || mnem == "RORW"
if isShift && len(ops) >= 2 && isImmOperand(ops[0]) {
return encodeARM64Bitfield(mnem, enc.op, ops)
}
return encodeARM64DPSR(mnem, enc.op, ops) return encodeARM64DPSR(mnem, enc.op, ops)
} }
@@ -269,7 +314,8 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
return encodeARM64CRC32(mnem, enc.op, ops) return encodeARM64CRC32(mnem, enc.op, ops)
} }
// Exclusive load/store (LDXR, STXR, LDAXR, STLXR). // Exclusive load/store (LDXR, STXR, LDAXR, STLXR and the register-pair
// forms LDXP, STXP).
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FExcl { if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FExcl {
return encodeARM64Excl(mnem, enc.op, ops) return encodeARM64Excl(mnem, enc.op, ops)
} }
@@ -306,8 +352,27 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
} }
op := ops[0] op := ops[0]
// External symbol reference: BL sym(SB). // Register-indirect: JMP (R0) is BR R0, CALL (R0) is BLR R0. The
if link && op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" { // toolchain's spelling carries no offset and no index; anything else
// is reported rather than silently dropped.
if op.Addr.Sym == nil && op.Addr.Base != "" {
if op.Addr.Offset != 0 || op.Addr.Index != "" {
return nil, fmt.Errorf("%s: invalid indirect branch operand %q", mnem, op.Raw)
}
rn := arm64RegNum(op.Addr.Base)
if rn < 0 {
return nil, fmt.Errorf("%s: unknown branch register %q", mnem, op.Addr.Base)
}
opc := uint32(0) // BR
if link {
opc = 1 // BLR
}
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
// relocation (R_CALLARM64 either way).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
if relocs != nil { if relocs != nil {
*relocs = append(*relocs, Reloc{ *relocs = append(*relocs, Reloc{
Off: 0, Off: 0,
@@ -317,8 +382,12 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
Kind: RelArm64Branch, Kind: RelArm64Branch,
}) })
} }
// Emit BL with zero offset; the linker fills in the target. // Emit B/BL with zero offset; the linker fills in the target.
return a64wordLE(a64Branch(1, 0)), nil bop := uint32(0) // B
if link {
bop = 1 // BL
}
return a64wordLE(a64Branch(bop, 0)), nil
} }
target := resolve(arm64Label(op)) target := resolve(arm64Label(op))
@@ -337,6 +406,19 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
return a64wordLE(a64Branch(bop, int32(rel))), nil return a64wordLE(a64Branch(bop, int32(rel))), nil
} }
// encodeARM64RegBranch encodes BR/BLR through a register operand:
// BR Xn = 0xd61f0000 | Rn<<5, BLR Xn = 0xd63f0000 | Rn<<5.
func encodeARM64RegBranch(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 1 {
return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(ops))
}
rn := arm64RegNum(operandRegName(ops[0]))
if rn < 0 {
return nil, fmt.Errorf("%s expects a register operand", mnem)
}
return a64wordLE(uint32(baseOp) | 31<<16 | uint32(rn)<<5), nil
}
// encodeARM64BranchCond encodes a conditional branch (B.cond) to a label. // encodeARM64BranchCond encodes a conditional branch (B.cond) to a label.
func encodeARM64BranchCond(mnem string, baseOp uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) { func encodeARM64BranchCond(mnem string, baseOp uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
if len(ops) != 1 { if len(ops) != 1 {
@@ -406,6 +488,88 @@ func encodeARM64DPSR(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, er
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
} }
// encodeARM64Shift encodes LSL/LSR/ASR/ROR in both widths. The operand order
// is source first, destination last: OP $sh|Rm, Rn, Rd or OP $sh|Rm, Rd.
// With an immediate the shift is the SBFM/UBFM (ROR: EXTR) alias, with a
// register it is the data-processing (2 source) LSLV/LSRV/ASRV/RORV; the
// two-source opcode rides the same 0xd6<<21 field as SDIV/UDIV, with
// LSLV=0b001000, LSRV=0b001001, ASRV=0b001010, RORV=0b001011 at bits 15:10.
func encodeARM64Shift(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 2 && len(ops) != 3 {
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
}
rn := arm64RegNum(operandRegName(ops[1]))
rd := arm64RegNum(operandRegName(ops[len(ops)-1]))
if rn < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem)
}
if isImmOperand(ops[0]) {
width := uint32(64)
if strings.HasSuffix(mnem, "W") {
width = 32
}
sh := arm64Imm64(ops[0])
if sh < 0 || uint32(sh) >= width {
return nil, fmt.Errorf("%s: shift amount %d out of range for %d-bit form", mnem, sh, width)
}
switch mnem {
case "LSL", "LSLW":
// UBFM Rd, Rn, #(-sh) mod W, #(W-1)-sh
immr := (width - uint32(sh)) % width
return a64wordLE(baseOp | immr<<16 | (width-1-uint32(sh))<<10 | uint32(rn)<<5 | uint32(rd)), nil
case "LSR", "LSRW":
// UBFM Rd, Rn, #sh, #(W-1)
return a64wordLE(baseOp | uint32(sh)<<16 | (width-1)<<10 | uint32(rn)<<5 | uint32(rd)), nil
case "ASR", "ASRW":
// SBFM Rd, Rn, #sh, #(W-1)
return a64wordLE(baseOp | uint32(sh)<<16 | (width-1)<<10 | uint32(rn)<<5 | uint32(rd)), nil
default:
// ROR, RORW: EXTR Rd, Rn, Rn, #sh (Rm = Rn, imms = sh).
return a64wordLE(baseOp | uint32(rn)<<16 | uint32(sh)<<10 | uint32(rn)<<5 | uint32(rd)), nil
}
}
rm := arm64RegNum(operandRegName(ops[0]))
if rm < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem)
}
op2 := uint32(8) // LSLV
switch mnem {
case "LSR", "LSRW":
op2 = 9 // LSRV
case "ASR", "ASRW":
op2 = 10 // ASRV
case "ROR", "RORW":
op2 = 11 // RORV
}
sf := uint32(1)
if strings.HasSuffix(mnem, "W") {
sf = 0
}
return a64wordLE(sf<<31 | 0xd6<<21 | op2<<10 | uint32(rm)<<16 | uint32(rn)<<5 | uint32(rd)), nil
}
// encodeARM64MAddSub encodes MADD/MSUB/MADDW/MSUBW. The toolchain's operand
// order is Rm, Ra, Rn, Rd (its optab case 15 comment says exactly that), so
// the accumulate register is the SECOND operand: base | Rm<<16 | Ra<<10 |
// Rn<<5 | Rd. The optab has no shorter row for these mnemonics, so all four
// operands are mandatory; MUL's two-operand spelling (Ra = ZR) belongs to the
// MUL mnemonic, not to these.
func encodeARM64MAddSub(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 4 {
return nil, fmt.Errorf("%s expects 4 operands (Rm, Ra, Rn, Rd), got %d", mnem, len(ops))
}
rm := arm64RegNum(operandRegName(ops[0]))
ra := arm64RegNum(operandRegName(ops[1]))
rn := arm64RegNum(operandRegName(ops[2]))
rd := arm64RegNum(operandRegName(ops[3]))
if rm < 0 || rn < 0 || ra < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem)
}
return a64wordLE(baseOp | uint32(rm)<<16 | uint32(ra)<<10 | uint32(rn)<<5 | uint32(rd)), nil
}
// ---- ADD/SUB immediate ---- // ---- ADD/SUB immediate ----
// encodeARM64AddSubImm encodes an ADD/SUB immediate instruction. // encodeARM64AddSubImm encodes an ADD/SUB immediate instruction.
@@ -413,7 +577,7 @@ func encodeARM64AddSubImm(mnem string, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 2 && len(ops) != 3 { if len(ops) != 2 && len(ops) != 3 {
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
} }
v := immFromOperand(ops[0]) v := arm64Imm64(ops[0])
rd := arm64RegNum(operandRegName(ops[len(ops)-1])) rd := arm64RegNum(operandRegName(ops[len(ops)-1]))
rn := rd rn := rd
if len(ops) == 3 { if len(ops) == 3 {
@@ -457,12 +621,14 @@ func encodeARM64AddSubImm(mnem string, ops []*ast.Operand) ([]byte, error) {
if v >= 0 && v <= 0xFFF000 && v&0xFFF == 0 { if v >= 0 && v <= 0xFFF000 && v&0xFFF == 0 {
return a64wordLE(a64AddSub(sf, op, S, 1, uint32(v>>12), uint32(rn), uint32(rd))), nil return a64wordLE(a64AddSub(sf, op, S, 1, uint32(v>>12), uint32(rn), uint32(rd))), nil
} }
// The imm12 field cannot carry the value; rejecting (rather than
// truncating) matches the toolchain, which reports the same shape.
return nil, fmt.Errorf("%s: immediate %d out of range for single instruction", mnem, v) return nil, fmt.Errorf("%s: immediate %d out of range for single instruction", mnem, v)
} }
// ---- MOV pseudo-instruction ---- // ---- MOV pseudo-instruction ----
// encodeARM64Mov encodes the MOV family — the load/store/immediate workhorse // encodeARM64Mov encodes the MOV family, the load/store/immediate workhorse
// of Go's arm64 assembly. MOV is an alias of MOVD (the width mnemonics // of Go's arm64 assembly. MOV is an alias of MOVD (the width mnemonics
// select the access width). The forms, mirroring the toolchain: // select the access width). The forms, mirroring the toolchain:
// //
@@ -543,14 +709,15 @@ func arm64MovSize(mnem string, ops []*ast.Operand, fi arm64FrameInfo) int {
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" { if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" {
return 8 // ADRP + ADD return 8 // ADRP + ADD
} }
v := arm64Imm64(src) // Size the immediate exactly as the encoder will emit it: multi-chunk
if v == 0 { // values expand to up to four words and the W forms truncate first.
// Anything else would desynchronise the label offsets of pass 1 from
// the bytes pass 2 lays down, corrupting every later branch.
b, err := encodeARM64LoadImm(31, arm64Imm64(src), mnem)
if err != nil {
return 4 return 4
} }
if arm64Movcon(v) >= 0 || arm64Movcon(^v) >= 0 { return len(b)
return 4
}
return 8 // MOVZ + MOVK
case src.Addr.Sym != nil && src.Addr.Sym.Pseudo == "SB": case src.Addr.Sym != nil && src.Addr.Sym.Pseudo == "SB":
return 8 // ADRP + LDR return 8 // ADRP + LDR
case dst.Addr.Sym != nil && dst.Addr.Sym.Pseudo == "SB": case dst.Addr.Sym != nil && dst.Addr.Sym.Pseudo == "SB":
@@ -562,69 +729,81 @@ func arm64MovSize(mnem string, ops []*ast.Operand, fi arm64FrameInfo) int {
} }
_, off := arm64MemWithFrame(mem, fi) _, off := arm64MemWithFrame(mem, fi)
// Scaled unsigned offset fits if aligned and in range. // Scaled unsigned offset fits if aligned and in range.
lt := a64LoadTable[mnem] lt, ok := a64LoadTable[mnem]
if lt.size == 0 { if !ok {
lt.size = 3 // default to64-bit for MOV lt = a64LoadTable["MOVD"] // the MOV pseudo is a 64-bit access
} }
scale := int32(1) << uint(lt.size) scale := int64(1) << uint(lt.size)
if off >= 0 && off%scale == 0 && off/scale < 4096 { if off >= 0 && off%scale == 0 && off/scale < 4096 {
return 4 return 4
} }
if off >= -256 && off <= 255 { if off >= -256 && off <= 255 {
return 4 // unscaled return 4 // unscaled
} }
return 12 // materialise offset + LDR/STR if _, _, _, ok := arm64SplitOffset(off, scale); ok {
return 8 // ADD base, REGTMP + access
}
return 12 // literal pool range: encoding reports it as unsupported
default: default:
return 4 // register move return 4 // register move
} }
} }
// encodeARM64LoadImm loads an immediate into a register, matching the // encodeARM64LoadImm loads an immediate into a register, matching the
// toolchain's MOVZ/MOVN/MOVK sequence. // toolchain's MOVZ/MOVN/MOVK sequence. W forms truncate to 32 bits first and
// every classification (movcon, complement, chunk count) runs on the truncated
// value, so a 32-bit immediate never reaches the 64-bit halves: MOVW $-1
// truncates to 0xFFFFFFFF, whose complement is a single zero chunk, and encodes
// as MOVN W, #0.
func encodeARM64LoadImm(rd int, v int64, mnem string) ([]byte, error) { func encodeARM64LoadImm(rd int, v int64, mnem string) ([]byte, error) {
d := v d := v
// For 32-bit MOVW, zero-extend. sf := uint32(1) // 64-bit
if mnem == "MOVW" || mnem == "MOVWU" { if mnem == "MOVW" || mnem == "MOVWU" {
d = int64(uint32(v)) d = int64(uint32(v))
sf = 0
} }
if d == 0 { if d == 0 {
// ORR Rd, ZR, ZR (MOV $0, Rd) // ORR Rd, ZR, ZR (MOV $0, Rd)
op := uint32(1<<31 | 1<<29 | 0x0a<<24) // ORR 64-bit op := uint32(1<<31 | 1<<29 | 0x0a<<24) // ORR 64-bit
if mnem == "MOVW" || mnem == "MOVWU" { if sf == 0 {
op = 0<<31 | 1<<29 | 0x0a<<24 // ORR 32-bit op = 0<<31 | 1<<29 | 0x0a<<24 // ORR 32-bit
} }
return a64wordLE(op | 31<<16 | 31<<5 | uint32(rd)), nil return a64wordLE(op | 31<<16 | 31<<5 | uint32(rd)), nil
} }
sf := uint32(1) // 64-bit // The Go toolchain classifies immediates (asm7.go conclass):
if mnem == "MOVW" || mnem == "MOVWU" { // - inside the imm12/shifted-imm12 "addcon" band (C_ABCON0/C_ABCON,
sf = 0 // 0 < v ≤ 4095 or a 4096 multiple up to 0xFFF000): bitmask first, so
} // `MOVD $4096, R27` is ORR $4096, not MOVZ $(1<<12)
// - outside that band: MOVZ/MOVN first (C_MOVCON before C_BITCON), and
// The Go toolchain classifies immediates: // negative values reach MOVN before the bitmask test
// - C_ABCON0 (0 < v ≤ 4095): bitmask first for positive values tryBitmaskFirst := d > 0 && (d <= 0xFFF || (d&0xFFF == 0 && d <= 0xFFF000))
// - Negative values: MOVN first, then bitmask
// - C_MOVCON (movcon-eligible, outside ABCON range): MOVZ/MOVN first
tryBitmaskFirst := d > 0 && d <= 0xFFF
if tryBitmaskFirst { if tryBitmaskFirst {
// Small immediate: try bitmask first (Go uses ORR for values like $1, $256). // Addcon-band immediate: try bitmask first (Go uses ORR for values
// like $1, $256 and $65536).
N, immr, imms, ok := arm64Bitmask(uint64(d), int(sf)) N, immr, imms, ok := arm64Bitmask(uint64(d), int(sf))
if ok { if ok {
return a64wordLE(sf<<31 | 1<<29 | 0x24<<23 | N<<22 | immr<<16 | imms<<10 | 31<<5 | uint32(rd)), nil return a64wordLE(sf<<31 | 1<<29 | 0x24<<23 | N<<22 | immr<<16 | imms<<10 | 31<<5 | uint32(rd)), nil
} }
} }
// Try MOVZ (single non-zero 16-bit chunk). // Try MOVZ (single non-zero 16-bit chunk) and MOVN (single non-0xFFFF
// chunk of the complement). The W forms must look inside the 32-bit
// window only, so the complement is masked to the operand width; d is
// already truncated and needs no mask.
width := uint64(0xFFFFFFFF)
if sf == 1 {
width = 0xFFFFFFFFFFFFFFFF
}
s := arm64Movcon(d) s := arm64Movcon(d)
if s >= 0 { if s >= 0 {
return a64wordLE(a64MoveWide(sf, 2, uint32(s>>4), uint32((d>>uint(s))&0xFFFF), uint32(rd))), nil return a64wordLE(a64MoveWide(sf, 2, uint32(s>>4), uint32((d>>uint(s))&0xFFFF), uint32(rd))), nil
} }
// Try MOVN (single non-0xFFFF 16-bit chunk of ^d). sn := arm64Movcon(^d & int64(width))
sn := arm64Movcon(^d)
if sn >= 0 { if sn >= 0 {
return a64wordLE(a64MoveWide(sf, 0, uint32(sn>>4), uint32((^d>>uint(sn))&0xFFFF), uint32(rd))), nil return a64wordLE(a64MoveWide(sf, 0, uint32(sn>>4), uint32(((^d)>>uint(sn))&0xFFFF), uint32(rd))), nil
} }
// For values outside the bitmask-first range that are not movcon: try bitmask. // For values outside the bitmask-first range that are not movcon: try bitmask.
@@ -635,7 +814,7 @@ func encodeARM64LoadImm(rd int, v int64, mnem string) ([]byte, error) {
} }
} }
// Multi-instruction: MOVZ + MOVK for each non-zero16-bit chunk. // Multi-instruction: MOVZ + MOVK for each non-zero 16-bit chunk.
var ws []uint32 var ws []uint32
first := true first := true
for i := range 4 { for i := range 4 {
@@ -730,7 +909,7 @@ func arm64Bitmask(v uint64, sf int) (N, immr, imms uint32, ok bool) {
// Integer → integer: ORR Rd, ZR, Rs. // Integer → integer: ORR Rd, ZR, Rs.
// FP → FP: FMOV Fd, Fn (FP data processing). // FP → FP: FMOV Fd, Fn (FP data processing).
// FP ↔ GP: FMOV general (FPCVTI encoding). // FP ↔ GP: FMOV general (FPCVTI encoding).
// Go Plan 9 syntax: MOV dst, src (first operand = destination). // Go Plan 9 syntax is source first, destination last: MOV src, dst.
func encodeARM64RegMove(mnem string, src, dst *ast.Operand) ([]byte, error) { func encodeARM64RegMove(mnem string, src, dst *ast.Operand) ([]byte, error) {
rs := arm64RegNum(operandRegName(src)) rs := arm64RegNum(operandRegName(src))
rd := arm64RegNum(operandRegName(dst)) rd := arm64RegNum(operandRegName(dst))
@@ -751,8 +930,7 @@ func encodeARM64RegMove(mnem string, src, dst *ast.Operand) ([]byte, error) {
} }
// GP ↔ FP: FMOV general (FPCVTI encoding). // GP ↔ FP: FMOV general (FPCVTI encoding).
// Go syntax: FMOV FPdst, GPsrc or FMOV GPdst, FPsrc. // Go syntax: FMOV GPsrc, FPdst or FMOV FPsrc, GPdst, source first.
// First operand = destination, second = source.
if sc == arm64ClsFP && dc == arm64ClsGR { if sc == arm64ClsFP && dc == arm64ClsGR {
// FP → GP: FMOV Wd/Xd, Sn/Dn. opcode bits[20:16]=6. // FP → GP: FMOV Wd/Xd, Sn/Dn. opcode bits[20:16]=6.
sf, typ := uint32(0), uint32(0) sf, typ := uint32(0), uint32(0)
@@ -792,34 +970,55 @@ func encodeARM64MemOp(mnem string, mem *ast.Operand, reg int, load bool, fi arm6
lt = a64LoadTable["MOVD"] lt = a64LoadTable["MOVD"]
} }
scale := int32(1) << uint(lt.size) scale := int64(1) << uint(lt.size)
if load {
// Try scaled unsigned offset first.
if off >= 0 && off%scale == 0 {
imm12 := uint32(off / scale)
if imm12 < 4096 {
return a64wordLE(a64LSU(uint32(lt.size), uint32(lt.V), uint32(lt.opc), imm12, uint32(rn), uint32(reg))), nil
}
}
// Try unscaled (9-bit signed).
if off >= -256 && off <= 255 {
return a64wordLE(a64LSUnscaled(lt.size, lt.V, lt.opc, off, rn, reg)), nil
}
// Large offset: materialise in R20 (TMP) and use register-offset.
return nil, fmt.Errorf("%s: offset %d out of range", mnem, off)
}
// Store: same encoding but opc bits indicate store.
storeOpc := a64StoreOpc(lt) storeOpc := a64StoreOpc(lt)
if off >= 0 && off%scale == 0 { var opc int
imm12 := uint32(off / scale) if load {
if imm12 < 4096 { opc = lt.opc
return a64wordLE(a64LSU(uint32(lt.size), uint32(lt.V), uint32(storeOpc), imm12, uint32(rn), uint32(reg))), nil } else {
} opc = storeOpc
}
// Scaled unsigned offset first, then the unscaled ±255 form.
if off >= 0 && off%scale == 0 && off/scale < 4096 {
return a64wordLE(a64LSU(uint32(lt.size), uint32(lt.V), uint32(opc), uint32(off/scale), uint32(rn), uint32(reg))), nil
} }
if off >= -256 && off <= 255 { if off >= -256 && off <= 255 {
return a64wordLE(a64LSUnscaled(lt.size, lt.V, storeOpc, off, rn, reg)), nil return a64wordLE(a64LSUnscaled(lt.size, lt.V, opc, int32(off), rn, reg)), nil
} }
return nil, fmt.Errorf("%s: offset %d out of range", mnem, off) // Large offset: materialise the base in REGTMP (R27) the way the
// toolchain does and access what remains. The ADD offsets from the
// operand's own base register, [SP] and [Rn] alike.
addImm, addShift, access, ok := arm64SplitOffset(off, scale)
if !ok {
return nil, fmt.Errorf("%s: offset %d out of range (literal pool not supported)", mnem, off)
}
return a64WordsLE(
a64AddSub(1, 0, 0, addShift, addImm, uint32(rn), 27), // ADD $addImm<<shift, Rn, R27
a64LSU(uint32(lt.size), uint32(lt.V), uint32(opc), uint32(access/scale), 27, uint32(reg)),
), nil
}
// arm64SplitOffset decomposes an out-of-range offset for a REGTMP base: an
// ADD (plain, or shifted left by 12) brings the base near the target and the
// access covers what remains. ok is false when no decomposition exists
// (negative offsets, or beyond 16 MiB, where the toolchain falls back to a
// literal pool).
func arm64SplitOffset(off int64, scale int64) (addImm, addShift uint32, access int64, ok bool) {
if off < 0 {
return 0, 0, 0, false
}
// Plain ADD: bring the base to within the largest scaled access.
l := min(off, 4095*scale)
l -= l % scale
if a := off - l; a <= 4095 {
return uint32(a), 0, l, true
}
// Shifted ADD: cover everything but the bits the access imm12 carries.
rest := off &^ (0xFFF * scale)
if rest >= 0 && rest>>12 <= 4095 {
return uint32(rest >> 12), 1, off - rest, true
}
return 0, 0, 0, false
} }
// ---- static symbol references (ADRP + offset) ---- // ---- static symbol references (ADRP + offset) ----
@@ -839,7 +1038,9 @@ func encodeARM64SBAddr(sym *ast.Symbol, rd int, relocs *[]Reloc) []byte {
) )
} }
// encodeARM64SBLoad emits ADRP R20, 0; LDR Rd, [R20, 0] with relocations. // encodeARM64SBLoad emits ADRP R27, 0; LDR Rd, [R27, 0] with relocations,
// matching the toolchain: the scratch register is REGTMP (R27) and the pair
// carries R_ARM64_PCREL_LDST64.
func encodeARM64SBLoad(sym *ast.Symbol, rd int, mnem string, relocs *[]Reloc) ([]byte, error) { func encodeARM64SBLoad(sym *ast.Symbol, rd int, mnem string, relocs *[]Reloc) ([]byte, error) {
lt, ok := a64LoadTable[mnem] lt, ok := a64LoadTable[mnem]
if !ok { if !ok {
@@ -847,17 +1048,17 @@ func encodeARM64SBLoad(sym *ast.Symbol, rd int, mnem string, relocs *[]Reloc) ([
} }
if relocs != nil { if relocs != nil {
*relocs = append(*relocs, *relocs = append(*relocs,
Reloc{Off: 0, After: 0, Name: sym.Name, Kind: RelArm64Addr, Addend: sym.Offset}, Reloc{Off: 0, After: 8, Name: sym.Name, Kind: RelArm64LDST64, Addend: sym.Offset},
Reloc{Off: 4, After: 4, Name: sym.Name, Kind: RelArm64Addr, Addend: sym.Offset},
) )
} }
return a64WordsLE( return a64WordsLE(
a64ADR(1, 0, 0, 20), // ADRP R20, 0 a64ADR(1, 0, 0, 27), // ADRP R27, 0
a64LSU(uint32(lt.size), uint32(lt.V), uint32(lt.opc), 0, 20, uint32(rd)), // LDR Rd, [R20, #0] a64LSU(uint32(lt.size), uint32(lt.V), uint32(lt.opc), 0, 27, uint32(rd)), // LDR Rd, [R27, #0]
), nil ), nil
} }
// encodeARM64SBStore emits ADRP R20, 0; STR Rs, [R20, 0] with relocations. // encodeARM64SBStore emits ADRP R27, 0; STR Rs, [R27, 0] with relocations,
// matching the toolchain's R27 scratch and R_ARM64_PCREL_LDST64 pair.
func encodeARM64SBStore(sym *ast.Symbol, rs int, mnem string, relocs *[]Reloc) ([]byte, error) { func encodeARM64SBStore(sym *ast.Symbol, rs int, mnem string, relocs *[]Reloc) ([]byte, error) {
lt, ok := a64LoadTable[mnem] lt, ok := a64LoadTable[mnem]
if !ok { if !ok {
@@ -866,13 +1067,12 @@ func encodeARM64SBStore(sym *ast.Symbol, rs int, mnem string, relocs *[]Reloc) (
storeOpc := a64StoreOpc(lt) storeOpc := a64StoreOpc(lt)
if relocs != nil { if relocs != nil {
*relocs = append(*relocs, *relocs = append(*relocs,
Reloc{Off: 0, After: 0, Name: sym.Name, Kind: RelArm64Addr, Addend: sym.Offset}, Reloc{Off: 0, After: 8, Name: sym.Name, Kind: RelArm64LDST64, Addend: sym.Offset},
Reloc{Off: 4, After: 4, Name: sym.Name, Kind: RelArm64Addr, Addend: sym.Offset},
) )
} }
return a64WordsLE( return a64WordsLE(
a64ADR(1, 0, 0, 20), // ADRP R20, 0 a64ADR(1, 0, 0, 27), // ADRP R27, 0
a64LSU(uint32(lt.size), uint32(lt.V), uint32(storeOpc), 0, 20, uint32(rs)), // STR Rs, [R20, #0] a64LSU(uint32(lt.size), uint32(lt.V), uint32(storeOpc), 0, 27, uint32(rs)), // STR Rs, [R27, #0]
), nil ), nil
} }
@@ -891,12 +1091,15 @@ func arm64Imm64(op *ast.Operand) int64 {
} }
// arm64MemWithFrame resolves a memory operand, translating FP/SP pseudo- // arm64MemWithFrame resolves a memory operand, translating FP/SP pseudo-
// registers via the frame mapping. // registers via the frame mapping. The offset stays 64-bit: the AST carries
func arm64MemWithFrame(op *ast.Operand, fi arm64FrameInfo) (rn int, off int32) { // int64 displacements and truncating here would wrap offsets beyond 2^31
// silently.
func arm64MemWithFrame(op *ast.Operand, fi arm64FrameInfo) (rn int, off int64) {
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo != "" { if op.Addr.Sym != nil && op.Addr.Sym.Pseudo != "" {
return arm64ResolvePseudo(op.Addr.Sym, fi) base, pseudo := arm64ResolvePseudo(op.Addr.Sym, fi)
return base, int64(pseudo)
} }
return arm64RegNum(op.Addr.Base), int32(op.Addr.Offset) return arm64RegNum(op.Addr.Base), op.Addr.Offset
} }
// arm64Label returns the label name of an operand. // arm64Label returns the label name of an operand.
@@ -972,7 +1175,7 @@ func encodeARM64FPCmp(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, e
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
} }
// Check if first operand is #0 (compare with zero): FCMP $0.0, Fn. // Check if first operand is #0 (compare with zero): FCMP $0.0, Fn.
if isImmOperand(ops[0]) && immFromOperand(ops[0]) == 0 { if isImmOperand(ops[0]) && arm64Imm64(ops[0]) == 0 {
rn := arm64RegNum(operandRegName(ops[1])) rn := arm64RegNum(operandRegName(ops[1]))
if rn < 0 { if rn < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem) return nil, fmt.Errorf("invalid register operand in %s", mnem)
@@ -1009,8 +1212,11 @@ func encodeARM64FPCCmp(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte,
if rm < 0 || rn < 0 { if rm < 0 || rn < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem) return nil, fmt.Errorf("invalid register operand in %s", mnem)
} }
nzcv := uint32(immFromOperand(ops[3])) nzcv := arm64Imm64(ops[3])
return a64wordLE(baseOp | uint32(rm)<<16 | cond<<12 | uint32(rn)<<5 | nzcv&0xF), nil if nzcv < 0 || nzcv > 0xF {
return nil, fmt.Errorf("%s: nzcv %d out of range (0..15)", mnem, nzcv)
}
return a64wordLE(baseOp | uint32(rm)<<16 | cond<<12 | uint32(rn)<<5 | uint32(nzcv)&0xF), nil
} }
// encodeARM64FPSel encodes a FP conditional select. // encodeARM64FPSel encodes a FP conditional select.
@@ -1092,7 +1298,7 @@ func encodeARM64CSEL(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, er
return a64wordLE(baseOp | uint32(rn)<<16 | invCond<<12 | uint32(rn)<<5 | uint32(rd)), nil return a64wordLE(baseOp | uint32(rn)<<16 | invCond<<12 | uint32(rn)<<5 | uint32(rd)), nil
} }
// CSEL cond, Rn, Rm, Rd (4 operands) — condition first. // CSEL cond, Rn, Rm, Rd (4 operands), condition first.
// Go assembler syntax: CSEL cond, Rn, Rm, Rd // Go assembler syntax: CSEL cond, Rn, Rm, Rd
// ARM64 encoding: Rm in bits[20:16], Rn in bits[9:5], Rd in bits[4:0]. // ARM64 encoding: Rm in bits[20:16], Rn in bits[9:5], Rd in bits[4:0].
if len(ops) != 4 { if len(ops) != 4 {
@@ -1137,32 +1343,94 @@ func encodeARM64CRC32(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, e
// ---- Atomics encoding ---- // ---- Atomics encoding ----
// encodeARM64Excl encodes an exclusive load/store instruction. // arm64ExclMem resolves the memory operand of an exclusive or atomic
// LDXR (Rn), Rt → LDXR Rt, [Rn] (2 operands: mem, reg or reg, mem) // instruction. These encodings have no immediate field: the toolchain
// STXR Rs, (Rn), Rt → STXR Rs, Rt, [Rn] (3 operands: Rs, mem, Rt-status) // rejects `LDXR 8(R1), R2` as an illegal combination, so a non-zero offset is
// reported rather than silently dropped (which would read the wrong address).
func arm64ExclMem(mnem string, op *ast.Operand) (int, error) {
rn, off := arm64MemWithFrame(op, arm64FrameInfo{})
if rn < 0 {
return 0, fmt.Errorf("invalid memory operand in %s", mnem)
}
if off != 0 {
return 0, fmt.Errorf("%s: offset %d not supported, exclusive and atomic accesses take a plain (Rn) operand", mnem, off)
}
return rn, nil
}
// arm64PairOf parses a register-pair operand `(R1, R2)`, reporting false
// when the operand is not a pair. The toolchain takes the second register of
// the pair from the operand's Offset (its C_PAIR class,
// cmd/internal/obj/arm64/asm7.go cases 58/59).
func arm64PairOf(op *ast.Operand) (int, int, bool) {
raw := strings.TrimSpace(op.Raw)
if !strings.HasPrefix(raw, "(") || !strings.HasSuffix(raw, ")") {
return -1, -1, false
}
parts := strings.Split(raw[1:len(raw)-1], ",")
if len(parts) != 2 {
return -1, -1, false
}
r1 := arm64RegNum(strings.TrimSpace(parts[0]))
r2 := arm64RegNum(strings.TrimSpace(parts[1]))
if r1 < 0 || r2 < 0 {
return -1, -1, false
}
return r1, r2, true
}
// encodeARM64Excl encodes the exclusive load/store family with the operand
// order the toolchain parses (cmd/internal/obj/arm64/asm7.go cases 58 and 59,
// and its own spellings in arm64enc.s):
//
// STXR Rt, (Rn), Rs store, single register
// STXP (Rt1, Rt2), (Rn), Rs store, register pair
// LDXR (Rn), Rt load, single register
// LDXP (Rn), (Rt1, Rt2) load, register pair
//
// Decoded toolchain evidence: `STXR R1, (R2), R3` assembles to 0xc8037c41,
// whose fields are Rs=3, Rn=2, Rt=1: the FIRST register operand is the data
// register and the LAST the status register.
func encodeARM64Excl(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) { func encodeARM64Excl(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
// LDXR/STXR have different operand forms.
isLoad := strings.HasPrefix(mnem, "LD") isLoad := strings.HasPrefix(mnem, "LD")
if isLoad { if isLoad {
// LDXR (Rn), Rt → 2 operands: mem, reg // LDXR (Rn), Rt / LDXP (Rn), (Rt1, Rt2): 2 operands.
if len(ops) != 2 { if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
} }
rn, _ := arm64MemWithFrame(ops[0], arm64FrameInfo{}) rn, err := arm64ExclMem(mnem, ops[0])
if err != nil {
return nil, err
}
if rt1, rt2, ok := arm64PairOf(ops[1]); ok {
// The single-register opcodes pre-set the unused Rs (bits 20:16)
// and Rt2 (bits 14:10) fields to 31; the pair forms carry a real
// Rt2 and keep Rs at 31.
return a64wordLE(baseOp | 0x1F<<16 | uint32(rt2)<<10 | uint32(rn)<<5 | uint32(rt1)), nil
}
rt := arm64RegNum(operandRegName(ops[1])) rt := arm64RegNum(operandRegName(ops[1]))
if rn < 0 || rt < 0 { if rt < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem) return nil, fmt.Errorf("invalid operand in %s", mnem)
} }
return a64wordLE(baseOp | uint32(rn)<<5 | uint32(rt)), nil return a64wordLE(baseOp | uint32(rn)<<5 | uint32(rt)), nil
} }
// STXR Rs, (Rn), Rt → 3 operands: Rs, mem, Rt // STXR Rt, (Rn), Rs / STXP (Rt1, Rt2), (Rn), Rs: 3 operands.
if len(ops) != 3 { if len(ops) != 3 {
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
} }
rs := arm64RegNum(operandRegName(ops[0])) rn, err := arm64ExclMem(mnem, ops[1])
rn, _ := arm64MemWithFrame(ops[1], arm64FrameInfo{}) if err != nil {
rt := arm64RegNum(operandRegName(ops[2])) return nil, err
if rs < 0 || rn < 0 || rt < 0 { }
rs := arm64RegNum(operandRegName(ops[2]))
if rs < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem)
}
if rt1, rt2, ok := arm64PairOf(ops[0]); ok {
return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rt2)<<10 | uint32(rn)<<5 | uint32(rt1)), nil
}
rt := arm64RegNum(operandRegName(ops[0]))
if rt < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem) return nil, fmt.Errorf("invalid operand in %s", mnem)
} }
return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rn)<<5 | uint32(rt)), nil return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rn)<<5 | uint32(rt)), nil
@@ -1176,9 +1444,15 @@ func encodeARM64LSEAtom(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte,
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
} }
rs := arm64RegNum(operandRegName(ops[0])) rs := arm64RegNum(operandRegName(ops[0]))
rn, _ := arm64MemWithFrame(ops[1], arm64FrameInfo{}) if rs < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem)
}
rn, err := arm64ExclMem(mnem, ops[1])
if err != nil {
return nil, err
}
rt := arm64RegNum(operandRegName(ops[2])) rt := arm64RegNum(operandRegName(ops[2]))
if rs < 0 || rn < 0 || rt < 0 { if rt < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem) return nil, fmt.Errorf("invalid operand in %s", mnem)
} }
return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rn)<<5 | uint32(rt)), nil return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rn)<<5 | uint32(rt)), nil
@@ -1187,43 +1461,25 @@ func encodeARM64LSEAtom(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte,
// ---- Bitfield/EXTR encoding ---- // ---- Bitfield/EXTR encoding ----
// encodeARM64Bitfield encodes a bitfield instruction. // encodeARM64Bitfield encodes a bitfield instruction.
// ASR/LSL/LSR/ROR $shamt, Rn, Rd → 3 operands: $imm, Rn, Rd
// BFI/BFXIL/SBFM/UBFM $immr, Rn, $imms, Rd → 4 operands // BFI/BFXIL/SBFM/UBFM $immr, Rn, $imms, Rd → 4 operands
func encodeARM64Bitfield(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) { func encodeARM64Bitfield(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
isShift := mnem == "ASR" || mnem == "ASRW" || mnem == "LSL" || mnem == "LSLW" ||
mnem == "LSR" || mnem == "LSRW" || mnem == "ROR" || mnem == "RORW"
if isShift {
// ASR $shamt, Rn, Rd → SBFM with immr=shamt, imms=31/63
if len(ops) != 3 {
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
}
shamt := int(immFromOperand(ops[0]))
rn := arm64RegNum(operandRegName(ops[1]))
rd := arm64RegNum(operandRegName(ops[2]))
if rn < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem)
}
// ASR: SBFM with immr=shamt, imms=31(32-bit) or 63(64-bit)
is64 := mnem == "ASR"
imms := 31
if is64 {
imms = 63
}
return a64wordLE(baseOp | uint32(shamt)<<16 | uint32(imms)<<10 | uint32(rn)<<5 | uint32(rd)), nil
}
// BFI/BFXIL/SBFM/UBFM: 4 operands ($immr, Rn, $imms, Rd) // BFI/BFXIL/SBFM/UBFM: 4 operands ($immr, Rn, $imms, Rd)
if len(ops) != 4 { if len(ops) != 4 {
return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
} }
immr := int(immFromOperand(ops[0])) immr := arm64Imm64(ops[0])
rn := arm64RegNum(operandRegName(ops[1])) rn := arm64RegNum(operandRegName(ops[1]))
imms := int(immFromOperand(ops[2])) imms := arm64Imm64(ops[2])
rd := arm64RegNum(operandRegName(ops[3])) rd := arm64RegNum(operandRegName(ops[3]))
if rn < 0 || rd < 0 { if rn < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem) return nil, fmt.Errorf("invalid register operand in %s", mnem)
} }
// The toolchain rejects bit numbers at or above the operand width, which
// sf (bit 31 of the base) selects: 64 when set, 32 otherwise.
width := uint32(32) << (baseOp >> 31 & 1)
if immr < 0 || uint32(immr) >= width || imms < 0 || uint32(imms) >= width {
return nil, fmt.Errorf("%s: bit number out of range (immr=%d imms=%d, width=%d)", mnem, immr, imms, width)
}
return a64wordLE(baseOp | uint32(immr)<<16 | uint32(imms)<<10 | uint32(rn)<<5 | uint32(rd)), nil return a64wordLE(baseOp | uint32(immr)<<16 | uint32(imms)<<10 | uint32(rn)<<5 | uint32(rd)), nil
} }
@@ -1233,13 +1489,19 @@ func encodeARM64Extr(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, er
if len(ops) != 4 { if len(ops) != 4 {
return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
} }
lsb := int(immFromOperand(ops[0])) lsb := arm64Imm64(ops[0])
rm := arm64RegNum(operandRegName(ops[1])) rm := arm64RegNum(operandRegName(ops[1]))
rn := arm64RegNum(operandRegName(ops[2])) rn := arm64RegNum(operandRegName(ops[2]))
rd := arm64RegNum(operandRegName(ops[3])) rd := arm64RegNum(operandRegName(ops[3]))
if rm < 0 || rn < 0 || rd < 0 { if rm < 0 || rn < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand in %s", mnem) return nil, fmt.Errorf("invalid register operand in %s", mnem)
} }
// The imms field is 6 bits and must stay below the operand width, which
// sf (bit 31 of the base) selects: 64 when set, 32 otherwise.
width := int64(32) << (baseOp >> 31 & 1)
if lsb < 0 || lsb >= width {
return nil, fmt.Errorf("%s: bit number %d out of range (width=%d)", mnem, lsb, width)
}
return a64wordLE(baseOp | uint32(rm)<<16 | uint32(lsb)<<10 | uint32(rn)<<5 | uint32(rd)), nil return a64wordLE(baseOp | uint32(rm)<<16 | uint32(lsb)<<10 | uint32(rn)<<5 | uint32(rd)), nil
} }
@@ -1271,7 +1533,7 @@ func AssembleFileARM64(f *ast.File) (*Image, error) {
return nil, err return nil, err
} }
img := &Image{Symbols: map[string]int{}} img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls { for _, d := range f.Decls {
t, ok := d.(*ast.Text) t, ok := d.(*ast.Text)
if !ok { if !ok {
+80 -74
View File
@@ -9,7 +9,7 @@ package asm
// an opcode constant, and the format selects the bit layout. The opcode // an opcode constant, and the format selects the bit layout. The opcode
// constants and formats are transcribed from the Go toolchain's own arm64 // constants and formats are transcribed from the Go toolchain's own arm64
// backend (cmd/internal/obj/arm64), so the emitted bytes match `go tool asm` // backend (cmd/internal/obj/arm64), so the emitted bytes match `go tool asm`
// exactly — the ground-truth oracle for the verify suite. // exactly, the ground-truth oracle for the verify suite.
// //
// All AArch64 instructions are 32 bits, little-endian. The formats used here // All AArch64 instructions are 32 bits, little-endian. The formats used here
// (per the ARM Architecture Reference Manual): // (per the ARM Architecture Reference Manual):
@@ -27,8 +27,10 @@ package asm
// Uncond-branch 0x6B<<25 | opc<<21 | Rn<<5 | Rd (BR/BLR/RET) // Uncond-branch 0x6B<<25 | opc<<21 | Rn<<5 | Rd (BR/BLR/RET)
// ADR/ADRP p<<31 | 0x10<<24 | immlo<<29 | immhi<<5 | Rd // ADR/ADRP p<<31 | 0x10<<24 | immlo<<29 | immhi<<5 | Rd
import "maps"
// arm64RegNum returns the 5-bit register number for an AArch64 register name: // arm64RegNum returns the 5-bit register number for an AArch64 register name:
// R0–R30 (integer), F0–F31 (floating point), and the ABI aliases the // R0-R30 (integer), F0-F31 (floating point), and the ABI aliases the
// runtime's assembly uses. Returns -1 for an unrecognised name. // runtime's assembly uses. Returns -1 for an unrecognised name.
func arm64RegNum(name string) int { func arm64RegNum(name string) int {
switch name { switch name {
@@ -96,10 +98,14 @@ func arm64RegNum(name string) int {
return 30 return 30
case "R31", "ZR": case "R31", "ZR":
return 31 return 31
case "SP": case "SP", "RSP":
return 31 // SP and ZR share encoding 31; context determines meaning // RSP is the toolchain's spelling for register 31 (it rejects
// R31 in an operand); SP stays for sources that spell it the
// amd64 way. SP and ZR share encoding 31; context determines
// the meaning.
return 31
} }
// F0–F31. // F0-F31.
if len(name) >= 1 && name[0] == 'F' { if len(name) >= 1 && name[0] == 'F' {
n := 0 n := 0
for i := 1; i < len(name); i++ { for i := 1; i < len(name); i++ {
@@ -262,19 +268,15 @@ type a64Format uint8
const ( const (
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc. a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
a64FDPIR // data-processing (immediate): ADD/SUB $imm
a64FLogImm // logical (immediate): AND/ORR/EOR $imm
a64FMovWide // move wide: MOVZ, MOVN, MOVK a64FMovWide // move wide: MOVZ, MOVN, MOVK
a64FLSU // load/store (unsigned immediate, scaled)
a64FLSUnscaled // load/store (unscaled immediate)
a64FLSPair // load/store pair
a64FBranch // unconditional branch (B/BL) a64FBranch // unconditional branch (B/BL)
a64FBranchCond // conditional branch (B.cond) a64FBranchCond // conditional branch (B.cond)
a64FUncondBranch // unconditional branch register (BR/BLR/RET) a64FUncondBranch // unconditional branch register (BR/BLR/RET)
a64FADR // ADR/ADRP a64FADR // ADR/ADRP
a64FEXTR // EXTR a64FEXTR // EXTR
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
a64FSystem // system: NOP, BRK, etc. a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc. a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT* a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc. a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
@@ -282,10 +284,9 @@ const (
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc. a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FFMovGR // FMOV between GP and FP registers
a64FCRC32 // CRC32 a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FSIMD3 // SIMD 3-operand: VADD, VSUB, VMUL a64FSIMD3 // SIMD 3-operand: VADD, VSUB, VMUL
) )
@@ -332,30 +333,14 @@ func init() {
"ANDSW": 0<<31 | 3<<29 | 0x0a<<24, "ANDSW": 0<<31 | 3<<29 | 0x0a<<24,
"BICS": 1<<31 | 3<<29 | 0x0a<<24 | 1<<21, "BICS": 1<<31 | 3<<29 | 0x0a<<24 | 1<<21,
"BICSW": 0<<31 | 3<<29 | 0x0a<<24 | 1<<21, "BICSW": 0<<31 | 3<<29 | 0x0a<<24 | 1<<21,
// Shift // Divide (data-processing 2 source): the opcode occupies bits 15:10
"LSL": 1<<31 | 0<<29 | 0x0a<<24, // alias of UBFM // of the 0xd6<<21 fixed field, UDIV=0b0010 and SDIV=0b0011 (ARM ARM
"LSLW": 0<<31 | 0<<29 | 0x0a<<24, // "Data-processing (2 source)"; the toolchain spells them OPDP2(2)
"LSR": 1<<31 | 0<<29 | 0x0a<<24, // and OPDP2(3)). sf=1 selects the X forms.
"LSRW": 0<<31 | 0<<29 | 0x0a<<24, "SDIV": 1<<31 | 0xd6<<21 | 3<<10,
"ASR": 1<<31 | 0<<29 | 0x0a<<24, "SDIVW": 0<<31 | 0xd6<<21 | 3<<10,
"ASRW": 0<<31 | 0<<29 | 0x0a<<24, "UDIV": 1<<31 | 0xd6<<21 | 2<<10,
"ROR": 1<<31 | 0<<29 | 0x0a<<24, "UDIVW": 0<<31 | 0xd6<<21 | 2<<10,
"RORW": 0<<31 | 0<<29 | 0x0a<<24,
// Multiply
"MADD": 1<<31 | 0<<29 | 0x1b<<24 | 0<<21,
"MADDW": 0<<31 | 0<<29 | 0x1b<<24 | 0<<21,
"MSUB": 1<<31 | 0<<29 | 0x1b<<24 | 1<<21,
"MSUBW": 0<<31 | 0<<29 | 0x1b<<24 | 1<<21,
// Divide
"SDIV": 1<<31 | 0<<29 | 0x0d<<24,
"SDIVW": 0<<31 | 0<<29 | 0x0d<<24,
"UDIV": 1<<31 | 0<<29 | 0x0d<<24 | 1<<10,
"UDIVW": 0<<31 | 0<<29 | 0x0d<<24 | 1<<10,
// CRC
"CRC32B": 0<<31 | 0<<29 | 0x1b<<24 | 4<<10,
"CRC32H": 0<<31 | 0<<29 | 0x1b<<24 | 5<<10,
"CRC32W": 0<<31 | 0<<29 | 0x1b<<24 | 6<<10,
"CRC32X": 1<<31 | 0<<29 | 0x1b<<24 | 7<<10,
// Conditional select // Conditional select
"CSEL": 1<<31 | 0<<29 | 0x1d<<24 | 0<<10, "CSEL": 1<<31 | 0<<29 | 0x1d<<24 | 0<<10,
"CSELW": 0<<31 | 0<<29 | 0x1d<<24 | 0<<10, "CSELW": 0<<31 | 0<<29 | 0x1d<<24 | 0<<10,
@@ -385,14 +370,37 @@ func init() {
a64InstrTable["MOV"] = a64Enc{format: a64FDPSR, op: dpsr["ORR"]} a64InstrTable["MOV"] = a64Enc{format: a64FDPSR, op: dpsr["ORR"]}
a64InstrTable["MOVW"] = a64Enc{format: a64FDPSR, op: dpsr["ORRW"]} a64InstrTable["MOVW"] = a64Enc{format: a64FDPSR, op: dpsr["ORRW"]}
// ---- data-processing (immediate) ---- // ---- shifts ----
// ADD/SUB $imm, Rn, Rd // The mnemonic serves both forms: with an immediate the aliases of the
a64InstrTable["ADDImm"] = a64Enc{format: a64FDPIR, op: 1<<31 | 0<<30 | 0<<29 | 0x11<<24} // data-processing (immediate) group apply (ARM ARM "Shifts"), with a
a64InstrTable["ADDWImm"] = a64Enc{format: a64FDPIR, op: 0<<31 | 0<<30 | 0<<29 | 0x11<<24} // register the data-processing (2 source) LSLV/LSRV/ASRV/RORV. The op
a64InstrTable["SUBImm"] = a64Enc{format: a64FDPIR, op: 1<<31 | 1<<30 | 0<<29 | 0x11<<24} // field carries the immediate-alias base; encodeARM64Shift derives both
a64InstrTable["SUBWImm"] = a64Enc{format: a64FDPIR, op: 0<<31 | 1<<30 | 0<<29 | 0x11<<24} // it and the two-source opcode. Identities, W = 64 (X) or 32 (W):
a64InstrTable["ADDSImm"] = a64Enc{format: a64FDPIR, op: 1<<31 | 0<<30 | 1<<29 | 0x11<<24} //
a64InstrTable["SUBSImm"] = a64Enc{format: a64FDPIR, op: 1<<31 | 1<<30 | 1<<29 | 0x11<<24} // LSL $sh, Rn, Rd = UBFM Rd, Rn, #(-sh) mod W, #(W-1)-sh
// LSR $sh, Rn, Rd = UBFM Rd, Rn, #sh, #(W-1)
// ASR $sh, Rn, Rd = SBFM Rd, Rn, #sh, #(W-1)
// ROR $sh, Rn, Rd = EXTR Rd, Rn, Rn, #sh
shifts := map[string]a64Enc{
"LSL": {format: a64FShift, op: 1<<31 | 2<<29 | 0x26<<23 | 1<<22}, // UBFM X
"LSLW": {format: a64FShift, op: 0<<31 | 2<<29 | 0x26<<23 | 0<<22}, // UBFM W
"LSR": {format: a64FShift, op: 1<<31 | 2<<29 | 0x26<<23 | 1<<22}, // UBFM X
"LSRW": {format: a64FShift, op: 0<<31 | 2<<29 | 0x26<<23 | 0<<22}, // UBFM W
"ASR": {format: a64FShift, op: 1<<31 | 0<<29 | 0x26<<23 | 1<<22}, // SBFM X
"ASRW": {format: a64FShift, op: 0<<31 | 0<<29 | 0x26<<23 | 0<<22}, // SBFM W
"ROR": {format: a64FShift, op: 1<<31 | 0x27<<23 | 1<<22}, // EXTR X
"RORW": {format: a64FShift, op: 0<<31 | 0x27<<23 | 0<<22}, // EXTR W
}
maps.Copy(a64InstrTable, shifts)
// ---- multiply accumulate ----
// MADD/MSUB Rm, Ra, Rn, Rd: sf 00 11011 o0(15) Rm Ra Rn Rd. The
// toolchain's optab has no shorter row, so all four operands are
// mandatory, and Ra is the SECOND operand.
a64InstrTable["MADD"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24}
a64InstrTable["MADDW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24}
a64InstrTable["MSUB"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<15}
a64InstrTable["MSUBW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24 | 1<<15}
// ---- move wide ---- // ---- move wide ----
// MOVZ/MOVN/MOVK // MOVZ/MOVN/MOVK
@@ -407,22 +415,9 @@ func init() {
a64InstrTable["ADR"] = a64Enc{format: a64FADR, op: 0} a64InstrTable["ADR"] = a64Enc{format: a64FADR, op: 0}
a64InstrTable["ADRP"] = a64Enc{format: a64FADR, op: 1} a64InstrTable["ADRP"] = a64Enc{format: a64FADR, op: 1}
// ---- load/store (unsigned immediate) ---- // Load/store mnemonics never enter this table: the MOV pseudo-instruction
a64InstrTable["MOVD"] = a64Enc{format: a64FLSU, op: 3<<30 | 7<<27 | 1<<22} // LDR 64-bit // dispatch handles them through a64LoadTable, which also carries the store
a64InstrTable["MOVWU"] = a64Enc{format: a64FLSU, op: 2<<30 | 7<<27 | 1<<22} // LDR 32-bit unsigned // opcode (integer and FP stores both use opc=00, differing only in V).
a64InstrTable["MOVHU"] = a64Enc{format: a64FLSU, op: 1<<30 | 7<<27 | 1<<22} // LDRH unsigned
a64InstrTable["MOVBU"] = a64Enc{format: a64FLSU, op: 0<<30 | 7<<27 | 1<<22} // LDRB unsigned
a64InstrTable["MOVW"] = a64Enc{format: a64FLSU, op: 2<<30 | 7<<27 | 2<<22} // LDRSW (signed 32→64)
a64InstrTable["MOVH"] = a64Enc{format: a64FLSU, op: 1<<30 | 7<<27 | 2<<22} // LDRSH (signed half)
a64InstrTable["MOVB"] = a64Enc{format: a64FLSU, op: 0<<30 | 7<<27 | 2<<22} // LDRSB (signed byte)
a64InstrTable["FMOVS"] = a64Enc{format: a64FLSU, op: 2<<30 | 7<<27 | 1<<26 | 1<<22} // FLDR 32-bit FP
a64InstrTable["FMOVD"] = a64Enc{format: a64FLSU, op: 3<<30 | 7<<27 | 1<<26 | 1<<22} // FLDR 64-bit FP
// Store opcodes (load ^ (1<<22)):
// STR 64-bit: size=3, V=0, opc=00 → 3<<30 | 7<<27 | 0<<22
// STR 32-bit: size=2, V=0, opc=00 → 2<<30 | 7<<27 | 0<<22
// STRH: size=1, V=0, opc=00 → 1<<30 | 7<<27 | 0<<22
// STRB: size=0, V=0, opc=00 → 0<<30 | 7<<27 | 0<<22
// ---- branches ---- // ---- branches ----
a64InstrTable["B"] = a64Enc{format: a64FBranch, op: 0<<31 | 5<<26} a64InstrTable["B"] = a64Enc{format: a64FBranch, op: 0<<31 | 5<<26}
@@ -445,10 +440,8 @@ func init() {
a64InstrTable["RET"] = a64Enc{format: a64FUncondBranch, op: 0x6B<<25 | 2<<21} a64InstrTable["RET"] = a64Enc{format: a64FUncondBranch, op: 0x6B<<25 | 2<<21}
// ---- system ---- // ---- system ----
a64InstrTable["NOP"] = a64Enc{format: a64FSystem, op: a64NOP} // NOP/NOOP/UNDEF are spelled out in encodeARM64Instr's pseudo switch,
a64InstrTable["NOOP"] = a64Enc{format: a64FSystem, op: a64NOP} // so they carry no table entry; a64NOP and a64BRK are the encoders.
a64InstrTable["BRK"] = a64Enc{format: a64FSystem, op: 0xd4200000}
a64InstrTable["UNDEF"] = a64Enc{format: a64FSystem, op: a64BRK(0)}
// ---- EXTR ---- // ---- EXTR ----
a64InstrTable["EXTR"] = a64Enc{format: a64FEXTR, op: 1<<31 | 0x27<<23 | 1<<22} a64InstrTable["EXTR"] = a64Enc{format: a64FEXTR, op: 1<<31 | 0x27<<23 | 1<<22}
@@ -549,8 +542,8 @@ func init() {
a64InstrTable[m] = a64Enc{format: a64FFPCvt, op: op} a64InstrTable[m] = a64Enc{format: a64FFPCvt, op: op}
} }
// ---- FMOV between GP and FP registers ---- // FMOV between GP and FP registers needs no table entry: the MOV
a64InstrTable["FMOVGR"] = a64Enc{format: a64FFMovGR, op: 0x1e260000} // placeholder, actual encoding depends on direction // pseudo-instruction dispatches it by operand class (encodeARM64RegMove).
// ---- conditional select: CSEL, CSINC, CSINV, CSNEG ---- // ---- conditional select: CSEL, CSINC, CSINV, CSNEG ----
csel := map[string]uint32{ csel := map[string]uint32{
@@ -586,6 +579,10 @@ func init() {
} }
// ---- exclusive load/store ---- // ---- exclusive load/store ----
// Single-register forms pre-set the unused Rs and Rt2 fields to 31 (the
// 0x7c00/0x1f0000 halves of the constants below); the register-pair
// forms carry a real Rt2 in bits 14:10, so their opcodes pre-set
// neither field.
a64InstrTable["LDXR"] = a64Enc{format: a64FExcl, op: 0xc85f7c00} a64InstrTable["LDXR"] = a64Enc{format: a64FExcl, op: 0xc85f7c00}
a64InstrTable["LDXRB"] = a64Enc{format: a64FExcl, op: 0x085f7c00} a64InstrTable["LDXRB"] = a64Enc{format: a64FExcl, op: 0x085f7c00}
a64InstrTable["LDXRH"] = a64Enc{format: a64FExcl, op: 0x485f7c00} a64InstrTable["LDXRH"] = a64Enc{format: a64FExcl, op: 0x485f7c00}
@@ -594,6 +591,12 @@ func init() {
a64InstrTable["LDAXRB"] = a64Enc{format: a64FExcl, op: 0x085ffc00} a64InstrTable["LDAXRB"] = a64Enc{format: a64FExcl, op: 0x085ffc00}
a64InstrTable["LDAXRH"] = a64Enc{format: a64FExcl, op: 0x485ffc00} a64InstrTable["LDAXRH"] = a64Enc{format: a64FExcl, op: 0x485ffc00}
a64InstrTable["LDAXRW"] = a64Enc{format: a64FExcl, op: 0x885ffc00} a64InstrTable["LDAXRW"] = a64Enc{format: a64FExcl, op: 0x885ffc00}
// Pair loads, LDSTX(sz, 0, l=1, o1=1, o0) in asm7.go: LDXP/ LDXPW have
// o0=0, LDAXP/LDAXPW o0=1 (bit 15). Rs (bits 20:16) stays 31.
a64InstrTable["LDXP"] = a64Enc{format: a64FExcl, op: 0xc8600000}
a64InstrTable["LDXPW"] = a64Enc{format: a64FExcl, op: 0x88600000}
a64InstrTable["LDAXP"] = a64Enc{format: a64FExcl, op: 0xc8608000}
a64InstrTable["LDAXPW"] = a64Enc{format: a64FExcl, op: 0x88608000}
a64InstrTable["STXR"] = a64Enc{format: a64FExcl, op: 0xc8007c00} a64InstrTable["STXR"] = a64Enc{format: a64FExcl, op: 0xc8007c00}
a64InstrTable["STXRB"] = a64Enc{format: a64FExcl, op: 0x08007c00} a64InstrTable["STXRB"] = a64Enc{format: a64FExcl, op: 0x08007c00}
a64InstrTable["STXRH"] = a64Enc{format: a64FExcl, op: 0x48007c00} a64InstrTable["STXRH"] = a64Enc{format: a64FExcl, op: 0x48007c00}
@@ -602,6 +605,12 @@ func init() {
a64InstrTable["STLXRB"] = a64Enc{format: a64FExcl, op: 0x0800fc00} a64InstrTable["STLXRB"] = a64Enc{format: a64FExcl, op: 0x0800fc00}
a64InstrTable["STLXRH"] = a64Enc{format: a64FExcl, op: 0x4800fc00} a64InstrTable["STLXRH"] = a64Enc{format: a64FExcl, op: 0x4800fc00}
a64InstrTable["STLXRW"] = a64Enc{format: a64FExcl, op: 0x8800fc00} a64InstrTable["STLXRW"] = a64Enc{format: a64FExcl, op: 0x8800fc00}
// Pair stores, LDSTX(sz, 0, l=0, o1=1, o0): STXP/STXPW have o0=0,
// STLXP/STLXPW o0=1 (bit 15). Both Rs and Rt2 are real fields.
a64InstrTable["STXP"] = a64Enc{format: a64FExcl, op: 0xc8200000}
a64InstrTable["STXPW"] = a64Enc{format: a64FExcl, op: 0x88200000}
a64InstrTable["STLXP"] = a64Enc{format: a64FExcl, op: 0xc8208000}
a64InstrTable["STLXPW"] = a64Enc{format: a64FExcl, op: 0x88208000}
// ---- LSE atomics ---- // ---- LSE atomics ----
a64InstrTable["LDADDD"] = a64Enc{format: a64FLSE, op: 3<<30 | 0x1c1<<21 | 0x00<<10} a64InstrTable["LDADDD"] = a64Enc{format: a64FLSE, op: 3<<30 | 0x1c1<<21 | 0x00<<10}
@@ -642,14 +651,11 @@ var a64LoadTable = map[string]a64LSType{
"FMOVD": {3, 1, 1}, // LDR D (64-bit FP) "FMOVD": {3, 1, 1}, // LDR D (64-bit FP)
} }
// a64StoreOpc returns the store opc for a given load type. // a64StoreOpc returns the store opc for a given load type: integer and FP
// For integer: store opc = 00 (the load opc bits cleared). // stores both encode opc=00 (the load's signedness bit sits in opc[1], which
// For FP: store opc = 00 (same pattern). // the store form clears; FP registers are selected by V, not opc).
func a64StoreOpc(t a64LSType) int { func a64StoreOpc(t a64LSType) int {
if t.V == 1 { return 0
return 0 // FP store
}
return 0 // integer store
} }
// arm64RegClass discriminates integer (R), floating-point (F) registers for // arm64RegClass discriminates integer (R), floating-point (F) registers for
+416
View File
@@ -572,3 +572,419 @@ func leWords(b []byte) []uint32 {
} }
return w return w
} }
// TestArm64IndirectBranch pins the indirect branch forms in a leaf function:
// JMP (Rn) lowers to BR Rn, matching the toolchain's spelling, and the raw
// BR/BLR mnemonics encode directly (a gasm superset the toolchain's front
// end does not accept). CALL (Rn) shares the BLR path and its non-leaf
// prologue parity is covered by the ground-truth kernel.
func TestArm64IndirectBranch(t *testing.T) {
src := `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (R0)
BR R5
BLR R6
RET
`
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
want := []uint32{
0xd61f0000, // BR R0
0xd61f00a0, // BR R5
0xd63f00c0, // BLR R6
0xd65f03c0, // RET (BR LR)
}
got := leWords(img.Code)
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// arm64Words assembles a single NOSPLIT leaf body and returns its words.
func arm64Words(t *testing.T, body string) []uint32 {
t.Helper()
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
return leWords(img.Code)
}
// TestArm64ShiftEncodings pins the shift words against `go tool asm -S`
// output (Go 1.27, arm64): immediate forms alias SBFM/UBFM with ROR as EXTR,
// register forms are the two-source LSLV/LSRV/ASRV/RORV.
func TestArm64ShiftEncodings(t *testing.T) {
got := arm64Words(t, "\tLSL $4, R0, R1\n\tLSR $8, R0, R2\n\tASR $4, R0, R3\n\tROR $12, R0, R4\n"+
"\tLSLW $4, R0, R5\n\tLSRW $8, R0, R6\n\tASRW $4, R0, R7\n\tRORW $12, R0, R8\n")
want := []uint32{
0xd37cec01, // LSL $4 = UBFM X1, X0, #60, #59
0xd348fc02, // LSR $8 = UBFM X2, X0, #8, #63
0x9344fc03, // ASR $4 = SBFM X3, X0, #4, #63
0x93c03004, // ROR $12 = EXTR X4, X0, X0, #12
0x531c6c05, // LSLW $4 = UBFM W5, W0, #28, #27
0x53087c06, // LSRW $8 = UBFM W6, W0, #8, #31
0x13047c07, // ASRW $4 = SBFM W7, W0, #4, #31
0x13803008, // RORW $12 = EXTR W8, W0, W0, #12
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("imm shift word %d = %08x, want %08x", i, got[i], want[i])
}
}
got = arm64Words(t, "\tLSL R9, R0, R10\n\tLSR R9, R0, R11\n\tASR R9, R0, R12\n\tROR R9, R0, R13\n"+
"\tLSLW R9, R0, R14\n\tLSRW R9, R0, R15\n\tASRW R9, R0, R16\n\tRORW R9, R0, R17\n")
want = []uint32{
0x9ac9200a, // LSLV X10, X0, X9
0x9ac9240b, // LSRV X11, X0, X9
0x9ac9280c, // ASRV X12, X0, X9
0x9ac92c0d, // RORV X13, X0, X9
0x1ac9200e, // LSLV W14, W0, W9
0x1ac9240f, // LSRV W15, W0, W9
0x1ac92810, // ASRV W16, W0, W9
0x1ac92c11, // RORV W17, W0, W9
0xd65f03c0, // RET
}
for i := range want {
if got[i] != want[i] {
t.Errorf("reg shift word %d = %08x, want %08x", i, got[i], want[i])
}
}
// Two-operand spellings fold to Rn = Rd.
got = arm64Words(t, "\tLSL $4, R1\n\tLSR R9, R1\n\tASR $4, R1\n\tROR R9, R1\n\tLSLW $4, R1\n\tRORW R9, R1\n")
want = []uint32{
0xd37cec21, // LSL $4, R1 = UBFM X1, X1, #60, #59
0x9ac92421, // LSRV X1, X1, X9
0x9344fc21, // ASR $4, R1 = SBFM X1, X1, #4, #63
0x9ac92c21, // RORV X1, X1, X9
0x531c6c21, // LSLW $4, R1 = UBFM W1, W1, #28, #27
0x1ac92c21, // RORV W1, W1, W9
0xd65f03c0, // RET
}
for i := range want {
if got[i] != want[i] {
t.Errorf("2op shift word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64ShiftRangeErrors: the toolchain reports "illegal bit number" for
// shift amounts at or above the operand width.
func TestArm64ShiftRangeErrors(t *testing.T) {
for _, src := range []string{
"\tLSL $64, R0, R1\n",
"\tLSRW $32, R0, R1\n",
"\tRORW $32, R0, R1\n",
"\tASR $-1, R0, R1\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+src+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", src)
}
}
}
// TestArm64DivEncodings pins SDIV/UDIV in both widths: the 2-source opcode
// field (bits 15:10 of the 0xd6<<21 fixed field) is UDIV=0b0010, SDIV=0b0011.
func TestArm64DivEncodings(t *testing.T) {
got := arm64Words(t, "\tSDIV R1, R2, R3\n\tUDIV R1, R2, R3\n\tSDIVW R1, R2, R3\n\tUDIVW R1, R2, R3\n")
want := []uint32{
0x9ac10c43, // SDIV X3, X2, X1
0x9ac10843, // UDIV X3, X2, X1
0x1ac10c43, // SDIV W3, W2, W1
0x1ac10843, // UDIV W3, W2, W1
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("div word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64MAddSub pins the four-operand MADD/MSUB words (Rm, Ra, Rn, Rd,
// with Ra in bits 14:10) and rejects the shorter spellings the toolchain
// also rejects.
func TestArm64MAddSub(t *testing.T) {
got := arm64Words(t, "\tMADD R1, R2, R3, R4\n\tMSUB R1, R2, R3, R4\n\tMADDW R1, R2, R3, R5\n\tMSUBW R1, R2, R3, R5\n")
want := []uint32{
0x9b010864, // MADD X4, X3, X1, X2 (Rm=1, Ra=2, Rn=3)
0x9b018864, // MSUB X4, X3, X1, X2
0x1b010865, // MADD W5, W3, W1, W2
0x1b018865, // MSUB W5, W3, W1, W2
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("madd word %d = %08x, want %08x", i, got[i], want[i])
}
}
// The accumulate operand is mandatory: 2- and 3-operand forms error
// rather than silently reading R0 or ZR as the accumulator.
for _, body := range []string{
"\tMADD R1, R2\n",
"\tMADD R1, R2, R3\n",
"\tMSUBW R1, R2, R3\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
}
// TestArm64MovImmWidth pins the immediate classifications whose size pass
// once disagreed with the encoder: negative and 0xFFFFFFFF W values go
// through MOVN after 32-bit truncation, and 3- to 4-chunk constants expand
// to one word per non-zero chunk.
func TestArm64MovImmWidth(t *testing.T) {
got := arm64Words(t, "\tMOVW $-1, R0\n\tMOVW $0xFFFFFFFF, R3\n")
want := []uint32{
0x12800000, // MOVN W0, #0
0x12800003, // MOVN W3, #0
0xd65f03c0, // RET
}
for i := range want {
if got[i] != want[i] {
t.Errorf("movw word %d = %08x, want %08x", i, got[i], want[i])
}
}
for _, tt := range []struct {
body string
words int
}{
{"\tMOVD $0x0001000200030000, R2\n", 3}, // three chunks
{"\tMOVD $0x0001000200030004, R1\n", 4}, // four chunks
{"\tMOVW $-1, R0\n", 1}, // MOVN after truncation
} {
if got := arm64Words(t, tt.body); len(got) != tt.words+1 {
t.Errorf("%s: %d words, want %d (including RET)", tt.body, len(got), tt.words+1)
}
}
}
// TestArm64ExclOffsetErrors: exclusive and atomic encodings carry no
// immediate field, so a non-zero offset is rejected the way the toolchain
// reports "illegal combination" for it, never silently dropped.
func TestArm64ExclOffsetErrors(t *testing.T) {
for _, body := range []string{
"\tLDXR 8(R1), R2\n",
"\tLDAXR 8(R1), R2\n",
"\tSTXR R3, 8(R1), R4\n",
"\tSTLXR R3, 8(R1), R4\n",
"\tCASD R3, 8(R1), R4\n",
"\tLDADDD R3, 8(R1), R4\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
}
// TestArm64ExclNoOffset pins the plain (Rn) forms, byte-for-byte against
// go tool asm. The toolchain parses the FIRST register of a store as the
// data register and the LAST as the status register (asm7.go case 59), and
// the pair forms as (Rt1, Rt2) (case 58/59):
//
// STXR R3, (R1), R4 → c8047c23 (Rt=3, Rn=1, Rs=4)
// STXP (R3, R4), (R1), R5 → c8251023 (Rt=3, Rt2=4, Rn=1, Rs=5)
// LDXP (R1), (R3, R4) → c87f1023 (Rn=1, Rt=3, Rt2=4)
func TestArm64ExclNoOffset(t *testing.T) {
got := arm64Words(t, "\tLDXR (R1), R2\n\tSTXR R3, (R1), R4\n"+
"\tSTXP (R3, R4), (R1), R5\n\tSTXPW (R3, R4), (R1), R5\n"+
"\tLDXP (R1), (R3, R4)\n\tLDXPW (R1), (R3, R4)\n"+
"\tSTXR R3, (RSP), R4\n\tLDXR (RSP), R2\n")
want := []uint32{
0xc85f7c22, // LDXR X2, [X1]
0xc8047c23, // STXR W3, [X1], W4 with Rt = R3, Rs = R4
0xc8251023, // STXP (R3, R4), [X1], R5
0x88251023, // STXPW (R3, R4), [X1], R5
0xc87f1023, // LDXP [X1], (R3, R4)
0x887f1023, // LDXPW [X1], (R3, R4)
0xc8047fe3, // STXR R3, [SP], R4
0xc85f7fe2, // LDXR [SP], R2
0xd65f03c0, // RET
}
for i := range want {
if got[i] != want[i] {
t.Errorf("excl word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64AddSubImmRange: immediates that cannot ride the imm12 field are
// rejected instead of wrapping through int32.
func TestArm64AddSubImmRange(t *testing.T) {
for _, body := range []string{
"\tADD $0x100000000, R0, R1\n",
"\tSUB $-0x100000000, R0, R1\n",
"\tCMP $0x100000000, R0\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
}
// TestArm64LargeRegisterOffset pins the large-offset path for a register
// base: the ADD offsets from the operand's own base, not from SP, matching
// the toolchain's `ADD $(256<<12), R2, R27; MOVD (R27), R3`.
func TestArm64LargeRegisterOffset(t *testing.T) {
got := arm64Words(t, "\tMOVD 0x100000(R2), R3\n\tMOVD R3, 0x100000(R2)\n")
want := []uint32{
0x9144005b, // ADD $(256<<12), R2, R27
0xf9400363, // MOVD (R27), R3
0x9144005b, // ADD $(256<<12), R2, R27
0xf9000363, // MOVD R3, (R27)
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("large offset word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64LargeFrameSpadj checks the stack-adjustment boundaries of a frame
// whose autosize must be materialised into REGTMP: $5000 rounds the autosize
// to 5024, so the prologue is [MOVD $5024, R27][SUB R27, RSP, R20][STP][ADD
// R20, SP][SUB $8] and SP moves only at its fourth word, while the RET's
// epilogue is [LDP][MOVD $5024, R27][ADD R27, RSP, RSP] before the final
// RET. These PCs feed the DWARF CFA rules and the goobj stack maps.
func TestArm64LargeFrameSpadj(t *testing.T) {
f, errs := parser.Parse("frame_arm64.s", "#include \"textflag.h\"\n\nTEXT ·framed(SB), $5000-0\n\tCALL ·other(SB)\n\tRET\n\nTEXT ·other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
fn := img.Funcs[0]
// autosize 5024: class-2 guard of 6 words (24 bytes), a 5-word prologue
// whose ADD R20, SP sits at byte 8 inside it, a one-instruction body,
// then a 3-word epilogue before the final RET.
wantSpadj := []SpadjStep{{PC: 24 + 12, Value: 5024}, {PC: 24 + 20 + 4 + 12, Value: 0}}
if len(fn.Spadj) != len(wantSpadj) {
t.Fatalf("spadj = %v, want %v", fn.Spadj, wantSpadj)
}
for i := range wantSpadj {
if fn.Spadj[i] != wantSpadj[i] {
t.Errorf("spadj[%d] = %v, want %v", i, fn.Spadj[i], wantSpadj[i])
}
}
// The words those PCs point between: the prologue's ADD R20, SP at byte
// 36, and the epilogue's materialised ADD R27, RSP, RSP right before the
// final RET at byte 60.
words := leWords(img.Code[fn.Offset : fn.Offset+fn.Size])
if got := words[(24+12)/4]; got != 0x9100029f {
t.Errorf("prologue word at byte 36 = %08x, want 9100029f (ADD R20, SP)", got)
}
if got := words[(24+20+4+8)/4]; got != 0x8b3b63ff {
t.Errorf("epilogue word at byte 56 = %08x, want 8b3b63ff (ADD R27, RSP, RSP)", got)
}
if got := words[(24+20+4+12)/4]; got != 0xd65f03c0 {
t.Errorf("final RET word at byte 60 = %08x, want d65f03c0", got)
}
}
// TestArm64SplitFrameSpadj pins the addcon2 band, where neither imm12 form
// nor a single MOVZ carries the autosize and the toolchain splits the
// prologue SUB into two imm12 instructions (asm7.go case 48) while the
// non-leaf RET still materialises the value into REGTMP (obj7.go ARET,
// issue 73259). $65664 rounds the autosize to 65680 = 144 + 16<<12:
//
// [SUB $144, RSP, R20][SUB $(16<<12), R20, R20][STP][MOVD R20, SP][SUB $8]
// [CALL]
// [LDP][MOVD $144, R27][MOVK $(1<<16), R27][ADD R27, RSP, RSP][RET]
//
// SP moves at the fourth word (byte 12) and returns to zero at the final
// RET (byte 40); the words are go tool asm's own for the same source.
func TestArm64SplitFrameSpadj(t *testing.T) {
f, errs := parser.Parse("frame_arm64.s", "#include \"textflag.h\"\n\nTEXT ·framed(SB), NOSPLIT, $65664-0\n\tCALL ·other(SB)\n\tRET\n\nTEXT ·other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
fn := img.Funcs[0]
wantSpadj := []SpadjStep{{PC: 12, Value: 65680}, {PC: 40, Value: 0}}
if len(fn.Spadj) != len(wantSpadj) {
t.Fatalf("spadj = %v, want %v", fn.Spadj, wantSpadj)
}
for i := range wantSpadj {
if fn.Spadj[i] != wantSpadj[i] {
t.Errorf("spadj[%d] = %v, want %v", i, fn.Spadj[i], wantSpadj[i])
}
}
want := []uint32{
0xd10243f4, // SUB $144, RSP, R20
0xd1404294, // SUB $(16<<12), R20, R20
0xa93ffa9d, // STP (R29, R30), -8(R20)
0x9100029f, // MOVD R20, RSP
0xd10023fd, // SUB $8, RSP, R29
0x94000000, // CALL (relocation masked at link time)
0xa97ffbfd, // LDP -8(RSP), (R29, R30)
0xd280121b, // MOVD $144, R27
0xf2a0003b, // MOVK $(1<<16), R27
0x8b3b63ff, // ADD R27, RSP, RSP
0xd65f03c0, // RET
}
words := leWords(img.Code[fn.Offset : fn.Offset+fn.Size])
if len(words) != len(want) {
t.Fatalf("framed = %d words, want %d", len(words), len(want))
}
for i, w := range want {
if words[i] != w {
t.Errorf("word %d = %08x, want %08x", i, words[i], w)
}
}
}
+256 -19
View File
@@ -63,6 +63,12 @@ type arm64FrameInfo struct {
args int // the declared -argsize args int // the declared -argsize
noSplit bool // the NOSPLIT flag noSplit bool // the NOSPLIT flag
leaf bool // no call instructions in the body leaf bool // no call instructions in the body
// Stack-split guard state: needSplit mirrors the toolchain, which skips
// the check for NOSPLIT functions and auto-marks leaf functions with an
// autosize below StackSmall as NOSPLIT.
needSplit bool
splitClass int // 0: <=StackSmall, 1: <=StackBig, 2: >StackBig
} }
// arm64ComputeFrame derives the frame layout for a TEXT function. // arm64ComputeFrame derives the frame layout for a TEXT function.
@@ -80,15 +86,68 @@ func arm64ComputeFrame(t *ast.Text) arm64FrameInfo {
if fi.frame != 0 || !fi.leaf { if fi.frame != 0 || !fi.leaf {
fi.autosize = fi.frame + 8 // space for the saved LR fi.autosize = fi.frame + 8 // space for the saved LR
if fi.autosize%16 != 0 { // The toolchain always adds an extrasize: 8 when the total leaves a
// The toolchain aligns to 16: if autosize%16 == 8, add 8; // 16-byte alignment gap, another 16 when already aligned.
// otherwise add whatever is needed. switch fi.autosize % 16 {
case 8:
fi.autosize += 8
case 0:
fi.autosize += 16
default:
// The toolchain rejects unaligned frames; round up so such
// sources still assemble.
fi.autosize += 16 - (fi.autosize % 16) fi.autosize += 16 - (fi.autosize % 16)
} }
} }
switch {
case fi.noSplit:
case fi.autosize < stackSmall && fi.leaf:
// Auto-NOSPLIT, as the toolchain's leaf mark concludes.
default:
fi.needSplit = true
switch {
case fi.autosize <= stackSmall:
fi.splitClass = 0
case fi.autosize <= stackBig:
fi.splitClass = 1
default:
fi.splitClass = 2
}
}
return fi return fi
} }
// arm64GuardLen returns the byte length of the stack-split guard prefix
// (zero when the function needs no guard). The big class materialises
// framesize-StackSmall into REGTMP, whose MOVZ/MOVK sequence length varies.
func arm64GuardLen(fi arm64FrameInfo) int {
if !fi.needSplit {
return 0
}
switch fi.splitClass {
case 0:
return 12
case 1:
return 16
default:
n, err := arm64LoadImmLen(int64(fi.autosize - stackSmall))
if err != nil {
return 0
}
return 4 + n + 4 + 4 + 4 + 4
}
}
// arm64LoadImmLen returns the byte length of the MOVZ/MOVK sequence that
// loads v into a register.
func arm64LoadImmLen(v int64) (int, error) {
b, err := encodeARM64LoadImm(27, v, "MOVD")
if err != nil {
return 0, err
}
return len(b), nil
}
// arm64IsLeaf reports whether a function contains no call instructions // arm64IsLeaf reports whether a function contains no call instructions
// (BL/CALL), matching the toolchain's LEAF mark. // (BL/CALL), matching the toolchain's LEAF mark.
func arm64IsLeaf(t *ast.Text) bool { func arm64IsLeaf(t *ast.Text) bool {
@@ -119,12 +178,99 @@ func arm64Prologue(fi arm64FrameInfo) []byte {
) )
} }
// Large frame: SUB $autosize, SP, R20; STP (FP,LR), -8(R20); ADD $0, R20, SP; SUB $8, SP, FP // Large frame: SUB $autosize, SP, R20; STP (FP,LR), -8(R20); ADD $0, R20, SP; SUB $8, SP, FP
return a64WordsLE( ws := arm64SubImmWords(uint32(fi.autosize), 20)
a64AddSub(1, 1, 0, 0, uint32(fi.autosize), 31, 20), // SUB $autosize, SP, R20 ws = append(ws,
a64LSP(2, 0, 0, -1, 30, 20, 29), // STP FP, LR, [R20, #-8] (opc=2 for 64-bit pair) a64LSP(2, 0, 0, -1, 30, 20, 29), // STP FP, LR, [R20, #-8] (opc=2 for 64-bit pair)
a64AddSub(1, 0, 0, 0, 0, 20, 31), // ADD $0, R20, SP (= MOV R20, SP) a64AddSub(1, 0, 0, 0, 0, 20, 31), // ADD $0, R20, SP (= MOV R20, SP)
a64AddSub(1, 1, 0, 0, 8, 31, 29), // SUB $8, SP, FP (op=1 for SUB) a64AddSub(1, 1, 0, 0, 8, 31, 29), // SUB $8, SP, FP (op=1 for SUB)
) )
return a64WordsLE(ws...)
}
// arm64SplitImm12 reports whether the toolchain decomposes ADD/SUB $imm into
// two imm12 instructions instead of materialising it into REGTMP
// (asm7.go case 48, the C_ADDCON2 class): the value must fit 24 bits
// unsigned and be neither encodable as one imm12 (checked by the callers
// first), nor loadable into a register in a single MOVZ/MOVN word, nor a
// logical immediate, because conclass tests all three before C_ADDCON2.
func arm64SplitImm12(imm uint32) bool {
if imm > 0xFFFFFF {
return false
}
if _, _, _, ok := arm64Bitmask(uint64(imm), 1); ok {
return false
}
return arm64Movcon(int64(imm)) < 0 && arm64Movcon(^int64(imm)) < 0
}
// arm64SubImmWords emits SUB $imm, SP, Rd with the toolchain's ladder for an
// ADD/SUB constant (asm7.go conclass and cases 2, 48, 62 and 13): the
// immediate form when the value fits imm12 (plain, or shifted left by 12
// when it is a multiple of 4096); a value with a single 16-bit chunk, a
// logical immediate, or one wider than 24 bits is materialised into REGTMP
// (R27) and subtracted in the extended-register form; everything else up to
// 0xFFFFFF is split into two imm12 instructions:
//
// SUB $(imm&0xfff), SP, Rd
// SUB $((imm&0xfff000)>>12)<<12, Rd, Rd
func arm64SubImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 0xFFF {
return []uint32{a64AddSub(1, 1, 0, 0, imm, 31, rd)}
}
if imm <= 4095<<12 && imm&0xFFF == 0 {
return []uint32{a64AddSub(1, 1, 0, 1, imm>>12, 31, rd)}
}
if !arm64SplitImm12(imm) {
mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD")
if err != nil {
mov = nil
}
return append(wordsOf(mov), arm64DPExtWords(arm64OpSub, 27, 31, rd))
}
return []uint32{
a64AddSub(1, 1, 0, 0, imm&0xFFF, 31, rd),
a64AddSub(1, 1, 0, 1, (imm&0xFFF000)>>12, rd, rd),
}
}
// arm64AddImmWords emits ADD $imm, SP, Rd with the same imm12, shifted-imm12,
// split and REGTMP ladder as arm64SubImmWords.
func arm64AddImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 0xFFF {
return []uint32{a64AddSub(1, 0, 0, 0, imm, 31, rd)}
}
if imm <= 4095<<12 && imm&0xFFF == 0 {
return []uint32{a64AddSub(1, 0, 0, 1, imm>>12, 31, rd)}
}
if !arm64SplitImm12(imm) {
mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD")
if err != nil {
mov = nil
}
return append(wordsOf(mov), arm64DPExtWords(arm64OpAdd, 27, 31, rd))
}
return []uint32{
a64AddSub(1, 0, 0, 0, imm&0xFFF, 31, rd),
a64AddSub(1, 0, 0, 1, (imm&0xFFF000)>>12, rd, rd),
}
}
// arm64RetAddWords emits the frame deallocation of a non-leaf RET with a
// large frame. The toolchain adds the frame back with a single instruction:
// a plain imm12 ADD when autosize fits 12 bits, otherwise the value is
// materialised into REGTMP and added as a register, so the epilogue never
// leaves a partially deallocated frame (obj7.go ARET, issue 73259). The
// shifted-imm12 and split-imm12 forms are therefore never used here, unlike
// the leaf epilogue's plain ADD instructions.
func arm64RetAddWords(autosize uint32) []uint32 {
if autosize < 1<<12 {
return []uint32{a64AddSub(1, 0, 0, 0, autosize, 31, 31)}
}
mov, err := encodeARM64LoadImm(27, int64(autosize), "MOVD")
if err != nil {
mov = nil
}
return append(wordsOf(mov), arm64DPExtWords(arm64OpAdd, 27, 31, 31))
} }
// arm64Return returns the bytes for a RET: the epilogue (restore FP/LR and // arm64Return returns the bytes for a RET: the epilogue (restore FP/LR and
@@ -134,10 +280,8 @@ func arm64Return(fi arm64FrameInfo) []byte {
if fi.autosize != 0 { if fi.autosize != 0 {
if fi.leaf { if fi.leaf {
// Leaf with frame: ADD $autosize-8, SP, FP; ADD $autosize, SP, SP // Leaf with frame: ADD $autosize-8, SP, FP; ADD $autosize, SP, SP
ws = append(ws, ws = append(ws, arm64AddImmWords(uint32(fi.autosize-8), 29)...)
a64AddSub(1, 0, 0, 0, uint32(fi.autosize-8), 31, 29), // ADD $autosize-8, SP, FP ws = append(ws, arm64AddImmWords(uint32(fi.autosize), 31)...)
a64AddSub(1, 0, 0, 0, uint32(fi.autosize), 31, 31), // ADD $autosize, SP, SP
)
} else if fi.autosize <= 0xf0 { } else if fi.autosize <= 0xf0 {
// Non-leaf small frame: LDR FP, [SP, #-8]; LDR.P LR, [SP], #autosize // Non-leaf small frame: LDR FP, [SP, #-8]; LDR.P LR, [SP], #autosize
ws = append(ws, ws = append(ws,
@@ -145,11 +289,11 @@ func arm64Return(fi arm64FrameInfo) []byte {
arm64PostLoad(3, 0, int32(fi.autosize), 31, 30), // LDR.P LR, [SP], #autosize arm64PostLoad(3, 0, int32(fi.autosize), 31, 30), // LDR.P LR, [SP], #autosize
) )
} else { } else {
// Large frame: LDP -8(SP), (FP, LR); ADD $autosize, SP, SP // Large frame: LDP -8(SP), (FP, LR), then deallocate.
ws = append(ws, ws = append(ws,
a64LSP(2, 0, 1, -1, 30, 31, 29), // LDP FP, LR, [SP, #-8] (opc=2 for 64-bit pair) a64LSP(2, 0, 1, -1, 30, 31, 29), // LDP FP, LR, [SP, #-8] (opc=2 for 64-bit pair)
a64AddSub(1, 0, 0, 0, uint32(fi.autosize), 31, 31), // ADD $autosize, SP, SP
) )
ws = append(ws, arm64RetAddWords(uint32(fi.autosize))...)
} }
} }
// RET: BR LR (0xd65f03c0) // RET: BR LR (0xd65f03c0)
@@ -166,22 +310,32 @@ func arm64PrologueSpadjPC(fi arm64FrameInfo) int {
if fi.autosize <= 0xf0 { if fi.autosize <= 0xf0 {
return 4 // MOVD.W instruction decrements SP return 4 // MOVD.W instruction decrements SP
} }
return 8 // SUB + STP + MOVD (3 instructions, SP updated at the MOVD) // Large frame: [SUB words][STP][ADD R20, SP]; SP moves at the ADD, whose
// position depends on how many words the SUB itself took (immediate,
// shifted immediate, the two-word imm12 split, or a materialised REGTMP
// sequence).
return 4 * (len(arm64SubImmWords(uint32(fi.autosize), 20)) + 1)
} }
// arm64ReturnEpilogueLen returns the byte length of the RET's epilogue up to // arm64ReturnEpilogueLen returns the byte length of the RET's epilogue up to
// (but not including) the final RET instruction. // (but not including) the final RET instruction. The lengths are read from
// the same word-emitting helpers the epilogue uses rather than assumed: the
// leaf path shares the prologue's immediate ladder, and a materialised
// autosize costs its MOV words plus the ADD itself.
func arm64ReturnEpilogueLen(fi arm64FrameInfo) int { func arm64ReturnEpilogueLen(fi arm64FrameInfo) int {
if fi.autosize == 0 { if fi.autosize == 0 {
return 0 return 0
} }
if fi.leaf { if fi.leaf {
return 8 // ADD + ADD return 4 * (len(arm64AddImmWords(uint32(fi.autosize-8), 29)) +
len(arm64AddImmWords(uint32(fi.autosize), 31)))
} }
if fi.autosize <= 0xf0 { if fi.autosize <= 0xf0 {
return 8 // LDR + LDR.P return 8 // LDR + LDR.P
} }
return 8 // LDP + ADD // LDP + the deallocation emitted by arm64RetAddWords, so the length
// tracks whatever the MOVD ladder needs.
return 4 + 4*len(arm64RetAddWords(uint32(fi.autosize)))
} }
// arm64ResolvePseudo translates a pseudo-register memory reference into a // arm64ResolvePseudo translates a pseudo-register memory reference into a
@@ -235,3 +389,86 @@ func arm64PostLoad(size, V int, imm9 int32, rn, rt int) uint32 {
return uint32(size)<<30 | 7<<27 | uint32(V)<<26 | 1<<22 | return uint32(size)<<30 | 7<<27 | uint32(V)<<26 | 1<<22 |
1<<10 | (uint32(imm9)&0x1FF)<<12 | uint32(rn&31)<<5 | uint32(rt&31) 1<<10 | (uint32(imm9)&0x1FF)<<12 | uint32(rn&31)<<5 | uint32(rt&31)
} }
// Data-processing (shifted register) base opcodes for the guard blocks.
const (
arm64OpAdd = 1<<31 | 0<<30 | 0<<29 | 0x0b<<24
arm64OpSub = 1<<31 | 1<<30 | 0<<29 | 0x0b<<24
arm64OpSubs = 1<<31 | 1<<30 | 1<<29 | 0x0b<<24
)
// arm64DPSRWords builds one data-processing (shifted register) word:
// OP Rm, Rn, Rd in the Go assembler's operand order.
func arm64DPSRWords(base uint32, rm, rn, rd uint32) uint32 {
return base | rm<<16 | rn<<5 | rd
}
// arm64DPExtWords builds one data-processing (extended register) word, the
// form the toolchain picks when a large immediate was materialised into
// REGTMP before the operation: base | 1<<21 | Rm<<16 | UXTX<<13 | Rn<<5 | Rd.
func arm64DPExtWords(base, rm, rn, rd uint32) uint32 {
return base | 1<<21 | rm<<16 | 3<<13 | rn<<5 | rd
}
// wordsOf converts little-endian instruction bytes back to words.
func wordsOf(b []byte) []uint32 {
ws := make([]uint32, 0, len(b)/4)
for i := 0; i+4 <= len(b); i += 4 {
ws = append(ws, uint32(b[i])|uint32(b[i+1])<<8|uint32(b[i+2])<<16|uint32(b[i+3])<<24)
}
return ws
}
// arm64GuardBytes emits the stack-split guard prefix; blockStart is the
// function-relative byte address of the morestack block the branches target.
func arm64GuardBytes(fi arm64FrameInfo, blockStart int) []byte {
// MOVD 16(R28), R16 (g.stackguard0)
ws := []uint32{a64LSU(3, 0, 1, 2, 28, 16)}
br := func(from int, cond uint32) uint32 {
return a64BranchCond(int32((blockStart-from)>>2), cond)
}
switch fi.splitClass {
case 0:
// CMP R16, RSP in the exact encoding go tool asm emits for it.
ws = append(ws, 0xeb3063ff)
ws = append(ws, br(8, a64CondLS))
case 1:
ws = append(ws, a64AddSub(1, 1, 0, 0, uint32(fi.autosize-stackSmall), 31, 17))
ws = append(ws, arm64DPSRWords(arm64OpSubs, 16, 17, 31)) // CMP R16, R17
ws = append(ws, br(12, a64CondLS))
default:
mov, err := encodeARM64LoadImm(27, int64(fi.autosize-stackSmall), "MOVD")
if err != nil {
mov = nil
}
ws = append(ws, wordsOf(mov)...)
ml := len(mov) / 4
ws = append(ws, arm64DPExtWords(arm64OpSubs, 27, 31, 17)) // SUBS R17, RSP, R27
// The branches sit at fixed byte offsets in the guard prefix: after
// the LDR (4), the ml MOV words (4*ml) and the SUBS (4) for B.LO,
// then a further B.LO word and the CMP for B.LS.
ws = append(ws, br(8+4*ml, a64CondLO))
ws = append(ws, arm64DPSRWords(arm64OpSubs, 16, 17, 31)) // CMP R16, R17
ws = append(ws, br(16+4*ml, a64CondLS))
}
return a64WordsLE(ws...)
}
// arm64MoreStackBlock emits the trailing block: MOVD R30, R3 (save LR),
// BL runtime.morestack_noctxt, B back to the function start. The BL carries
// the R_CALLARM64 relocation.
func arm64MoreStackBlock(blockStart int) ([]byte, Reloc) {
ws := []uint32{
1<<31 | 1<<29 | 0x0a<<24 | 30<<16 | 31<<5 | 3, // MOVD R30, R3
a64Branch(1, 0), // BL, patched by the linker
}
bPC := blockStart + 8
ws = append(ws, a64Branch(0, int32(-bPC>>2))) // B back to the entry
reloc := Reloc{
Off: blockStart + 4,
After: blockStart + 8,
Name: "runtime\u00b7morestack_noctxt",
Kind: RelArm64Branch,
}
return a64WordsLE(ws...), reloc
}
+126
View File
@@ -0,0 +1,126 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// parseArm64File is a helper assembling one arm64 source file.
func parseArm64File(t *testing.T, src string) *Image {
t.Helper()
f, errs := parser.Parse("k_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
return img
}
// TestArm64RelocOffsetsIncludePrologue pins the function-relative relocation
// offsets of a framed function: the offsets used to exclude the prologue, so
// every relocation landed on a prologue instruction in the GOOBJ/ELF output.
// The function calls an external, so it is a non-leaf and carries the
// stack-split guard (12 bytes, small class) before the prologue.
func TestArm64RelocOffsetsIncludePrologue(t *testing.T) {
img := parseArm64File(t, "TEXT \u00b7f(SB), $16-0\n"+
"\tBL ext\u00b7foo(SB)\n"+
"\tMOVD $gdata(SB), R5\n"+
"\tMOVD $extsym(SB), R6\n"+
"\tRET\n"+
"GLOBL gdata(SB), $8\n")
fn := img.Funcs[0]
// Layout: 12-byte guard, 12-byte prologue, BL (24), ADRP+ADD (28, 32),
// ADRP+ADD (36, 40), 12-byte epilogue with RET, 12-byte morestack block.
want := []struct {
off int
after int
name string
kind RelocKind
external bool
}{
{24, 28, "foo", RelArm64Branch, true},
{28, 28, "gdata", RelArm64Addr, false},
{32, 32, "gdata", RelArm64Addr, false},
{36, 36, "extsym", RelArm64Addr, true},
{40, 40, "extsym", RelArm64Addr, true},
{60, 64, "runtime\u00b7morestack_noctxt", RelArm64Branch, true},
}
if len(fn.Relocs) != len(want) {
t.Fatalf("relocs = %d, want %d", len(fn.Relocs), len(want))
}
for i, w := range want {
r := fn.Relocs[i]
if r.Off != w.off || r.After != w.after || r.Name != w.name || r.Kind != w.kind || r.External != w.external {
t.Errorf("reloc %d = {off %d after %d name %q kind %d ext %v}, want {off %d after %d name %q kind %d ext %v}",
i, r.Off, r.After, r.Name, r.Kind, r.External, w.off, w.after, w.name, w.kind, w.external)
}
}
// The BL with a zero offset sits exactly at the first reloc site.
code := img.Code[fn.Offset : fn.Offset+fn.Size]
if w := binary.LittleEndian.Uint32(code[24:28]); w != 0x94000000 {
t.Errorf("BL word = %08x, want 94000000", w)
}
}
// TestArm64SBLoadStoreMatchesToolchain pins the ADRP scratch register
// (REGTMP, R27) and the LDST64 relocation kind for sym loads and stores,
// against the bytes go tool asm emits for MOVD sym(SB), R5.
func TestArm64SBLoadStoreMatchesToolchain(t *testing.T) {
img := parseArm64File(t, "TEXT \u00b7ld(SB), NOSPLIT, $0\n"+
"\tMOVD sym(SB), R5\n"+
"\tMOVD R5, sym(SB)\n"+
"\tRET\n"+
"GLOBL sym(SB), $8\n")
fn := img.Funcs[0]
code := img.Code[fn.Offset : fn.Offset+fn.Size]
// go tool asm: ADRP 0(PC), R27 (9000001b); MOVD (R27), R5 (f9400365);
// ADRP 0(PC), R27; MOVD R5, (R27) (f9000365).
for off, want := range map[int]uint32{0: 0x9000001b, 4: 0xf9400365, 8: 0x9000001b, 12: 0xf9000365} {
if got := binary.LittleEndian.Uint32(code[off : off+4]); got != want {
t.Errorf("word at %d = %08x, want %08x", off, got, want)
}
}
if len(fn.Relocs) != 2 {
t.Fatalf("relocs = %d, want 2", len(fn.Relocs))
}
for i, w := range []struct{ off, after int }{{0, 8}, {8, 16}} {
r := fn.Relocs[i]
if r.Kind != RelArm64LDST64 {
t.Errorf("reloc %d kind = %d, want RelArm64LDST64 (%d)", i, r.Kind, RelArm64LDST64)
}
if r.Off != w.off || r.After != w.after {
t.Errorf("reloc %d = {off %d after %d}, want {off %d after %d}", i, r.Off, r.After, w.off, w.after)
}
}
}
// TestArm64GOObjRelocTypes checks that GOOBJ emission succeeds with the new
// relocation kinds in play; the detailed layout is covered by the goobj tests.
func TestArm64GOObjRelocTypes(t *testing.T) {
img := parseArm64File(t, "TEXT \u00b7ld(SB), NOSPLIT, $0\n"+
"\tMOVD sym(SB), R5\n"+
"\tMOVD R5, sym(SB)\n"+
"\tRET\n"+
"GLOBL sym(SB), $8\n")
obj, err := img.GOObjectAARCH64("testpkg", "k_arm64.s")
if err != nil {
t.Fatalf("GOObjectAARCH64: %v", err)
}
if len(obj) == 0 {
t.Fatal("empty object")
}
// The detailed layout is covered by the goobj tests; here we only pin
// that emission succeeds with the new relocation kinds in play.
}
+365 -15
View File
@@ -22,6 +22,11 @@ import (
// FP/SP frame-relative operands, and local-label jumps. SB (global symbol) // FP/SP frame-relative operands, and local-label jumps. SB (global symbol)
// operands require relocations and are not yet supported; the SIMD (VEX/AVX2) // operands require relocations and are not yet supported; the SIMD (VEX/AVX2)
// integer and shuffle/extract/permute/move set is in. // integer and shuffle/extract/permute/move set is in.
//
// Like the other architectures, the stack-growth guard (the morestack check
// in the prologue and the call back into the runtime in the epilogue) is not
// emitted: the bytes match go tool asm only for NOSPLIT functions or
// zero-frame leaves, where the toolchain emits no guard either.
func Assemble(t *ast.Text) ([]byte, map[string]int, error) { func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
code, _, labels, _, _, err := assemble(t, nil) code, _, labels, _, _, err := assemble(t, nil)
return code, labels, err return code, labels, err
@@ -31,7 +36,7 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
// the set of static symbols a GLOBL in the same file defines. A nil link // the set of static symbols a GLOBL in the same file defines. A nil link
// rejects SB operands outright (single-function assembly cannot resolve // rejects SB operands outright (single-function assembly cannot resolve
// them). When allowExternal is set, a reference to a symbol no GLOBL in the // them). When allowExternal is set, a reference to a symbol no GLOBL in the
// file defines is recorded as an external relocation instead of failing — // file defines is recorded as an external relocation instead of failing
// the object-file emitters resolve it at link time. // the object-file emitters resolve it at link time.
type linkInfo struct { type linkInfo struct {
symbols map[string]bool symbols map[string]bool
@@ -46,6 +51,7 @@ type sbPatch struct {
after int after int
name string name string
addend int64 addend int64
kind RelocKind
} }
// spadjStep is one stack-adjustment boundary within a function: Value is the // spadjStep is one stack-adjustment boundary within a function: Value is the
@@ -70,13 +76,18 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
return name return name
} }
// Layout: iterate jump sizes to a fixed point. // Layout: iterate jump sizes to a fixed point. The stack-split guard
// prefix and the trailing morestack block participate in the iteration:
// their conditional branches relax from rel8 to rel32 when the body
// outgrows the short form.
long := make([]bool, len(t.Body)) long := make([]bool, len(t.Body))
sizes := make([]int, len(t.Body)) sizes := make([]int, len(t.Body))
offsets := map[string]int{} offsets := map[string]int{}
pcs := make([]int, len(t.Body)) pcs := make([]int, len(t.Body))
var guardJBlong, guardJBElong, moreJMPlong bool
for { for {
pos := len(fi.prologue) guard := fi.guardLen(guardJBlong, guardJBElong)
pos := guard + len(fi.prologue)
for i, stmt := range t.Body { for i, stmt := range t.Body {
switch s := stmt.(type) { switch s := stmt.(type) {
case *ast.Label: case *ast.Label:
@@ -91,6 +102,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
pos += sz pos += sz
} }
} }
bodyLen := pos - (guard + len(fi.prologue))
// Expand any short jump whose displacement no longer fits rel8. // Expand any short jump whose displacement no longer fits rel8.
changed := false changed := false
for i, stmt := range t.Body { for i, stmt := range t.Body {
@@ -102,6 +114,11 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
if !isJumpMnemonic(mnem) || mnem == "CALL" || long[i] { if !isJumpMnemonic(mnem) || mnem == "CALL" || long[i] {
continue continue
} }
// A zero-operand jump parses; its arity is reported during
// emission (encodeJump), so the layout must not index Operands.
if len(s.Operands) != 1 {
continue
}
name, ok := labelName(s.Operands[0]) name, ok := labelName(s.Operands[0])
if !ok { if !ok {
continue // reported during emission continue // reported during emission
@@ -116,25 +133,77 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
changed = true changed = true
} }
} }
// The guard's conditional branches target the morestack block, which
// starts right after the body: the JBE measures from the end of the
// guard, so its displacement is the prologue plus the body.
if !guardJBElong && !fits8(int64(len(fi.prologue)+bodyLen)) {
guardJBElong = true
changed = true
}
if fi.splitClass == 2 && !guardJBlong {
// The underflow JB sits before the CMPQ; its displacement spans
// the rest of the guard plus the prologue and the body. The JB
// is still the short form this branch tests (relaxing it is this
// branch's job), so guardLen is taken with a short JB and the
// subtraction drops the prefix and the JB's own 2 bytes.
rest := fi.guardLen(false, guardJBElong) - (9 + 3 + 7 + 2)
if !fits8(int64(rest + len(fi.prologue) + bodyLen)) {
guardJBlong = true
changed = true
}
}
// The morestack JMP returns to the function start, so its
// displacement is the negated distance from its own end; while it is
// still short, its own length is 2 bytes.
if !moreJMPlong && !fits8(-int64(guard+len(fi.prologue)+bodyLen+5+2)) {
moreJMPlong = true
changed = true
}
if !changed { if !changed {
break break
} }
} }
// Pass 2: emit. // Pass 2: emit. The guard comes first, then the prologue, the body and
out := append([]byte(nil), fi.prologue...) // the morestack block.
guardLen := fi.guardLen(guardJBlong, guardJBElong)
bodyLen := 0
{
pos := guardLen + len(fi.prologue)
for i, stmt := range t.Body {
if _, ok := stmt.(*ast.Instr); ok {
pos += sizes[i]
}
}
bodyLen = pos - (guardLen + len(fi.prologue))
}
var out []byte
var patches []sbPatch var patches []sbPatch
if fi.needSplit {
// The JBE ends the guard, so its displacement is the prologue plus
// the body; the underflow JB additionally spans the trailing CMPQ and
// JBE, whose combined length is guardLen minus the prefix and the
// JB's own length (2 short, 6 long).
jbLen := 2
if guardJBlong {
jbLen = 6
}
guard, tlsPatch := buildGuard(fi, int32(len(fi.prologue)+bodyLen), int32(fi.guardLen(guardJBlong, guardJBElong)-(9+3+7+jbLen)+len(fi.prologue)+bodyLen))
out = append(out, guard...)
patches = append(patches, tlsPatch)
}
out = append(out, fi.prologue...)
var steps []spadjStep var steps []spadjStep
var lines []LineEntry var lines []LineEntry
if fi.useFP { if fi.useFP {
// PUSHQ BP saves the return-address-relative base (+8); the MOVQ // PUSHQ BP saves the return-address-relative base (+8); the MOVQ
// changes nothing; SUBQ $size, SP completes the frame. // changes nothing; SUBQ $size, SP completes the frame.
steps = append(steps, steps = append(steps,
spadjStep{1, 8}, spadjStep{guardLen + 1, 8},
spadjStep{len(fi.prologue), 8 + fi.size}, spadjStep{guardLen + len(fi.prologue), 8 + fi.size},
) )
} }
pos := len(fi.prologue) pos := guardLen + len(fi.prologue)
for i, stmt := range t.Body { for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr) s, ok := stmt.(*ast.Instr)
if !ok { if !ok {
@@ -156,11 +225,32 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
if len(code) != sizes[i] { if len(code) != sizes[i] {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i]) return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
} }
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
for k := range ps {
ps[k].kind = RelCall
}
}
patches = append(patches, ps...) patches = append(patches, ps...)
lines = append(lines, LineEntry{Offset: pos, Line: s.Pos().Line}) lines = append(lines, LineEntry{Offset: pos, Line: s.Pos().Line})
out = append(out, code...) out = append(out, code...)
pos += len(code) pos += len(code)
} }
if fi.needSplit {
// The morestack block: CALL runtime.morestack_noctxt, then a JMP
// back to the function entry.
jmpLen := 2
if moreJMPlong {
jmpLen = 5
}
jmpDisp := -int64(pos + 5 + jmpLen)
suffix, callPatch := buildMoreStack(int32(jmpDisp))
callPatch.off += pos
callPatch.after = pos + 5
patches = append(patches, callPatch)
out = append(out, suffix...)
pos += len(suffix)
}
_ = pos
return out, patches, offsets, steps, lines, nil return out, patches, offsets, steps, lines, nil
} }
@@ -224,16 +314,51 @@ type frameInfo struct {
spAdjust int64 // x-N(SP) becomes (spAdjust - N)(SP) spAdjust int64 // x-N(SP) becomes (spAdjust - N)(SP)
prologue []byte prologue []byte
epilogue []byte epilogue []byte
// Stack-split guard state (matching the toolchain's stacksplit): needSplit
// is false for NOSPLIT functions and for leaf functions whose frame is
// below StackSmall, which the toolchain auto-marks NOSPLIT.
needSplit bool
splitClass int // 0: <=StackSmall, 1: <=StackBig, 2: >StackBig
framesize int // the size the guard checks: frame+8 for framed functions
} }
// Stack-frame size classes from runtime/stack.go.
const (
stackSmall = 128
stackBig = 4096
)
// sbPatch gains a kind so the emitters can tell CALL and TLS patches from
// plain PC-relative displacements.
// computeFrame derives the frame layout, matching the Go assembler's default // computeFrame derives the frame layout, matching the Go assembler's default
// (a frame pointer is used whenever the function has a non-zero frame). // (a frame pointer is used whenever the function has a non-zero frame). It
// also decides whether the function needs the stack-split guard, mirroring
// obj6: a NOSPLIT function never splits, and a leaf function whose frame is
// below StackSmall is auto-marked NOSPLIT. One deliberate deviation: the
// toolchain treats zero-argument runtime calls (duffcopy and friends) as
// leaf-compatible; here any CALL makes the function a non-leaf.
func computeFrame(t *ast.Text) frameInfo { func computeFrame(t *ast.Text) frameInfo {
fi := frameInfo{} fi := frameInfo{}
if t.Frame != nil && t.Frame.Imm.HasVal { if t.Frame != nil && t.Frame.Imm.HasVal {
fi.size = int(t.Frame.Imm.Val) fi.size = int(t.Frame.Imm.Val)
} }
if fi.size > 0 { if fi.size == 0 && hasCall(t) {
// The toolchain gives a frameless function containing a CALL an
// 8-byte frame for the pushed base pointer: the prologue saves BP
// with no stack adjustment, every RET pops it back, FP references
// pass one extra slot, and the virtual SP is the hardware SP.
fi.size = 8
fi.useFP = true
// The push is the frame: the saved BP sits at SP+0 and the
// return address at SP+8, so arguments begin at SP+16. Unlike
// a SUBQ frame, the 8-byte size must not be added again.
fi.fpAdjust = 16
fi.spAdjust = 0
fi.prologue = []byte{0x55, 0x48, 0x89, 0xE5} // PUSHQ BP; MOVQ SP, BP
fi.epilogue = []byte{0x5D} // POPQ BP
} else if fi.size > 0 {
fi.useFP = true fi.useFP = true
fi.fpAdjust = int64(fi.size) + 16 // frame + saved BP + return address fi.fpAdjust = int64(fi.size) + 16 // frame + saved BP + return address
fi.spAdjust = int64(fi.size) fi.spAdjust = int64(fi.size)
@@ -242,9 +367,131 @@ func computeFrame(t *ast.Text) frameInfo {
} else { } else {
fi.fpAdjust = 8 // return address only fi.fpAdjust = 8 // return address only
} }
noSplit := false
for _, f := range t.Flags {
if strings.EqualFold(f, "NOSPLIT") {
noSplit = true
}
}
// The toolchain's autoffset: the frame plus the saved base pointer.
framesize := fi.size
if framesize > 0 {
framesize += 8
}
switch {
case noSplit:
case framesize < stackSmall && !hasCall(t):
// Auto-NOSPLIT, as the toolchain's leaf search concludes.
default:
fi.needSplit = true
fi.framesize = framesize
switch {
case framesize <= stackSmall:
fi.splitClass = 0
case framesize <= stackBig:
fi.splitClass = 1
default:
fi.splitClass = 2
}
}
return fi return fi
} }
// hasCall reports whether the function body contains a CALL instruction.
func hasCall(t *ast.Text) bool {
for _, stmt := range t.Body {
in, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if strings.ToUpper(in.Mnemonic.Text) == "CALL" {
return true
}
}
return false
}
// guardLen returns the byte length of the stack-split guard prefix. The
// final conditional branch (JBE, and JB in the big class) is 2 bytes in the
// short form and 6 in the long form.
func (fi frameInfo) guardLen(jbLong, jbeLong bool) int {
if !fi.needSplit {
return 0
}
jb, jbe := 2, 2
if jbLong {
jb = 6
}
if jbeLong {
jbe = 6
}
switch fi.splitClass {
case 0:
return 9 + 4 + jbe
case 1:
return 9 + 8 + 4 + jbe
default:
return 9 + 3 + 7 + jb + 4 + jbe
}
}
// buildGuard emits the stack-split guard prefix. jbeDisp and jbDisp are the
// already-computed displacements of the conditional branches that jump to the
// morestack block (unused in classes without them). The TLS load carries a
// R_TLS_LE patch site at offset 5.
func buildGuard(fi frameInfo, jbeDisp, jbDisp int32) ([]byte, sbPatch) {
out := []byte{
0x64, 0x4c, 0x8b, 0x34, 0x25, // MOVQ FS:0, R14
0, 0, 0, 0, // TLS slot offset, filled by the linker
}
tls := sbPatch{off: 5, after: 9, kind: RelTLSLE}
jmp := func(op8, op32 byte, disp int32) []byte {
if disp >= -128 && disp <= 127 {
return []byte{op8, byte(disp)}
}
return append([]byte{0x0F, op32}, le32(int64(disp))...)
}
switch fi.splitClass {
case 0:
// CMPQ SP, 16(R14)
out = append(out, 0x49, 0x3b, 0x66, 0x10)
out = append(out, jmp(0x76, 0x86, jbeDisp)...)
case 1:
// LEAQ -(framesize-StackSmall)(SP), R12; CMPQ R12, 16(R14)
out = append(out, 0x4c, 0x8d, 0xa4, 0x24)
out = append(out, le32(-int64(fi.framesize-stackSmall))...)
out = append(out, 0x4d, 0x3b, 0x66, 0x10)
out = append(out, jmp(0x76, 0x86, jbeDisp)...)
default:
// MOVQ SP, R12; SUBQ $(framesize-StackSmall), R12; JB; CMPQ R12, 16(R14)
out = append(out, 0x49, 0x89, 0xe4)
out = append(out, 0x49, 0x81, 0xec)
out = append(out, le32(int64(fi.framesize-stackSmall))...)
out = append(out, jmp(0x72, 0x82, jbDisp)...)
out = append(out, 0x4d, 0x3b, 0x66, 0x10)
out = append(out, jmp(0x76, 0x86, jbeDisp)...)
}
return out, tls
}
// buildMoreStack emits the trailing block: CALL runtime.morestack_noctxt
// (patched by the linker) and a JMP back to the function start.
func buildMoreStack(jmpDisp int32) ([]byte, sbPatch) {
out := []byte{0xE8, 0, 0, 0, 0}
call := sbPatch{off: 1, after: 5, name: "runtime\u00b7morestack_noctxt", kind: RelCall}
out = append(out, jmpBytes(jmpDisp)...)
return out, call
}
// jmpBytes encodes a near JMP in the short or long form.
func jmpBytes(disp int32) []byte {
if disp >= -128 && disp <= 127 {
return []byte{0xEB, byte(disp)}
}
return append([]byte{0xE9}, le32(int64(disp))...)
}
// prologueBytes emits: PUSHQ BP; MOVQ SP, BP; SUBQ $size, SP. // prologueBytes emits: PUSHQ BP; MOVQ SP, BP; SUBQ $size, SP.
func prologueBytes(size int) []byte { func prologueBytes(size int) []byte {
out := []byte{0x55, 0x48, 0x89, 0xE5} // PUSHQ BP; MOVQ SP, BP out := []byte{0x55, 0x48, 0x89, 0xE5} // PUSHQ BP; MOVQ SP, BP
@@ -258,14 +505,11 @@ func epilogueBytes(size int) []byte {
} }
func subSP(size int) []byte { // SUBQ $size, SP func subSP(size int) []byte { // SUBQ $size, SP
// imm8 holds -128..127; anything larger takes the imm32 form, exactly as
// the Go assembler encodes it (verified for 8, 128, 200 and 255).
if size >= -128 && size <= 127 { if size >= -128 && size <= 127 {
return []byte{0x48, 0x83, 0xEC, byte(int8(size))} return []byte{0x48, 0x83, 0xEC, byte(int8(size))}
} }
// 128..255 do not fit SUB's unsigned imm8, but the Go assembler
// switches to ADDQ $-size, SP whose sign-extended imm8 does.
if size >= -255 && size <= 255 {
return []byte{0x48, 0x83, 0xC4, byte(int8(-size))}
}
return append([]byte{0x48, 0x81, 0xEC}, le32(int64(size))...) return append([]byte{0x48, 0x81, 0xEC}, le32(int64(size))...)
} }
@@ -282,6 +526,16 @@ func addSP(size int) []byte { // ADDQ $size, SP
func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, error) { func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, error) {
mnem := strings.ToUpper(s.Mnemonic.Text) mnem := strings.ToUpper(s.Mnemonic.Text)
if isJumpMnemonic(mnem) { if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
return 5, nil // opcode + rel32, always the long form
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
code, err := encodeIndirectJump(s, mnem)
if err != nil {
return 0, err
}
return len(code), nil
}
return jumpSize(mnem, long), nil return jumpSize(mnem, long), nil
} }
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link) code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
@@ -331,6 +585,33 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
var ps []sbPatch var ps []sbPatch
var err error var err error
if isJumpMnemonic(mnem) { if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
// CALL/JMP sym(SB): a rel32 call (or tail call) against a
// static or external symbol, resolved by the file-level layout
// or the linker.
code, ps, err = encodeSBCall(s, link)
if err != nil {
return nil, nil, err
}
for i := range ps {
ps[i].kind = RelCall
}
body := pc + len(prefix)
for i := range ps {
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
}
return append(prefix, code...), nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve) code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
} else { } else {
code, ps, err = encodeNormal(s, fi, link) code, ps, err = encodeNormal(s, fi, link)
@@ -412,6 +693,37 @@ func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long
} }
} }
// isSBCall reports whether the CALL operand is a symbol reference.
func isSBCall(s *ast.Instr) bool {
return len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpAddr &&
s.Operands[0].Addr.Sym != nil && s.Operands[0].Addr.Sym.Pseudo == "SB"
}
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
if err != nil {
return nil, nil, err
}
m, ok := o.(sbMem)
if !ok {
return nil, nil, fmt.Errorf("CALL: unsupported operand")
}
opcode := []byte{0xE8}
if strings.ToUpper(s.Mnemonic.Text) == "JMP" {
opcode = []byte{0xE9} // a tail call, no return address pushed
}
e := &enc{}
if err := e.emit(&instr{opcode: opcode, modrm: -1, sib: -1, disp: le32(0), sb: &sbRef{name: m.name, addend: m.addend}}); err != nil {
return nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelCall}
}
return e.out, ps, nil
}
// labelName extracts a local-label name from a jump operand. // labelName extracts a local-label name from a jump operand.
func labelName(op *ast.Operand) (string, bool) { func labelName(op *ast.Operand) (string, bool) {
if op.Kind == ast.OpAddr && op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" && if op.Kind == ast.OpAddr && op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" &&
@@ -421,6 +733,44 @@ func labelName(op *ast.Operand) (string, bool) {
return "", false return "", false
} }
// indirectJumpTarget reports whether the JMP/CALL operand addresses a
// register or a memory location rather than a label or a static symbol.
// A bare identifier is a register when the register table knows the name and
// a label otherwise, which is exactly how the parser cannot distinguish them.
func indirectJumpTarget(s *ast.Instr) bool {
if len(s.Operands) != 1 || s.Operands[0].Kind != ast.OpAddr {
return false
}
a := s.Operands[0].Addr
if a.Base != "" || a.Index != "" {
return true
}
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if _, ok := ParseReg(a.Sym.Name); ok {
return true
}
}
return false
}
// encodeIndirectJump assembles a JMP/CALL through a register or memory
// operand, which carries no relocation and no label to resolve.
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encodeIndirectBranch(mnem, ops); err != nil {
return nil, err
}
return e.out, nil
}
// spReg is the hardware stack pointer used to realise FP/SP pseudo-operands. // spReg is the hardware stack pointer used to realise FP/SP pseudo-operands.
var spReg = Reg{idx: 4, size: 8} var spReg = Reg{idx: 4, size: 8}
+90 -2
View File
@@ -4,6 +4,7 @@
package asm package asm
import ( import (
"bytes"
"strings" "strings"
"testing" "testing"
@@ -159,6 +160,57 @@ TEXT ·loadarg(SB), NOSPLIT, $0-24
} }
} }
// TestAssembleFramelessCall verifies the forced base-pointer frame a $0-frame
// function containing a CALL receives: the PUSHQ BP prologue with no stack
// adjustment and the x+N(FP) → (N+16)(SP) translation, against the bytes the
// Go assembler produces. The push is the frame, so the offset must not count
// it twice.
func TestAssembleFramelessCall(t *testing.T) {
f, errs := parser.Parse("frameless_call_amd64.s", `
#include "textflag.h"
TEXT ·withcall(SB), NOSPLIT, $0-16
MOVQ x+0(FP), AX
CALL ·other(SB)
MOVQ AX, ret+8(FP)
RET
TEXT ·other(SB), NOSPLIT, $0-0
RET
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
code := append([]byte(nil), img.Code[img.Funcs[0].Offset:img.Funcs[0].Offset+img.Funcs[0].Size]...)
for _, r := range img.Funcs[0].Relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
// From `go tool objdump` of the Go-assembled function:
// PUSHQ BP 55
// MOVQ SP, BP 4889e5
// MOVQ 0x10(SP), AX 488b442410
// CALL other e800000000
// MOVQ AX, 0x18(SP) 4889442418
// POPQ BP 5d
// RET c3
want := []byte{
0x55,
0x48, 0x89, 0xe5,
0x48, 0x8b, 0x44, 0x24, 0x10,
0xe8, 0x00, 0x00, 0x00, 0x00,
0x48, 0x89, 0x44, 0x24, 0x18,
0x5d,
0xc3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("frameless CALL FP translation mismatch:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
// TestAssembleFrame verifies a function with a non-zero frame: the Go-style // TestAssembleFrame verifies a function with a non-zero frame: the Go-style
// prologue/epilogue and the x+N(FP) → (N+frame+16)(SP) translation, against // prologue/epilogue and the x+N(FP) → (N+frame+16)(SP) translation, against
// the bytes the Go assembler produces. // the bytes the Go assembler produces.
@@ -201,8 +253,8 @@ TEXT ·withframe(SB), NOSPLIT, $16-16
} }
// TestAssembleVexKernel assembles the horizontal-sum reduction the go-flac // TestAssembleVexKernel assembles the horizontal-sum reduction the go-flac
// kernels end with — exercising the VEX moves, shuffle and extract forms // kernels end with; exercising the VEX moves, shuffle and extract forms
// through the full parser → encoder path — and checks the output is // through the full parser → encoder path; and checks the output is
// byte-identical to the Go assembler's. // byte-identical to the Go assembler's.
func TestAssembleVexKernel(t *testing.T) { func TestAssembleVexKernel(t *testing.T) {
fn := firstText(t, ` fn := firstText(t, `
@@ -351,3 +403,39 @@ TEXT ·pf(SB), NOSPLIT, $0
t.Errorf("PREFETCHT0 bytes: got %s, want 0f 18 0b", hex) t.Errorf("PREFETCHT0 bytes: got %s, want 0f 18 0b", hex)
} }
} }
// TestAssembleBareJump checks that a zero-operand jump (which parses, because
// the parser does not arity-check mnemonics) is rejected with an error rather
// than panicking in the layout loop, which indexes Operands[0] before the
// emission pass gets a chance to diagnose the arity.
func TestAssembleBareJump(t *testing.T) {
for _, mnem := range []string{"JE", "JMP", "JLT", "CALL"} {
fn := firstText(t, "TEXT ·bare(SB), $16-0\n\t"+mnem+"\n")
if _, _, err := Assemble(fn); err == nil {
t.Errorf("%s with no operand: expected an error, got none", mnem)
}
}
}
// TestSubSPEncodings pins the prologue SUB against the bytes go tool asm
// emits for SUBQ $size, SP: imm8 for -128..127, the imm32 form for anything
// larger. The intermediate 129..255 range used to encode an ADD with a
// truncated immediate, moving SP the wrong way.
func TestSubSPEncodings(t *testing.T) {
for _, tt := range []struct {
size int
want []byte
}{
{8, []byte{0x48, 0x83, 0xEC, 0x08}},
{127, []byte{0x48, 0x83, 0xEC, 0x7F}},
{128, []byte{0x48, 0x81, 0xEC, 0x80, 0x00, 0x00, 0x00}},
{200, []byte{0x48, 0x81, 0xEC, 0xC8, 0x00, 0x00, 0x00}},
{255, []byte{0x48, 0x81, 0xEC, 0xFF, 0x00, 0x00, 0x00}},
{4096, []byte{0x48, 0x81, 0xEC, 0x00, 0x10, 0x00, 0x00}},
} {
got := subSP(tt.size)
if !bytes.Equal(got, tt.want) {
t.Errorf("subSP(%d) = %x, want %x", tt.size, got, tt.want)
}
}
}
+59 -7
View File
@@ -12,7 +12,7 @@ import (
// Image: a .text section holding the function bodies, a .data section // Image: a .text section holding the function bodies, a .data section
// holding the GLOBL initialisers, a symbol table with one symbol per TEXT // holding the GLOBL initialisers, a symbol table with one symbol per TEXT
// and GLOBL (file-local <> symbols are STB_LOCAL, the rest STB_GLOBAL), and // and GLOBL (file-local <> symbols are STB_LOCAL, the rest STB_GLOBAL), and
// a .rela.text relocation table — one R_X86_64_PC32 entry per static-symbol // a .rela.text relocation table, one R_X86_64_PC32 entry per static-symbol
// reference, internal references resolving against the local data symbols // reference, internal references resolving against the local data symbols
// and external ones against undefined globals. The output links with the // and external ones against undefined globals. The output links with the
// system toolchain (cc/ld) the way a hand-assembled .o would. // system toolchain (cc/ld) the way a hand-assembled .o would.
@@ -43,6 +43,9 @@ const (
stInfoShift = 4 stInfoShift = 4
rX8664PC32 = 2 rX8664PC32 = 2
// R_X86_64_TPOFF32 (debug/elf): the local-exec TLS offset the stack
// guard loads from FS. 20 is R_X86_64_TLSLD, a different relocation.
rX8664TPOFF32 = 23
) )
// elfSym is one symbol-table entry in construction. // elfSym is one symbol-table entry in construction.
@@ -71,7 +74,7 @@ func (img *Image) ELFObject() ([]byte, error) {
// Build the symbol table: the null entry and the two section symbols // Build the symbol table: the null entry and the two section symbols
// come first, then the local symbols (static TEXT and GLOBL), then the // come first, then the local symbols (static TEXT and GLOBL), then the
// globals (exported TEXT and GLOBL, and the undefined externals) — ELF // globals (exported TEXT and GLOBL, and the undefined externals), ELF
// requires every local to precede every global, and sh_info records the // requires every local to precede every global, and sh_info records the
// boundary. symIdx maps a symbol name to its index for the relocations. // boundary. symIdx maps a symbol name to its index for the relocations.
var locals, globals []elfSym var locals, globals []elfSym
@@ -125,11 +128,19 @@ func (img *Image) ELFObject() ([]byte, error) {
type elfRela struct { type elfRela struct {
off uint64 off uint64
sym int sym int
typ uint32
addend int64 addend int64
} }
var relas []elfRela var relas []elfRela
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
for _, r := range fn.Relocs { for _, r := range fn.Relocs {
var typ uint32 = rX8664PC32
if r.Kind == RelTLSLE {
// R_X86_64_TPOFF32 resolves to the local-exec TLS offset and
// carries no symbol.
relas = append(relas, elfRela{off: uint64(fn.Offset + r.Off), sym: 0, typ: rX8664TPOFF32})
continue
}
idx, ok := symIdx[r.Name] idx, ok := symIdx[r.Name]
if !ok { if !ok {
return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name) return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name)
@@ -137,6 +148,7 @@ func (img *Image) ELFObject() ([]byte, error) {
relas = append(relas, elfRela{ relas = append(relas, elfRela{
off: uint64(fn.Offset + r.Off), off: uint64(fn.Offset + r.Off),
sym: idx, sym: idx,
typ: typ,
// R_X86_64_PC32 computes S + A − P with P the patch site; the // R_X86_64_PC32 computes S + A − P with P the patch site; the
// assembler measures the symbol from the instruction end, // assembler measures the symbol from the instruction end,
// After − Off bytes past the field, so the addend carries // After − Off bytes past the field, so the addend carries
@@ -209,7 +221,7 @@ func (img *Image) ELFObject() ([]byte, error) {
for _, r := range relas { for _, r := range relas {
var b [24]byte var b [24]byte
le.PutUint64(b[0:], r.off) le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|rX8664PC32) le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend)) le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...) out = append(out, b[:]...)
} }
@@ -218,15 +230,32 @@ func (img *Image) ELFObject() ([]byte, error) {
shstrOff := len(out) shstrOff := len(out)
out = append(out, stSections.bytes()...) out = append(out, stSections.bytes()...)
// DWARF debug sections (no relocations — the linker resolves DWARF fixups). // DWARF debug sections; the address placeholders they leave are carried
// as .rela.debug_info/.rela.debug_line entries the system linker applies.
dwAlign := func(n int) { dwAlign := func(n int) {
for len(out)%n != 0 { for len(out)%n != 0 {
out = append(out, 0) out = append(out, 0)
} }
} }
dw := appendDWARFSections(&out, img, "gasm.s", symIdx, dwAlign) dw := appendDWARFSections(&out, img, dwarfSourceName(img), symIdx, dwAlign, cfiAMD64)
dwarfStart := 0 // section index of .debug_abbrev, set when DWARF is present
if dw != nil { if dw != nil {
nSections += 4 // .debug_abbrev, .debug_info, .debug_line, .debug_line_str // Five DWARF sections: .debug_abbrev, .debug_info, .debug_line,
// .debug_line_str and .debug_frame (the CIE is unconditional, so
// the frame section is always present), plus the relocation
// sections below when they carry entries.
dwarfStart = nSections
nSections += 5
appendDWARFRelas(&out, dw, rX8664Abs64, dwAlign)
if dw.infoRelaCount > 0 {
nSections++
}
if dw.lineRelaCount > 0 {
nSections++
}
if dw.frameRelaCount > 0 {
nSections++
}
} }
align(8) align(8)
@@ -257,14 +286,37 @@ func (img *Image) ELFObject() ([]byte, error) {
} }
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0) putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers. // DWARF section headers; their indices follow the write order.
if dw != nil { if dw != nil {
// secIdx is a running section index: each putSh below emits the
// next header, and the sh_info of a .rela section names the index
// of the section it relocates.
secIdx := dwarfStart
putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0) putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0)
secIdx++
putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0) putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0)
secInfoIdx := secIdx
secIdx++
if dw.infoRelaCount > 0 {
putSh(".rela.debug_info", shtRela, 0, dw.infoRelaOff, 24*dw.infoRelaCount, secSymtab, secInfoIdx, 8, 24)
secIdx++
}
putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0) putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0)
secLineIdx := secIdx
secIdx++
if dw.lineRelaCount > 0 {
putSh(".rela.debug_line", shtRela, 0, dw.lineRelaOff, 24*dw.lineRelaCount, secSymtab, secLineIdx, 8, 24)
secIdx++
}
putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0) putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0)
secIdx++
if dw.frameSize > 0 { if dw.frameSize > 0 {
putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0) putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0)
secFrameIdx := secIdx
secIdx++
if dw.frameRelaCount > 0 {
putSh(".rela.debug_frame", shtRela, 0, dw.frameRelaOff, 24*dw.frameRelaCount, secSymtab, secFrameIdx, 8, 24)
}
} }
} }
+162 -75
View File
@@ -12,43 +12,74 @@ import (
// self-contained sections because the system linker only performs fixup // self-contained sections because the system linker only performs fixup
// relocations, not assembly. // relocations, not assembly.
// DWARF5 attribute, form and line-table constants (the values the
// toolchain uses, cmd/internal/dwarf/dwarf_defs.go; the DIE streams below
// are written against these forms).
const (
dwAtName = 0x03 // DW_AT_name
dwAtStmtList = 0x10 // DW_AT_stmt_list
dwAtLowPC = 0x11 // DW_AT_low_pc
dwAtHighPC = 0x12 // DW_AT_high_pc
dwAtDeclFile = 0x3a // DW_AT_decl_file
dwAtDeclLine = 0x3b // DW_AT_decl_line
dwAtExternal = 0x3f // DW_AT_external
dwAtFrameBase = 0x40 // DW_AT_frame_base
dwTagSubprog = 0x2e // DW_TAG_subprogram
dwTagCompUnit = 0x11 // DW_TAG_compile_unit
dwFormAddr = 0x01 // DW_FORM_addr
dwFormData8 = 0x07 // DW_FORM_data8
dwFormString = 0x08 // DW_FORM_string
dwFormData1 = 0x0b // DW_FORM_data1
dwFormUdata = 0x0f // DW_FORM_udata
dwFormSecOff = 0x17 // DW_FORM_sec_offset
dwFormExprloc = 0x18 // DW_FORM_exprloc
dwFormLineStrp = 0x1f // DW_FORM_line_strp
dwLnctPath = 0x01 // DW_LNCT_path
dwLnctDirIndex = 0x02 // DW_LNCT_directory_index
)
// dwarfAbbrevTable returns the .debug_abbrev content: a single compilation // dwarfAbbrevTable returns the .debug_abbrev content: a single compilation
// unit with DW_TAG_compile_unit and DW_TAG_subprogram entries. // unit with DW_TAG_compile_unit and DW_TAG_subprogram entries. The
// attribute/form pairs must match the DIE streams dwarfBuildInfoSection
// writes byte for byte, in the same order, or every consumer's parse of
// .debug_info desynchronises.
func dwarfAbbrevTable() []byte { func dwarfAbbrevTable() []byte {
var b []byte var b []byte
// Abbrev 1: DW_TAG_compile_unit // Abbrev 1: DW_TAG_compile_unit.
b = append(b, 1) // abbreviation code b = append(b, 1) // abbreviation code
b = append(b, 0x11) // DW_TAG_compile_unit b = appendUleb(b, dwTagCompUnit) // DW_TAG_compile_unit
b = append(b, 1) // DW_CHILDREN_yes b = append(b, 1) // DW_CHILDREN_yes
b = appendUleb(b, 0x1b) // DW_AT_low_pc b = appendUleb(b, dwAtLowPC) // DW_AT_low_pc
b = appendUleb(b, 0x01) // DW_FORM_addr b = appendUleb(b, dwFormAddr) // DW_FORM_addr
b = appendUleb(b, 0x29) // DW_AT_high_pc b = appendUleb(b, dwAtHighPC) // DW_AT_high_pc
b = appendUleb(b, 0x07) // DW_FORM_data8 b = appendUleb(b, dwFormData8) // DW_FORM_data8
b = appendUleb(b, 0x10) // DW_AT_stmt_list b = appendUleb(b, dwAtStmtList) // DW_AT_stmt_list
b = appendUleb(b, 0x25) // DW_FORM_sec_offset b = appendUleb(b, dwFormSecOff) // DW_FORM_sec_offset (4 bytes here)
b = appendUleb(b, 0x01) // DW_AT_name b = appendUleb(b, dwAtName) // DW_AT_name
b = appendUleb(b, 0x08) // DW_FORM_string b = appendUleb(b, dwFormString) // DW_FORM_string
b = appendUleb(b, 0) // end of attributes b = appendUleb(b, 0) // end of attributes: attr 0
b = appendUleb(b, 0) // ... paired with form 0
// Abbrev 2: DW_TAG_subprogram // Abbrev 2: DW_TAG_subprogram.
b = append(b, 2) // abbreviation code b = append(b, 2) // abbreviation code
b = append(b, 0x2e) // DW_TAG_subprogram b = appendUleb(b, dwTagSubprog) // DW_TAG_subprogram
b = append(b, 0) // DW_CHILDREN_no b = append(b, 0) // DW_CHILDREN_no
b = appendUleb(b, 0x03) // DW_AT_name b = appendUleb(b, dwAtName) // DW_AT_name
b = appendUleb(b, 0x08) // DW_FORM_string b = appendUleb(b, dwFormString) // DW_FORM_string
b = appendUleb(b, 0x11) // DW_AT_low_pc b = appendUleb(b, dwAtLowPC) // DW_AT_low_pc
b = appendUleb(b, 0x01) // DW_FORM_addr b = appendUleb(b, dwFormAddr) // DW_FORM_addr
b = appendUleb(b, 0x29) // DW_AT_high_pc b = appendUleb(b, dwAtHighPC) // DW_AT_high_pc
b = appendUleb(b, 0x07) // DW_FORM_data8 b = appendUleb(b, dwFormData8) // DW_FORM_data8
b = appendUleb(b, 0x3f) // DW_AT_frame_base b = appendUleb(b, dwAtFrameBase) // DW_AT_frame_base
b = appendUleb(b, 0x18) // DW_FORM_exprloc b = appendUleb(b, dwFormExprloc) // DW_FORM_exprloc
b = appendUleb(b, 0x3b) // DW_AT_decl_file b = appendUleb(b, dwAtDeclFile) // DW_AT_decl_file
b = appendUleb(b, 0x0b) // DW_FORM_data1 b = appendUleb(b, dwFormData1) // DW_FORM_data1
b = appendUleb(b, 0x37) // DW_AT_decl_line b = appendUleb(b, dwAtDeclLine) // DW_AT_decl_line
b = appendUleb(b, 0x0b) // DW_FORM_data1 b = appendUleb(b, dwFormData1) // DW_FORM_data1
b = appendUleb(b, 0x63) // DW_AT_external b = appendUleb(b, dwAtExternal) // DW_AT_external
b = appendUleb(b, 0x0b) // DW_FORM_flag b = appendUleb(b, 0x0c) // DW_FORM_flag (one byte, 0 or 1)
b = appendUleb(b, 0) // end of attributes b = appendUleb(b, 0) // end of attributes: attr 0
b = appendUleb(b, 0) // ... paired with form 0
// End of table. // End of table.
b = append(b, 0) b = append(b, 0)
@@ -68,6 +99,9 @@ type dwarfSections struct {
infoRelocs []dwarfReloc infoRelocs []dwarfReloc
// Relocations for .debug_line: (offset, symbol name, addend). // Relocations for .debug_line: (offset, symbol name, addend).
lineRelocs []dwarfReloc lineRelocs []dwarfReloc
// Relocations for .debug_frame: (offset, symbol name, addend), one per
// FDE initial_location.
frameRelocs []dwarfReloc
} }
type dwarfReloc struct { type dwarfReloc struct {
@@ -76,8 +110,9 @@ type dwarfReloc struct {
addend int64 addend int64
} }
// emitDWARF generates complete DWARF5 sections for the image. // emitDWARF generates complete DWARF5 sections for the image. cfi carries
func emitDWARF(img *Image, srcFile string) *dwarfSections { // the architecture's .debug_frame register conventions.
func emitDWARF(img *Image, srcFile string, cfi cfiArch) *dwarfSections {
ds := &dwarfSections{} ds := &dwarfSections{}
ds.debugAbbrev = dwarfAbbrevTable() ds.debugAbbrev = dwarfAbbrevTable()
@@ -86,19 +121,21 @@ func emitDWARF(img *Image, srcFile string) *dwarfSections {
lineStr.add(srcFile) lineStr.add(srcFile)
ds.debugLineStr = lineStr.bytes() ds.debugLineStr = lineStr.bytes()
// Build .debug_line. // Build .debug_line; the file table references the source name through
ds.debugLine = dwarfBuildLineSection(img, ds) // its offset in .debug_line_str.
ds.debugLine = dwarfBuildLineSection(img, uint32(lineStr.at(srcFile)), ds)
// Build .debug_info. // Build .debug_info.
ds.debugInfo = dwarfBuildInfoSection(img, srcFile, ds) ds.debugInfo = dwarfBuildInfoSection(img, srcFile, ds)
// Build .debug_frame. // Build .debug_frame.
ds.debugFrame = dwarfBuildFrameSection(img) ds.debugFrame = dwarfBuildFrameSection(img, cfi, ds)
return ds return ds
} }
// dwarfBuildLineSection builds a complete .debug_line section. // dwarfBuildLineSection builds a complete .debug_line section. srcStrOff is
func dwarfBuildLineSection(img *Image, ds *dwarfSections) []byte { // the source file name's offset in .debug_line_str.
func dwarfBuildLineSection(img *Image, srcStrOff uint32, ds *dwarfSections) []byte {
var b []byte var b []byte
le := binary.LittleEndian le := binary.LittleEndian
@@ -120,15 +157,25 @@ func dwarfBuildLineSection(img *Image, ds *dwarfSections) []byte {
// Standard opcode lengths (opcode 1..opcode_base-1). // Standard opcode lengths (opcode 1..opcode_base-1).
b = append(b, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0) b = append(b, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0)
// Directory table (DWARF5 format). // Directory table (DWARF5 §6.2.4): entry format descriptors followed by
b = append(b, 0) // one directory entry (index 0 = empty) // the entries. One directory, the compilation directory, whose path is
// File table. // the empty string at .debug_line_str offset 0.
b = appendUleb(b, 1) // file count b = append(b, 1) // directory_entry_format_count
// File 1: name index into .debug_line_str, dir index, time, size. b = appendUleb(b, dwLnctPath) // DW_LNCT_path
b = appendUleb(b, 0) // name (index 0 in line_str) b = appendUleb(b, dwFormLineStrp) // DW_FORM_line_strp
b = appendUleb(b, 0) // directory index b = appendUleb(b, 1) // directories_count
b = appendUleb(b, 0) // last modification time b = le.AppendUint32(b, 0) // .debug_line_str offset of ""
b = appendUleb(b, 0) // file size
// File table (DWARF5 §6.2.5). v5 indexes files from 0, so the source
// file is entry 0, matching the DW_AT_decl_file value 0 the DIEs carry.
b = append(b, 2) // file_name_entry_format_count
b = appendUleb(b, dwLnctPath) // DW_LNCT_path
b = appendUleb(b, dwFormLineStrp) // DW_FORM_line_strp
b = appendUleb(b, dwLnctDirIndex) // DW_LNCT_directory_index
b = appendUleb(b, dwFormUdata) // DW_FORM_udata
b = appendUleb(b, 1) // file_names_count
b = le.AppendUint32(b, srcStrOff) // .debug_line_str offset of the source name
b = appendUleb(b, 0) // directory index 0 (the compilation directory)
headerEnd := len(b) headerEnd := len(b)
@@ -176,8 +223,11 @@ func dwarfBuildLineSection(img *Image, ds *dwarfSections) []byte {
// Patch unit_length. // Patch unit_length.
le.PutUint32(b[headerStart:], uint32(len(b)-headerStart-4)) le.PutUint32(b[headerStart:], uint32(len(b)-headerStart-4))
// Patch header_length. // Patch header_length. In the v5 header it follows the one-byte
le.PutUint32(b[headerStart+6:], uint32(headerEnd-headerStart-10)) // address_size and segment_selector_size (offset 8, not the DWARF2-4
// offset 6), and counts from just past itself to the first program
// byte.
le.PutUint32(b[headerStart+8:], uint32(headerEnd-headerStart-12))
return b return b
} }
@@ -195,14 +245,16 @@ func dwarfBuildInfoSection(img *Image, srcFile string, ds *dwarfSections) []byte
// DW_TAG_compile_unit (abbrev 1). // DW_TAG_compile_unit (abbrev 1).
b = append(b, 1) // abbreviation code b = append(b, 1) // abbreviation code
// DW_AT_low_pc: address of .text start. // DW_AT_low_pc: address of .text start. A data-only image has no
infoRelocBase := len(b) // functions to relocate against; its CU covers no code, so the base
// stays zero (the DWARF "no base address" value) with no relocation.
b = le.AppendUint64(b, 0) // placeholder b = le.AppendUint64(b, 0) // placeholder
ds.infoRelocs = append(ds.infoRelocs, dwarfReloc{ if len(img.Funcs) > 0 {
off: uint64(infoRelocBase), ds.infoRelocs = append(ds.infoRelocs, dwarfReloc{
name: img.Funcs[0].Name, off: uint64(len(b) - 8),
addend: 0, name: img.Funcs[0].Name,
}) })
}
// DW_AT_high_pc: size of .text. // DW_AT_high_pc: size of .text.
b = le.AppendUint64(b, uint64(len(img.Code))) b = le.AppendUint64(b, uint64(len(img.Code)))
// DW_AT_stmt_list: offset into .debug_line (0). // DW_AT_stmt_list: offset into .debug_line (0).
@@ -229,8 +281,9 @@ func dwarfBuildInfoSection(img *Image, srcFile string, ds *dwarfSections) []byte
b = le.AppendUint64(b, uint64(fn.Size)) b = le.AppendUint64(b, uint64(fn.Size))
// DW_AT_frame_base: DW_OP_call_frame_cfa. // DW_AT_frame_base: DW_OP_call_frame_cfa.
b = append(b, 1, 0x9c) b = append(b, 1, 0x9c)
// DW_AT_decl_file: file index 1. // DW_AT_decl_file: the single file-table entry, index 0 (v5 indexes
b = append(b, 1) // files from 0).
b = append(b, 0)
// DW_AT_decl_line. // DW_AT_decl_line.
b = append(b, uint8(fn.Line)) b = append(b, uint8(fn.Line))
// DW_AT_external. // DW_AT_external.
@@ -253,41 +306,75 @@ func appendUleb(b []byte, v uint64) []byte {
return binary.AppendUvarint(b, v) return binary.AppendUvarint(b, v)
} }
// appendSleb appends v in signed LEB128, the encoding DWARF specifies:
// two's-complement sign extension, which is NOT Go's zigzag varint
// (binary.AppendVarint(-8) encodes 15, where DWARF wants 0x78).
func appendSleb(b []byte, v int64) []byte { func appendSleb(b []byte, v int64) []byte {
return binary.AppendVarint(b, v) for {
c := byte(v & 0x7f)
v >>= 7
if (v == 0 && c&0x40 == 0) || (v == -1 && c&0x40 != 0) {
return append(b, c)
}
b = append(b, c|0x80)
}
} }
// cfiArch carries the .debug_frame CIE parameters that differ per
// architecture: the DWARF register numbers of the stack pointer the initial
// CFA rule names and of the return address. The values are the ones the Go
// linker writes into its own CIE (cmd/link/internal/ld/dwarf.go uses
// Dwarfregsp and Dwarfreglr; the per-architecture constants live in
// cmd/link/internal/<arch>/l.go).
type cfiArch struct {
name string
cfaReg byte // the stack-pointer register the initial CFA rule names
raReg byte // the return-address register
}
var (
cfiAMD64 = cfiArch{"amd64", 7, 16} // RSP, RIP
cfiARM64 = cfiArch{"arm64", 31, 30} // SP (X31), LR (X30)
cfiRISCV64 = cfiArch{"riscv64", 2, 1} // X2 (sp), X1 (ra)
cfiLOONG64 = cfiArch{"loong64", 3, 1} // $r3 (sp), $r1 (ra)
)
// dwarfBuildFrameSection builds a .debug_frame section with CFI for stack // dwarfBuildFrameSection builds a .debug_frame section with CFI for stack
// unwinding. It emits one CIE and one FDE per function, encoding the // unwinding. It emits one CIE and one FDE per function, encoding the
// CFA (Canonical Frame Address) rule changes at each stack-adjustment // CFA (Canonical Frame Address) rule changes at each stack-adjustment
// boundary recorded in FuncLayout.Spadj. // boundary recorded in FuncLayout.Spadj.
func dwarfBuildFrameSection(img *Image) []byte { func dwarfBuildFrameSection(img *Image, cfi cfiArch, ds *dwarfSections) []byte {
var b []byte var b []byte
le := binary.LittleEndian le := binary.LittleEndian
// CIE (Common Information Entry). // CIE (Common Information Entry).
cieStart := len(b) cieStart := len(b)
b = append(b, 0, 0, 0, 0) // length (placeholder) b = append(b, 0, 0, 0, 0) // length (placeholder)
b = le.AppendUint32(b, 0xFFFFFFFF) // CIE marker b = le.AppendUint32(b, 0xFFFFFFFF) // CIE marker
b = append(b, 3) // version (DWARF3, widely supported) b = append(b, 3) // version (DWARF3, widely supported)
b = append(b, 0) // augmentation (empty) b = append(b, 0) // augmentation (empty)
b = appendUleb(b, 1) // code alignment b = appendUleb(b, 1) // code alignment
b = appendSleb(b, -8) // data alignment (-8 for 64-bit) b = appendSleb(b, -8) // data alignment (-8 for 64-bit)
b = appendUleb(b, 16) // return address register (LR on arm64, RIP on amd64) b = appendUleb(b, uint64(cfi.raReg)) // return address register
// Initial CFA rule: DW_CFA_def_cfa (SP, 0) // Initial CFA rule: DW_CFA_def_cfa (SP, 0)
b = append(b, 0x0c) // DW_CFA_def_cfa b = append(b, 0x0c) // DW_CFA_def_cfa
b = appendUleb(b, 31) // register: SP (RSP=7 on amd64, SP=31 on arm64) b = appendUleb(b, uint64(cfi.cfaReg)) // the architecture's stack pointer
b = appendUleb(b, 0) // offset: 0 b = appendUleb(b, 0) // offset: 0
b = append(b, 0) // DW_CFA_nop (padding) b = append(b, 0) // DW_CFA_nop (padding)
// Patch CIE length. // Patch CIE length.
le.PutUint32(b[cieStart:], uint32(len(b)-cieStart-4)) le.PutUint32(b[cieStart:], uint32(len(b)-cieStart-4))
// FDEs (Frame Description Entries) — one per function. // FDEs (Frame Description Entries), one per function.
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
fdeStart := len(b) fdeStart := len(b)
b = append(b, 0, 0, 0, 0) // length (placeholder) b = append(b, 0, 0, 0, 0) // length (placeholder)
b = le.AppendUint32(b, uint32(cieStart)) // CIE pointer (offset from start) b = le.AppendUint32(b, uint32(cieStart)) // CIE pointer (offset from start)
// Initial location: function offset in .text (relocated by linker). // Initial location: function offset in .text, referenced through
// the function's symbol so the linker relocates it.
ds.frameRelocs = append(ds.frameRelocs, dwarfReloc{
off: uint64(fdeStart + 8),
name: fn.Name,
})
b = le.AppendUint64(b, uint64(fn.Offset)) b = le.AppendUint64(b, uint64(fn.Offset))
// Address range: function size. // Address range: function size.
b = le.AppendUint64(b, uint64(fn.Size)) b = le.AppendUint64(b, uint64(fn.Size))
+84 -24
View File
@@ -3,6 +3,17 @@
package asm package asm
import "encoding/binary"
// Absolute 64-bit relocation types for the DWARF address fixups, one per
// supported architecture (the numbers debug/elf carries).
const (
rX8664Abs64 = 1 // R_X86_64_64
rAARCH64Abs64 = 257 // R_AARCH64_ABS64
rRISCVAbs64 = 2 // R_RISCV_64
rLarchAbs64 = 2 // R_LARCH_64
)
// dwarfELFSections holds the laid-out DWARF sections ready for inclusion // dwarfELFSections holds the laid-out DWARF sections ready for inclusion
// in an ELF file. // in an ELF file.
type dwarfELFSections struct { type dwarfELFSections struct {
@@ -11,15 +22,23 @@ type dwarfELFSections struct {
lineOff, lineSize int lineOff, lineSize int
lineStrOff, lineStrSize int lineStrOff, lineStrSize int
frameOff, frameSize int frameOff, frameSize int
// Relocations for .debug_info address references. // .rela.debug_info and .rela.debug_line contents: file offsets and
// entry counts (zero count: the section is absent).
infoRelaOff, infoRelaCount int
lineRelaOff, lineRelaCount int
frameRelaOff, frameRelaCount int
// Relocations for .debug_info address references, offsets relative to
// the section start (what an r_offset in .rela.debug_info means).
infoRelocs []elfDwarfReloc infoRelocs []elfDwarfReloc
// Relocations for .debug_line address references. // Relocations for .debug_line address references, section-relative.
lineRelocs []elfDwarfReloc lineRelocs []elfDwarfReloc
// Relocations for .debug_frame FDE initial locations, section-relative.
frameRelocs []elfDwarfReloc
} }
type elfDwarfReloc struct { type elfDwarfReloc struct {
off uint64 off uint64 // offset within the target section
sym int // symbol index in .symtab sym int // symbol index in .symtab
addend int64 addend int64
} }
@@ -29,8 +48,9 @@ type elfDwarfReloc struct {
// //
// symIdx maps function names to their .symtab indices (needed for relocations // symIdx maps function names to their .symtab indices (needed for relocations
// against .text symbols). The map uses objectName format (pkg.name); the // against .text symbols). The map uses objectName format (pkg.name); the
// DWARF code uses bare function names, so we build a reverse lookup. // DWARF code uses bare function names, so we build a reverse lookup. cfi
func appendDWARFSections(out *[]byte, img *Image, srcFile string, symIdx map[string]int, align func(int)) *dwarfELFSections { // carries the architecture's .debug_frame register conventions.
func appendDWARFSections(out *[]byte, img *Image, srcFile string, symIdx map[string]int, align func(int), cfi cfiArch) *dwarfELFSections {
// Build a lookup from bare function name to symbol index. // Build a lookup from bare function name to symbol index.
nameToIdx := make(map[string]int, len(symIdx)) nameToIdx := make(map[string]int, len(symIdx))
for name, idx := range symIdx { for name, idx := range symIdx {
@@ -45,7 +65,7 @@ func appendDWARFSections(out *[]byte, img *Image, srcFile string, symIdx map[str
} }
nameToIdx[name] = idx nameToIdx[name] = idx
} }
ds := emitDWARF(img, srcFile) ds := emitDWARF(img, srcFile, cfi)
if ds == nil || len(ds.debugAbbrev) == 0 { if ds == nil || len(ds.debugAbbrev) == 0 {
return nil return nil
} }
@@ -68,15 +88,11 @@ func appendDWARFSections(out *[]byte, img *Image, srcFile string, symIdx map[str
align(1) align(1)
result.lineOff = len(*out) result.lineOff = len(*out)
result.lineSize = len(ds.debugLine) result.lineSize = len(ds.debugLine)
lineBase := len(*out)
*out = append(*out, ds.debugLine...) *out = append(*out, ds.debugLine...)
// Patch .debug_line relocations: replace placeholder addresses with
// actual .text offsets via symbol lookup.
for _, dr := range ds.lineRelocs { for _, dr := range ds.lineRelocs {
if idx, ok := nameToIdx[dr.name]; ok { if idx, ok := nameToIdx[dr.name]; ok {
result.lineRelocs = append(result.lineRelocs, elfDwarfReloc{ result.lineRelocs = append(result.lineRelocs, elfDwarfReloc{
off: uint64(lineBase) + dr.off, off: dr.off,
sym: idx, sym: idx,
addend: dr.addend, addend: dr.addend,
}) })
@@ -87,33 +103,77 @@ func appendDWARFSections(out *[]byte, img *Image, srcFile string, symIdx map[str
align(1) align(1)
result.infoOff = len(*out) result.infoOff = len(*out)
result.infoSize = len(ds.debugInfo) result.infoSize = len(ds.debugInfo)
infoBase := len(*out)
*out = append(*out, ds.debugInfo...) *out = append(*out, ds.debugInfo...)
// .debug_frame
if len(ds.debugFrame) > 0 {
align(1)
result.frameOff = len(*out)
result.frameSize = len(ds.debugFrame)
*out = append(*out, ds.debugFrame...)
}
// Patch .debug_info relocations.
for _, dr := range ds.infoRelocs { for _, dr := range ds.infoRelocs {
if idx, ok := nameToIdx[dr.name]; ok { if idx, ok := nameToIdx[dr.name]; ok {
result.infoRelocs = append(result.infoRelocs, elfDwarfReloc{ result.infoRelocs = append(result.infoRelocs, elfDwarfReloc{
off: uint64(infoBase) + dr.off, off: dr.off,
sym: idx, sym: idx,
addend: dr.addend, addend: dr.addend,
}) })
} }
} }
// .debug_frame: the section header declares alignment 8, so the data is
// padded to 8, matching it.
if len(ds.debugFrame) > 0 {
align(8)
result.frameOff = len(*out)
result.frameSize = len(ds.debugFrame)
*out = append(*out, ds.debugFrame...)
for _, dr := range ds.frameRelocs {
if idx, ok := nameToIdx[dr.name]; ok {
result.frameRelocs = append(result.frameRelocs, elfDwarfReloc{
off: dr.off,
sym: idx,
addend: dr.addend,
})
}
}
}
return result return result
} }
// appendDWARFRelas writes the .rela.debug_info and .rela.debug_line section
// bodies from the relocations appendDWARFSections recorded, with the
// architecture's absolute 64-bit relocation type, and records their file
// offsets and entry counts on dw. Called after the DWARF sections
// themselves so the r_offsets (section-relative) need no adjustment.
func appendDWARFRelas(out *[]byte, dw *dwarfELFSections, abs64 uint32, align func(int)) {
le := binary.LittleEndian
write := func(relas []elfDwarfReloc) (off, count int) {
if len(relas) == 0 {
return 0, 0
}
align(8)
off = len(*out)
for _, r := range relas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(abs64))
le.PutUint64(b[16:], uint64(r.addend))
*out = append(*out, b[:]...)
}
return off, len(relas)
}
dw.infoRelaOff, dw.infoRelaCount = write(dw.infoRelocs)
dw.lineRelaOff, dw.lineRelaCount = write(dw.lineRelocs)
dw.frameRelaOff, dw.frameRelaCount = write(dw.frameRelocs)
}
// dwarfSourceName returns the source name the DWARF sections record: the
// image's source path when the assembler captured one, "gasm.s" otherwise.
func dwarfSourceName(img *Image) string {
if img.SourcePath != "" {
return img.SourcePath
}
return "gasm.s"
}
// dwarfSectionNames returns the DWARF section names for the string table. // dwarfSectionNames returns the DWARF section names for the string table.
var dwarfSectionNames = []string{ var dwarfSectionNames = []string{
".debug_abbrev", ".debug_info", ".debug_line", ".debug_line_str", ".debug_abbrev", ".debug_info", ".debug_line", ".debug_line_str",
".debug_frame", ".rela.debug_info", ".rela.debug_line", ".debug_frame", ".rela.debug_info", ".rela.debug_line",
".rela.debug_frame",
} }
+303 -12
View File
@@ -4,11 +4,313 @@
package asm package asm
import ( import (
"bytes"
"encoding/binary"
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
// ulebIter reads ULEB128 values, the .debug_abbrev and line-header
// encoding.
type ulebIter struct {
b []byte
i int
}
func (r *ulebIter) uleb(t *testing.T) uint64 {
t.Helper()
v, n := binary.Uvarint(r.b[r.i:])
if n <= 0 {
t.Fatalf("bad ULEB at %d", r.i)
}
r.i += n
return v
}
func (r *ulebIter) byteAt(t *testing.T) byte {
t.Helper()
if r.i >= len(r.b) {
t.Fatalf("read past end at %d", r.i)
}
c := r.b[r.i]
r.i++
return c
}
func (r *ulebIter) uint32At(t *testing.T) uint32 {
t.Helper()
v := binary.LittleEndian.Uint32(r.b[r.i:])
r.i += 4
return v
}
// sleb reads a signed LEB128, the DWARF encoding (sign-extended two's
// complement, not Go's zigzag varint).
func (r *ulebIter) sleb(t *testing.T) int64 {
t.Helper()
var v int64
var shift uint
for {
c := r.byteAt(t)
v |= int64(c&0x7f) << shift
shift += 7
if c&0x80 == 0 {
if c&0x40 != 0 {
v |= -1 << shift
}
return v
}
}
}
// dwarfAttr is one attribute/form pair of an abbreviation.
type dwarfAttr struct{ attr, form uint64 }
// dwarfAbbrev is one parsed abbreviation declaration.
type dwarfAbbrev struct {
code uint64
tag uint64
children bool
attrs []dwarfAttr
}
// parseAbbrevs walks a .debug_abbrev table: abbreviation code, tag,
// children flag, then attr/form ULEB pairs terminated by a double zero.
func parseAbbrevs(t *testing.T, b []byte) map[uint64]dwarfAbbrev {
t.Helper()
out := map[uint64]dwarfAbbrev{}
r := &ulebIter{b: b}
for {
code := r.uleb(t)
if code == 0 {
return out
}
ab := dwarfAbbrev{code: code, tag: r.uleb(t)}
ab.children = r.byteAt(t) == 1
for {
attr := r.uleb(t)
form := r.uleb(t)
if attr == 0 && form == 0 {
break
}
if attr == 0 || form == 0 {
t.Fatalf("abbrev %d: half-terminated attr/form pair (%d, %d)", code, attr, form)
}
ab.attrs = append(ab.attrs, dwarfAttr{attr, form})
}
out[code] = ab
}
}
func eqAttrs(t *testing.T, ab dwarfAbbrev, want []dwarfAttr) {
t.Helper()
if len(ab.attrs) != len(want) {
t.Fatalf("abbrev %d attrs = %v, want %v", ab.code, ab.attrs, want)
}
for i, w := range want {
if ab.attrs[i] != w {
t.Fatalf("abbrev %d attr %d = (%#x, %#x), want (%#x, %#x)", ab.code, i, ab.attrs[i].attr, ab.attrs[i].form, w.attr, w.form)
}
}
}
// TestDwarfAbbrevTable walks the abbreviation table as a consumer does and
// checks the attribute/form sets against the constants the toolchain uses
// (cmd/internal/dwarf/dwarf_defs.go). A wrong constant here renames an
// attribute (0x1b is comp_dir, not low_pc; 0x29 and 0x37 are bounds and
// count) and a wrong form desynchronises the DIE parse: 0x25 is strx1, one
// byte, where the writer emits four for a section offset.
func TestDwarfAbbrevTable(t *testing.T) {
abbrev := dwarfAbbrevTable()
if len(abbrev) == 0 {
t.Fatal("empty abbrev table")
}
// Must end with a zero byte (end of table).
if abbrev[len(abbrev)-1] != 0 {
t.Fatalf("abbrev table last byte = %d, want 0", abbrev[len(abbrev)-1])
}
abs := parseAbbrevs(t, abbrev)
if len(abs) != 2 {
t.Fatalf("abbreviations = %d, want 2", len(abs))
}
cu, ok := abs[1]
if !ok {
t.Fatal("missing abbreviation 1 (compile unit)")
}
if cu.tag != dwTagCompUnit || !cu.children {
t.Errorf("abbrev 1: tag %#x children %v, want compile unit with children", cu.tag, cu.children)
}
eqAttrs(t, cu, []dwarfAttr{
{dwAtLowPC, dwFormAddr},
{dwAtHighPC, dwFormData8},
{dwAtStmtList, dwFormSecOff},
{dwAtName, dwFormString},
})
sp, ok := abs[2]
if !ok {
t.Fatal("missing abbreviation 2 (subprogram)")
}
if sp.tag != dwTagSubprog || sp.children {
t.Errorf("abbrev 2: tag %#x children %v, want subprogram without children", sp.tag, sp.children)
}
eqAttrs(t, sp, []dwarfAttr{
{dwAtName, dwFormString},
{dwAtLowPC, dwFormAddr},
{dwAtHighPC, dwFormData8},
{dwAtFrameBase, dwFormExprloc},
{dwAtDeclFile, dwFormData1},
{dwAtDeclLine, dwFormData1},
{dwAtExternal, 0x0c}, // DW_FORM_flag
})
}
// TestDwarfLineHeaderV5 parses the .debug_line header under DWARF5 rules:
// the directory and file tables are format-descriptor lists, not the
// DWARF2-4 shape of null-terminated strings, and the file entry references
// the source name through .debug_line_str.
func TestDwarfLineHeaderV5(t *testing.T) {
src := `#include "textflag.h"
TEXT ·add(SB), NOSPLIT, $0-24
MOVQ a+0(FP), AX
MOVQ b+8(FP), BX
ADDQ BX, AX
MOVQ AX, ret+16(FP)
RET
`
f, errs := parser.Parse("test_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
ds := emitDWARF(img, "test_amd64.s", cfiAMD64)
r := &ulebIter{b: ds.debugLine}
r.uint32At(t) // unit_length
if v := binary.LittleEndian.Uint16(ds.debugLine[4:]); v != 5 {
t.Fatalf("version = %d, want 5", v)
}
r.i = 6
r.byteAt(t) // address_size
r.byteAt(t) // segment_selector_size
r.uint32At(t) // header_length
r.byteAt(t) // minimum_instruction_length
r.byteAt(t) // maximum_ops_per_instruction
r.byteAt(t) // default_is_stmt
r.byteAt(t) // line_base
r.byteAt(t) // line_range
opcodeBase := r.byteAt(t)
for range int(opcodeBase) - 1 {
r.byteAt(t) // standard opcode lengths
}
// Directory table (DWARF5 §6.2.4).
if n := r.byteAt(t); n != 1 {
t.Fatalf("directory_entry_format_count = %d, want 1", n)
}
if lnct := r.uleb(t); lnct != dwLnctPath {
t.Errorf("directory content type = %#x, want DW_LNCT_path", lnct)
}
if form := r.uleb(t); form != dwFormLineStrp {
t.Errorf("directory form = %#x, want DW_FORM_line_strp", form)
}
if n := r.uleb(t); n != 1 {
t.Fatalf("directories_count = %d, want 1", n)
}
if off := r.uint32At(t); off != 0 {
t.Errorf("compilation directory line_strp = %d, want 0 (the empty string)", off)
}
// File table (DWARF5 §6.2.5).
if n := r.byteAt(t); n != 2 {
t.Fatalf("file_name_entry_format_count = %d, want 2", n)
}
if lnct := r.uleb(t); lnct != dwLnctPath {
t.Errorf("file content type = %#x, want DW_LNCT_path", lnct)
}
if form := r.uleb(t); form != dwFormLineStrp {
t.Errorf("file path form = %#x, want DW_FORM_line_strp", form)
}
if lnct := r.uleb(t); lnct != dwLnctDirIndex {
t.Errorf("file content type = %#x, want DW_LNCT_directory_index", lnct)
}
if form := r.uleb(t); form != dwFormUdata {
t.Errorf("file dir-index form = %#x, want DW_FORM_udata", form)
}
if n := r.uleb(t); n != 1 {
t.Fatalf("file_names_count = %d, want 1", n)
}
strOff := r.uint32At(t)
if dirIdx := r.uleb(t); dirIdx != 0 {
t.Errorf("file directory index = %d, want 0", dirIdx)
}
// The file entry's line_strp must resolve to the source name.
end := int(strOff) + len("test_amd64.s")
if int(strOff) >= len(ds.debugLineStr) || !bytes.Equal(ds.debugLineStr[strOff:end], []byte("test_amd64.s")) {
t.Errorf("file entry line_strp %d does not name the source: %q", strOff, ds.debugLineStr)
}
// The fixed header fields: address_size 8 and a header_length that
// points just past the file table (the patch site is offset 8 in the
// v5 header, and the field counts from its own end).
if ds.debugLine[6] != 8 || ds.debugLine[7] != 0 {
t.Errorf("address_size/segment_selector = %d/%d, want 8/0", ds.debugLine[6], ds.debugLine[7])
}
if hl := binary.LittleEndian.Uint32(ds.debugLine[8:]); hl != uint32(r.i-12) {
t.Errorf("header_length = %d, want %d (the byte after the file table is %d)", hl, r.i-12, r.i)
}
}
// TestDwarfFrameCIEArch checks the shared CIE carries each architecture's
// stack-pointer and return-address registers: the values the Go linker
// writes (cmd/link/internal/<arch>/l.go dwarfRegSP/dwarfRegLR).
func TestDwarfFrameCIEArch(t *testing.T) {
for _, tc := range []struct {
name string
cfi cfiArch
}{
{"amd64", cfiAMD64},
{"arm64", cfiARM64},
{"riscv64", cfiRISCV64},
{"loong64", cfiLOONG64},
} {
frame := dwarfBuildFrameSection(&Image{}, tc.cfi, &dwarfSections{})
r := &ulebIter{b: frame}
r.uint32At(t) // length
if cid := r.uint32At(t); cid != 0xFFFFFFFF {
t.Errorf("%s: CIE id = %#x, want 0xffffffff", tc.name, cid)
}
if v := r.byteAt(t); v != 3 {
t.Errorf("%s: CIE version = %d, want 3", tc.name, v)
}
if aug := r.byteAt(t); aug != 0 {
t.Errorf("%s: CIE augmentation = %d, want 0", tc.name, aug)
}
if ca := r.uleb(t); ca != 1 {
t.Errorf("%s: code alignment = %d, want 1", tc.name, ca)
}
if da := r.sleb(t); da != -8 {
t.Errorf("%s: data alignment = %d, want -8 (signed LEB128, not zigzag)", tc.name, da)
}
if ra := r.uleb(t); ra != uint64(tc.cfi.raReg) {
t.Errorf("%s: return-address register = %d, want %d", tc.name, ra, tc.cfi.raReg)
}
if op := r.byteAt(t); op != 0x0c {
t.Errorf("%s: expected DW_CFA_def_cfa, got opcode %#x", tc.name, op)
}
if cfa := r.uleb(t); cfa != uint64(tc.cfi.cfaReg) {
t.Errorf("%s: CFA register = %d, want %d", tc.name, cfa, tc.cfi.cfaReg)
}
if off := r.uleb(t); off != 0 {
t.Errorf("%s: CFA offset = %d, want 0", tc.name, off)
}
}
}
func TestEmitDWARF(t *testing.T) { func TestEmitDWARF(t *testing.T) {
src := `#include "textflag.h" src := `#include "textflag.h"
TEXT ·add(SB), NOSPLIT, $0-24 TEXT ·add(SB), NOSPLIT, $0-24
@@ -27,7 +329,7 @@ TEXT ·add(SB), NOSPLIT, $0-24
t.Fatalf("assemble: %v", err) t.Fatalf("assemble: %v", err)
} }
ds := emitDWARF(img, "test_amd64.s") ds := emitDWARF(img, "test_amd64.s", cfiAMD64)
// .debug_abbrev must not be empty and must start with abbrev code 1. // .debug_abbrev must not be empty and must start with abbrev code 1.
if len(ds.debugAbbrev) == 0 { if len(ds.debugAbbrev) == 0 {
@@ -68,14 +370,3 @@ TEXT ·add(SB), NOSPLIT, $0-24
t.Fatal("no .debug_info relocations") t.Fatal("no .debug_info relocations")
} }
} }
func TestDwarfAbbrevTable(t *testing.T) {
abbrev := dwarfAbbrevTable()
if len(abbrev) == 0 {
t.Fatal("empty abbrev table")
}
// Must end with a zero byte (end of table).
if abbrev[len(abbrev)-1] != 0 {
t.Fatalf("abbrev table last byte = %d, want 0", abbrev[len(abbrev)-1])
}
}
+443 -1
View File
@@ -12,6 +12,7 @@ import (
"path/filepath" "path/filepath"
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
@@ -52,7 +53,7 @@ func elfTestImage(t *testing.T) *Image {
} }
// TestAssembleFileExternals checks that a reference to a symbol no GLOBL // TestAssembleFileExternals checks that a reference to a symbol no GLOBL
// defines is recorded as an external relocation instead of failing — the // defines is recorded as an external relocation instead of failing; the
// raw image leaves the displacement zero, the object emitters carry it. // raw image leaves the displacement zero, the object emitters carry it.
func TestAssembleFileExternals(t *testing.T) { func TestAssembleFileExternals(t *testing.T) {
img := elfTestImage(t) img := elfTestImage(t)
@@ -211,6 +212,75 @@ func TestELFObject(t *testing.T) {
} }
} }
// TestELFObjectTLSGuardReloc checks that a non-NOSPLIT function's stack
// guard carries an R_X86_64_TPOFF32 relocation against the null symbol in
// .rela.text. The serialisation must honour the record's type field: a
// hardcoded R_X86_64_PC32 mislinks the TLS load as an ordinary
// PC-relative reference.
func TestELFObjectTLSGuardReloc(t *testing.T) {
f, errs := parser.Parse("g_amd64.s", `
#include "textflag.h"
TEXT ·grow(SB), $0
CALL ·other(SB)
RET
TEXT ·other(SB), NOSPLIT, $0
RET
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
var haveTLS bool
for _, fn := range img.Funcs {
for _, r := range fn.Relocs {
if r.Kind == RelTLSLE {
haveTLS = true
}
}
}
if !haveTLS {
t.Fatal("test source produced no RelTLSLE relocation")
}
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaSec := ef.Section(".rela.text")
if relaSec == nil {
t.Fatal("missing .rela.text")
}
raw, err := relaSec.Data()
if err != nil {
t.Fatal(err)
}
found := false
for i := 0; i+24 <= len(raw); i += 24 {
e := raw[i:]
info := binary.LittleEndian.Uint64(e[8:])
typ := info & 0xffffffff
sym := int(info >> 32)
if typ == uint64(elf.R_X86_64_TPOFF32) {
found = true
if sym != 0 {
t.Errorf("TPOFF32 relocation against symbol %d, want 0 (the null symbol)", sym)
}
}
}
if !found {
t.Errorf("no R_X86_64_TPOFF32 relocation in .rela.text (%d bytes)", len(raw))
}
}
// TestELFObjectNoRelocations checks a file with no static-symbol references // TestELFObjectNoRelocations checks a file with no static-symbol references
// emits a valid object without a .rela.text section. // emits a valid object without a .rela.text section.
func TestELFObjectNoRelocations(t *testing.T) { func TestELFObjectNoRelocations(t *testing.T) {
@@ -253,6 +323,238 @@ TEXT ·nop(SB), NOSPLIT, $0
} }
} }
// elfSectionHeaderCount returns the e_shnum the ELF header declares.
func elfSectionHeaderCount(t *testing.T, obj []byte) int {
t.Helper()
return int(binary.LittleEndian.Uint16(obj[60:]))
}
// checkELFSectionAccounting verifies the number of section headers the
// writer physically laid out equals e_shnum: every DWARF section written
// after .shstrtab must be counted, or the last ones (always .debug_frame)
// are invisible to every consumer, debug/elf included.
func checkELFSectionAccounting(t *testing.T, obj []byte) {
t.Helper()
shoff := int(binary.LittleEndian.Uint64(obj[40:]))
shentsize := int(binary.LittleEndian.Uint16(obj[58:]))
shnum := elfSectionHeaderCount(t, obj)
if shentsize != 64 {
t.Fatalf("e_shentsize = %d, want 64", shentsize)
}
if (len(obj)-shoff)%shentsize != 0 {
t.Fatalf("section header table is not a whole number of entries: shoff=%d len=%d", shoff, len(obj))
}
if present := (len(obj) - shoff) / shentsize; present != shnum {
t.Errorf("e_shnum = %d but %d section headers are laid out", shnum, present)
}
}
// TestELFDWARFSectionAccounting runs the header accounting check over all
// four architecture emitters, and additionally checks the .debug_frame
// section is visible (its data aligned as its header declares).
func TestELFDWARFSectionAccounting(t *testing.T) {
parse := func(name, src string) *ast.File {
f, errs := parser.Parse(name, src)
if len(errs) > 0 {
t.Fatalf("parse %s: %v", name, errs)
}
return f
}
cases := []struct {
name string
img *Image
emit func(*Image) ([]byte, error)
}{
{"amd64", elfTestImage(t), (*Image).ELFObject},
{"arm64", mustImage(t, func() (*Image, error) {
return AssembleFileARM64(parse("k_arm64.s", `
#include "textflag.h"
TEXT ·add(SB), NOSPLIT, $0-24
MOVD a+0(FP), R4
MOVD b+8(FP), R5
ADD R5, R4, R4
MOVD R4, ret+16(FP)
RET
`))
}), (*Image).ELFAARCH64Object},
{"riscv64", mustImage(t, func() (*Image, error) {
return AssembleFileRISCV(parse("k_riscv64.s", `
#include "textflag.h"
TEXT ·sb(SB), NOSPLIT, $0-0
MOV $answer<>(SB), X10
RET
GLOBL answer<>(SB), RODATA, $8
DATA answer<>+0(SB)/8, $42
`))
}), (*Image).ELFRISCVObject},
{"loong64", mustImage(t, func() (*Image, error) {
return AssembleFileLOONG64(parse("k_loong64.s", `
#include "textflag.h"
TEXT ·add(SB), NOSPLIT, $0-24
MOVV a+0(FP), R4
MOVV b+8(FP), R5
ADDV R5, R4, R4
MOVV R4, ret+16(FP)
RET
`))
}), (*Image).ELFLOONG64Object},
}
for _, tc := range cases {
obj, err := tc.emit(tc.img)
if err != nil {
t.Fatalf("%s: emit: %v", tc.name, err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("%s: parse emitted object: %v", tc.name, err)
}
frame := ef.Section(".debug_frame")
if frame == nil {
t.Errorf("%s: .debug_frame invisible to debug/elf (e_shnum too small?)", tc.name)
ef.Close()
continue
}
if frame.Offset%8 != 0 || frame.Addralign != 8 {
t.Errorf("%s: .debug_frame offset %d align %d, want offset%%8==0 align 8", tc.name, frame.Offset, frame.Addralign)
}
ef.Close()
}
}
func mustImage(t *testing.T, f func() (*Image, error)) *Image {
t.Helper()
img, err := f()
if err != nil {
t.Fatal(err)
}
return img
}
// TestELFDWARFRelocations checks the .rela.debug_info and .rela.debug_line
// sections exist and carry absolute 64-bit relocations against the
// function symbols, with r_offsets inside their target sections.
func TestELFDWARFRelocations(t *testing.T) {
img := elfTestImage(t)
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
// The DWARF must record the assembled file's path (threaded through
// Image.SourcePath), not a placeholder name.
info, err := ef.Section(".debug_info").Data()
if err != nil {
t.Fatal(err)
}
if img.SourcePath != "t_amd64.s" || !bytes.Contains(info, []byte(img.SourcePath)) {
t.Errorf("DWARF compilation unit does not name the source %q", img.SourcePath)
}
for _, tc := range []struct {
rela string
target string
want uint32
}{
{".rela.debug_info", ".debug_info", rX8664Abs64},
{".rela.debug_line", ".debug_line", rX8664Abs64},
{".rela.debug_frame", ".debug_frame", rX8664Abs64},
} {
rs := ef.Section(tc.rela)
if rs == nil {
t.Fatalf("missing %s", tc.rela)
}
if rs.Type != elf.SHT_RELA {
t.Errorf("%s: type %v, want SHT_RELA", tc.rela, rs.Type)
}
target := ef.Section(tc.target)
if target == nil {
t.Fatalf("missing %s", tc.target)
}
if rs.Link == 0 || ef.Sections[rs.Info] != target {
t.Errorf("%s: link %d info %d, want the symtab and %s", tc.rela, rs.Link, rs.Info, tc.target)
}
b, err := rs.Data()
if err != nil {
t.Fatal(err)
}
// .debug_line has one address per function; .debug_info adds the
// compile unit's own low_pc.
want := len(img.Funcs)
if tc.target == ".debug_info" {
want++
}
if len(b)/24 != want {
t.Errorf("%s: %d entries, want %d", tc.rela, len(b)/24, want)
}
for i := 0; i+24 <= len(b); i += 24 {
r_offset := binary.LittleEndian.Uint64(b[i:])
info := binary.LittleEndian.Uint64(b[i+8:])
typ := uint32(info)
sym := int(info >> 32)
if typ != tc.want {
t.Errorf("%s entry %d: type %d, want R_X86_64_64 (%d)", tc.rela, i/24, typ, tc.want)
}
if r_offset >= uint64(target.Size) {
t.Errorf("%s entry %d: r_offset %d outside %s (%d bytes)", tc.rela, i/24, r_offset, tc.target, target.Size)
}
if sym == 0 {
t.Errorf("%s entry %d: against the null symbol", tc.rela, i/24)
}
}
}
}
// TestELFDataOnly checks a source with GLOBL data and no TEXT emits a valid
// ELF object: the DWARF compilation unit of a code-less image has no
// function to relocate against and must not reach for one.
func TestELFDataOnly(t *testing.T) {
f, errs := parser.Parse("d0_amd64.s", `
GLOBL table<>(SB), RODATA, $8
DATA table<>+0(SB)/8, $12345
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
found := false
for _, s := range syms {
if s.Name == "table" && s.Size == 8 {
found = true
}
}
if !found {
t.Errorf("data symbol table missing: %v", syms)
}
if ef.Section(".rela.debug_info") != nil || ef.Section(".rela.debug_line") != nil {
t.Error("data-only image must not emit DWARF address relocations")
}
}
// TestELFLinkAndRun is the end-to-end check: assemble the test functions, // TestELFLinkAndRun is the end-to-end check: assemble the test functions,
// link the emitted object with a C driver that defines the external symbol, // link the emitted object with a C driver that defines the external symbol,
// and run the result. Skipped when no C compiler is available. // and run the result. Skipped when no C compiler is available.
@@ -307,4 +609,144 @@ int main(void) {
if got := string(run); got != "42 42 7\n" { if got := string(run); got != "42 42 7\n" {
t.Errorf("output %q, want \"42 42 7\\n\"", got) t.Errorf("output %q, want \"42 42 7\\n\"", got)
} }
// The DWARF addresses must have resolved at link time: the .debug_info
// placeholders were carried by .rela.debug_info, so every subprogram's
// low_pc must now equal its linked symbol address.
bin, err := os.ReadFile(appPath)
if err != nil {
t.Fatal(err)
}
lef, err := elf.NewFile(bytes.NewReader(bin))
if err != nil {
t.Fatalf("parse linked binary: %v", err)
}
defer lef.Close()
syms, err := lef.Symbols()
if err != nil {
t.Fatal(err)
}
addrByName := map[string]uint64{}
for _, s := range syms {
if elf.ST_TYPE(s.Info) == elf.STT_FUNC && s.Value != 0 {
addrByName[s.Name] = s.Value
}
}
lowPCs := dwarfSubprogramLowPCs(t, lef)
if len(lowPCs) == 0 {
t.Fatal("no subprogram DW_AT_low_pc parsed from the linked binary")
}
for name, pc := range lowPCs {
addr, ok := addrByName[name]
if !ok {
t.Errorf("subprogram %q not in the linked symbol table", name)
continue
}
if pc != addr {
t.Errorf("subprogram %q: DW_AT_low_pc = %#x, linked address %#x (DWARF relocation unresolved)", name, pc, addr)
}
}
}
// dwarfSubprogramLowPCs walks the linked binary's .debug_info with its own
// .debug_abbrev and returns each DW_TAG_subprogram's DW_AT_low_pc by name.
func dwarfSubprogramLowPCs(t *testing.T, ef *elf.File) map[string]uint64 {
t.Helper()
abbrevSec := ef.Section(".debug_abbrev")
infoSec := ef.Section(".debug_info")
if abbrevSec == nil || infoSec == nil {
t.Fatal("linked binary lacks .debug_abbrev or .debug_info")
}
abbrev, err := abbrevSec.Data()
if err != nil {
t.Fatal(err)
}
info, err := infoSec.Data()
if err != nil {
t.Fatal(err)
}
abs := parseAbbrevs(t, abbrev)
le := binary.LittleEndian
out := map[string]uint64{}
r := &ulebIter{b: info}
r.uint32At(t) // unit_length
if v := le.Uint16(info[4:]); v != 5 {
t.Fatalf(".debug_info version %d, want 5", v)
}
r.i = 6
r.byteAt(t) // unit_type
r.byteAt(t) // address_size
r.uint32At(t) // debug_abbrev_offset
var name string
var lowPC uint64
for r.i < len(r.b) {
code := r.uleb(t)
if code == 0 {
continue // end of the CU's children
}
ab, ok := abs[code]
if !ok {
t.Fatalf("unknown abbreviation code %d", code)
}
name, lowPC = "", 0
for _, a := range ab.attrs {
switch a.attr {
case dwAtName:
readFormKeep(t, r, a.form, &name, nil)
case dwAtLowPC:
readFormKeep(t, r, a.form, nil, &lowPC)
default:
readFormSkip(t, r, a.form)
}
}
if ab.tag == dwTagSubprog && name != "" {
out[name] = lowPC
}
}
return out
}
// readFormKeep reads one DIE attribute value, keeping a string or an
// address into the pointer it was given (nil keeps nothing).
func readFormKeep(t *testing.T, r *ulebIter, form uint64, name *string, addr *uint64) {
t.Helper()
switch form {
case dwFormString:
end := r.i
for end < len(r.b) && r.b[end] != 0 {
end++
}
if name != nil {
*name = string(r.b[r.i:end])
}
r.i = end + 1
case dwFormAddr:
if addr != nil {
*addr = binary.LittleEndian.Uint64(r.b[r.i:])
}
r.i += 8
default:
readFormSkip(t, r, form)
}
}
func readFormSkip(t *testing.T, r *ulebIter, form uint64) {
t.Helper()
switch form {
case dwFormString:
for r.i < len(r.b) && r.b[r.i] != 0 {
r.i++
}
r.i++
case dwFormAddr, dwFormData8:
r.i += 8
case dwFormSecOff:
r.i += 4
case dwFormExprloc:
r.i += int(r.uleb(t))
case dwFormData1, 0x0c:
r.i++
default:
t.Fatalf("unsupported form %#x", form)
}
} }
+81 -20
View File
@@ -15,8 +15,9 @@ const (
// AArch64 relocation types (the ELF psABI). // AArch64 relocation types (the ELF psABI).
rArm64PrelPgHi21 = 275 // R_AARCH64_ADR_PREL_PG_HI21 (ADRP page) rArm64PrelPgHi21 = 275 // R_AARCH64_ADR_PREL_PG_HI21 (ADRP page)
rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD/STR/LDR page offset) rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD page offset)
rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction) rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction)
rArm64Ldst64Lo12NC = 286 // R_AARCH64_LDST64_ABS_LO12_NC (64-bit LDR/STR page offset)
) )
// ELFAARCH64Object returns the image as an ELF64 relocatable object file for // ELFAARCH64Object returns the image as an ELF64 relocatable object file for
@@ -80,8 +81,19 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
} }
// Build relocations. Each SB reference is an ADRP pair: // Build relocations. Each SB reference is an ADRP pair:
// ADRP Rd, 0 → R_AARCH64_ADR_PREL_PG_HI21 // ADRP Rd, 0 → R_AARCH64_ADR_PREL_PG_HI21 at the ADRP
// ADD/LDR/STR → R_AARCH64_ADD_ABS_LO12_NC // ADD → R_AARCH64_ADD_ABS_LO12_NC at the ADD word
// LDR/STR X → R_AARCH64_LDST64_ABS_LO12_NC at the LDR/STR word
// BL → R_AARCH64_CALL26
// cmd/link's own conversion emits the HI21 at sectoff and the LO12 at
// sectoff+4 (cmd/link/internal/arm64/asm.go), so the ADD or load word
// carries the page-offset relocation, never a second HI21. The
// assembler records two RelArm64Addr relocs per ADRP+ADD pair (one per
// word), so the second of the pair is consumed here.
// Addends stay raw: ADR_PREL_PG_HI21 and the ABS_LO12_NC forms resolve
// against S+A, and CALL26 branches take the branch instruction's own
// place as the PC-relative base, so subtracting the field width (the
// amd64 R_PCREL convention) would misplace every branch by 4 bytes.
type elfRela struct { type elfRela struct {
off uint64 off uint64
typ uint32 typ uint32
@@ -90,26 +102,34 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
} }
var relas []elfRela var relas []elfRela
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
for _, r := range fn.Relocs { for i := 0; i < len(fn.Relocs); i++ {
r := fn.Relocs[i]
idx, ok := symIdx[r.Name] idx, ok := symIdx[r.Name]
if !ok { if !ok {
return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name) return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name)
} }
var typ uint32 switch r.Kind {
switch { case RelArm64Branch:
case r.Kind == RelArm64Branch: relas = append(relas, elfRela{
typ = rArm64Call26 off: uint64(fn.Offset + r.Off), typ: rArm64Call26, sym: idx, addend: r.Addend,
case r.Kind == RelArm64Addr && r.Off%4 == 4: })
typ = rArm64AddAbsLo12NC case RelArm64Addr:
// ADRP+ADD: the pair's second reloc (at Off+4) is the
// assembler's twin of the same pair; skip it.
relas = append(relas,
elfRela{off: uint64(fn.Offset + r.Off), typ: rArm64PrelPgHi21, sym: idx, addend: r.Addend},
elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rArm64AddAbsLo12NC, sym: idx, addend: r.Addend},
)
i++
case RelArm64LDST64:
// ADRP+LDR/STR: one assembler reloc covers the pair.
relas = append(relas,
elfRela{off: uint64(fn.Offset + r.Off), typ: rArm64PrelPgHi21, sym: idx, addend: r.Addend},
elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rArm64Ldst64Lo12NC, sym: idx, addend: r.Addend},
)
default: default:
typ = rArm64PrelPgHi21 return nil, fmt.Errorf("relocation kind %v unsupported in ELF emission", r.Kind)
} }
relas = append(relas, elfRela{
off: uint64(fn.Offset + r.Off),
typ: typ,
sym: idx,
addend: r.Addend - int64(r.After-r.Off),
})
} }
} }
@@ -184,15 +204,32 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
shstrOff := len(out) shstrOff := len(out)
out = append(out, stSections.bytes()...) out = append(out, stSections.bytes()...)
// DWARF debug sections. // DWARF debug sections; the address placeholders they leave are carried
// as .rela.debug_info/.rela.debug_line entries the system linker applies.
dwAlign := func(n int) { dwAlign := func(n int) {
for len(out)%n != 0 { for len(out)%n != 0 {
out = append(out, 0) out = append(out, 0)
} }
} }
dw := appendDWARFSections(&out, img, "gasm.s", symIdx, dwAlign) dw := appendDWARFSections(&out, img, dwarfSourceName(img), symIdx, dwAlign, cfiARM64)
dwarfStart := 0 // section index of .debug_abbrev, set when DWARF is present
if dw != nil { if dw != nil {
nSections += 4 // Five DWARF sections: .debug_abbrev, .debug_info, .debug_line,
// .debug_line_str and .debug_frame (the CIE is unconditional, so
// the frame section is always present), plus the relocation
// sections below when they carry entries.
dwarfStart = nSections
nSections += 5
appendDWARFRelas(&out, dw, rAARCH64Abs64, dwAlign)
if dw.infoRelaCount > 0 {
nSections++
}
if dw.lineRelaCount > 0 {
nSections++
}
if dw.frameRelaCount > 0 {
nSections++
}
} }
align(8) align(8)
@@ -221,13 +258,37 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24) putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
} }
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0) putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil { if dw != nil {
// secIdx is a running section index: each putSh below emits the
// next header, and the sh_info of a .rela section names the index
// of the section it relocates.
secIdx := dwarfStart
putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0) putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0)
secIdx++
putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0) putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0)
secInfoIdx := secIdx
secIdx++
if dw.infoRelaCount > 0 {
putSh(".rela.debug_info", shtRela, 0, dw.infoRelaOff, 24*dw.infoRelaCount, secSymtab, secInfoIdx, 8, 24)
secIdx++
}
putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0) putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0)
secLineIdx := secIdx
secIdx++
if dw.lineRelaCount > 0 {
putSh(".rela.debug_line", shtRela, 0, dw.lineRelaOff, 24*dw.lineRelaCount, secSymtab, secLineIdx, 8, 24)
secIdx++
}
putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0) putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0)
secIdx++
if dw.frameSize > 0 { if dw.frameSize > 0 {
putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0) putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0)
secFrameIdx := secIdx
secIdx++
if dw.frameRelaCount > 0 {
putSh(".rela.debug_frame", shtRela, 0, dw.frameRelaOff, 24*dw.frameRelaCount, secSymtab, secFrameIdx, 8, 24)
}
} }
} }
+58 -1
View File
@@ -6,6 +6,7 @@ package asm
import ( import (
"bytes" "bytes"
"debug/elf" "debug/elf"
"encoding/binary"
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
@@ -28,6 +29,7 @@ TEXT ·add(SB), NOSPLIT, $0-24
TEXT ·getanswer(SB), NOSPLIT, $0-8 TEXT ·getanswer(SB), NOSPLIT, $0-8
MOVD answer<>(SB), R4 MOVD answer<>(SB), R4
MOVD $answer<>(SB), R5
MOVD R4, ret+0(FP) MOVD R4, ret+0(FP)
RET RET
@@ -102,8 +104,63 @@ DATA answer<>+0(SB)/8, $42
// Check that .rela.text exists (getanswer has SB reference). // Check that .rela.text exists (getanswer has SB reference).
relaText := ef.Section(".rela.text") relaText := ef.Section(".rela.text")
if relaText == nil { if relaText == nil {
t.Error("missing .rela.text section") t.Fatal("missing .rela.text section")
} }
// The SB references of getanswer form two ADRP pairs: the load
// (MOVD answer<>(SB), R4) is ADRP+LDR carrying HI21 at the ADRP and
// LDST64_ABS_LO12_NC at the LDR word, and the address-of
// (MOVD $answer<>(SB), R5) is ADRP+ADD carrying HI21 and
// ADD_ABS_LO12_NC. cmd/link's own conversion emits exactly this
// sectoff / sectoff+4 pairing; a second HI21 at the ADD or LDR word
// corrupts the pair.
raw, err := relaText.Data()
if err != nil {
t.Fatal(err)
}
if len(raw)%24 != 0 || len(raw)/24 != 4 {
t.Fatalf(".rela.text has %d bytes, want four 24-byte entries", len(raw))
}
wantRela := []struct {
typ elf.R_AARCH64
off uint64 // relative to the getanswer function start
}{
{elf.R_AARCH64_ADR_PREL_PG_HI21, 0},
{elf.R_AARCH64_LDST64_ABS_LO12_NC, 4},
{elf.R_AARCH64_ADR_PREL_PG_HI21, 8},
{elf.R_AARCH64_ADD_ABS_LO12_NC, 12},
}
getanswer := byNameElf(t, ef, "getanswer")
for i, w := range wantRela {
e := raw[i*24 : (i+1)*24]
off := binary.LittleEndian.Uint64(e[0:])
info := binary.LittleEndian.Uint64(e[8:])
typ := elf.R_AARCH64(info & 0xffffffff)
sym := int(info >> 32)
if typ != w.typ || off != getanswer.Value+w.off {
t.Errorf("reloc %d: type %v off %d, want %v at %d", i, typ, off, w.typ, getanswer.Value+w.off)
}
if sym != 3 { // NULL, .text, .data, then the first local: answer
t.Errorf("reloc %d: symbol index %d, want 3 (answer)", i, sym)
}
}
}
// byNameElf returns the symbol table entry for name from the raw .symtab,
// which carries every entry including the null and section symbols in order.
func byNameElf(t *testing.T, ef *elf.File, name string) elf.Symbol {
t.Helper()
syms, err := ef.Symbols()
if err != nil {
t.Fatalf("symbols: %v", err)
}
for _, s := range syms {
if s.Name == name {
return s
}
}
t.Fatalf("symbol %q not found", name)
return elf.Symbol{}
} }
// TestELFAARCH64ObjectNoRelocations checks the ELF output when there are no // TestELFAARCH64ObjectNoRelocations checks the ELF output when there are no
+55 -5
View File
@@ -13,9 +13,16 @@ import (
const ( const (
emLOONGARCH = 258 // EM_LOONGARCH emLOONGARCH = 258 // EM_LOONGARCH
// EF_LOONGARCH_ABI_DOUBLE_FLOAT | EF_LOONGARCH_OBJABI_V1: the flags the
// Go toolchain writes (cmd/link/internal/ld/elf.go: Flags = 0x43 for
// Loong64). System linkers refuse to merge ET_REL objects whose float
// ABI differs, so 0 (soft-float) would make the object unlinkable.
efLarchAbiDoubleObjV1 = 0x43
// LoongArch relocation types (the ELF psABI). // LoongArch relocation types (the ELF psABI).
rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i) rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i)
rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st) rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st)
rLarchB26 = 66 // R_LARCH_B26 (b/bl, matches the Go linker's mapping)
) )
// ELFLOONG64Object returns the image as an ELF64 relocatable object file for // ELFLOONG64Object returns the image as an ELF64 relocatable object file for
@@ -95,14 +102,17 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name) return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name)
} }
typ := uint32(rLarchPCALAHI20) typ := uint32(rLarchPCALAHI20)
if r.Kind == RelLoong64AddrLo { switch r.Kind {
case RelLoong64AddrLo:
typ = rLarchPCALALO12 typ = rLarchPCALALO12
case RelLoong64Branch:
typ = rLarchB26
} }
relas = append(relas, elfRela{ relas = append(relas, elfRela{
off: uint64(fn.Offset + r.Off), off: uint64(fn.Offset + r.Off),
typ: typ, typ: typ,
sym: idx, sym: idx,
addend: r.Addend - int64(r.After-r.Off), addend: r.Addend,
}) })
} }
} }
@@ -183,9 +193,25 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
out = append(out, 0) out = append(out, 0)
} }
} }
dw := appendDWARFSections(&out, img, "gasm.s", symIdx, dwAlign) dw := appendDWARFSections(&out, img, dwarfSourceName(img), symIdx, dwAlign, cfiLOONG64)
dwarfStart := 0 // section index of .debug_abbrev, set when DWARF is present
if dw != nil { if dw != nil {
nSections += 4 // Five DWARF sections: .debug_abbrev, .debug_info, .debug_line,
// .debug_line_str and .debug_frame (the CIE is unconditional, so
// the frame section is always present), plus the relocation
// sections below when they carry entries.
dwarfStart = nSections
nSections += 5
appendDWARFRelas(&out, dw, rLarchAbs64, dwAlign)
if dw.infoRelaCount > 0 {
nSections++
}
if dw.lineRelaCount > 0 {
nSections++
}
if dw.frameRelaCount > 0 {
nSections++
}
} }
align(8) align(8)
@@ -214,13 +240,37 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24) putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
} }
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0) putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil { if dw != nil {
// secIdx is a running section index: each putSh below emits the
// next header, and the sh_info of a .rela section names the index
// of the section it relocates.
secIdx := dwarfStart
putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0) putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0)
secIdx++
putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0) putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0)
secInfoIdx := secIdx
secIdx++
if dw.infoRelaCount > 0 {
putSh(".rela.debug_info", shtRela, 0, dw.infoRelaOff, 24*dw.infoRelaCount, secSymtab, secInfoIdx, 8, 24)
secIdx++
}
putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0) putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0)
secLineIdx := secIdx
secIdx++
if dw.lineRelaCount > 0 {
putSh(".rela.debug_line", shtRela, 0, dw.lineRelaOff, 24*dw.lineRelaCount, secSymtab, secLineIdx, 8, 24)
secIdx++
}
putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0) putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0)
secIdx++
if dw.frameSize > 0 { if dw.frameSize > 0 {
putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0) putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0)
secFrameIdx := secIdx
secIdx++
if dw.frameRelaCount > 0 {
putSh(".rela.debug_frame", shtRela, 0, dw.frameRelaOff, 24*dw.frameRelaCount, secSymtab, secFrameIdx, 8, 24)
}
} }
} }
@@ -233,7 +283,7 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
le.PutUint64(hdr[24:], 0) le.PutUint64(hdr[24:], 0)
le.PutUint64(hdr[32:], 0) le.PutUint64(hdr[32:], 0)
le.PutUint64(hdr[40:], uint64(shoff)) le.PutUint64(hdr[40:], uint64(shoff))
le.PutUint32(hdr[48:], 0) le.PutUint32(hdr[48:], efLarchAbiDoubleObjV1)
le.PutUint16(hdr[52:], 64) le.PutUint16(hdr[52:], 64)
le.PutUint16(hdr[54:], 0) le.PutUint16(hdr[54:], 0)
le.PutUint16(hdr[56:], 0) le.PutUint16(hdr[56:], 0)
+47
View File
@@ -55,6 +55,11 @@ DATA answer<>+0(SB)/8, $42
if ef.Type != elf.ET_REL || ef.Machine != elf.EM_LOONGARCH { if ef.Type != elf.ET_REL || ef.Machine != elf.EM_LOONGARCH {
t.Errorf("type/machine = %v/%v, want ET_REL/EM_LOONGARCH", ef.Type, ef.Machine) t.Errorf("type/machine = %v/%v, want ET_REL/EM_LOONGARCH", ef.Type, ef.Machine)
} }
// The double-float ABI plus OBJABI_V1 flags the Go toolchain writes;
// system linkers refuse ABI-mismatched merges.
if flags := binary.LittleEndian.Uint32(obj[48:]); flags != efLarchAbiDoubleObjV1 {
t.Errorf("e_flags = %#x, want %#x (double-float, OBJABI_V1)", flags, efLarchAbiDoubleObjV1)
}
text := ef.Section(".text") text := ef.Section(".text")
data := ef.Section(".data") data := ef.Section(".data")
@@ -198,3 +203,45 @@ TEXT ·nop(SB), NOSPLIT, $0
t.Error("function symbol nop not found") t.Error("function symbol nop not found")
} }
} }
// TestELFLOONG64BranchRelocation checks that the morestack call and an
// internal CALL both carry R_LARCH_B26 in the emitted object, matching the
// Go linker's mapping of its call relocation.
func TestELFLOONG64BranchRelocation(t *testing.T) {
f, errs := parser.Parse("k_loong64.s", "TEXT \u00b7callbig(SB), $8192-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("AssembleFileLOONG64: %v", err)
}
obj, err := img.ELFLOONG64Object()
if err != nil {
t.Fatalf("ELFLOONG64Object: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaSec := ef.Section(".rela.text")
if relaSec == nil {
t.Fatal("missing .rela.text")
}
raw, err := relaSec.Data()
if err != nil {
t.Fatal(err)
}
// The guard's morestack call plus the body's CALL to other.
if len(raw)%24 != 0 || len(raw)/24 != 2 {
t.Fatalf(".rela.text has %d bytes, want two 24-byte entries", len(raw))
}
le := binary.LittleEndian
for i := range 2 {
info := le.Uint64(raw[i*24+8:])
if elf.R_LARCH(info&0xffffffff) != elf.R_LARCH_B26 {
t.Errorf("relocation %d type = %v, want R_LARCH_B26", i, elf.R_LARCH(info&0xffffffff))
}
}
}
+60 -11
View File
@@ -13,8 +13,13 @@ import (
const ( const (
emRISCV = 243 // EM_RISCV emRISCV = 243 // EM_RISCV
// EF_RISCV_FLOAT_ABI_DOUBLE: the double-precision float ABI the Go
// toolchain targets (cmd/link/internal/ld/elf.go writes Flags = 0x4 for
// RISCV64). System linkers refuse to merge ET_REL objects whose float
// ABI differs, so 0 (soft-float) would make the object unlinkable.
efRISCVFloatAbiDouble = 0x4
// RISC-V relocation types. // RISC-V relocation types.
rRISCV32 = 1
rRISCVJAL = 17 // R_RISCV_JAL rRISCVJAL = 17 // R_RISCV_JAL
rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20 rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20
rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I
@@ -83,9 +88,14 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
// Build relocations. Each SB reference is an AUIPC + second-instruction // Build relocations. Each SB reference is an AUIPC + second-instruction
// pair carrying a single relocation kind; the ELF writer expands it into // pair carrying a single relocation kind; the ELF writer expands it into
// the R_RISCV_PCREL_HI20 + R_RISCV_PCREL_LO12_I/S pair the psABI expects. // the R_RISCV_PCREL_HI20 + R_RISCV_PCREL_LO12_I/S pair the psABI expects.
// The HI20 carries the symbol addend; the LO12 addend is zero, matching // The HI20 carries the symbol and its addend. The LO12's symbol must
// cmd/link's own ELF conversion (the LO12 resolves against the HI20's // denote the AUIPC site the HI20 relocates (psABI §8.4.9: the pair is
// AUIPC location). // resolved against the label of the AUIPC, not the target symbol;
// cmd/link generates one local text symbol per AUIPC for exactly this,
// cmd/link/internal/riscv64/asm.go). The .text section symbol with the
// AUIPC's section-relative offset as addend gives S + A = the AUIPC
// address, which is that label.
const secSymText = 1 // syms[1], the .text section symbol
type elfRela struct { type elfRela struct {
off uint64 off uint64
typ uint32 typ uint32
@@ -99,21 +109,20 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
if !ok { if !ok {
return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name) return nil, fmt.Errorf("relocation references unknown symbol %q", r.Name)
} }
auipc := int64(fn.Offset + r.Off)
switch r.Kind { switch r.Kind {
case RelRISCVPCRELIType: case RelRISCVPCRELIType:
relas = append(relas, relas = append(relas,
elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVPCRELHI20, sym: idx, addend: r.Addend}, elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVPCRELHI20, sym: idx, addend: r.Addend},
elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rRISCVPCRELLO12I, sym: idx, addend: 0}, elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rRISCVPCRELLO12I, sym: secSymText, addend: auipc},
) )
case RelRISCVPCRELSType: case RelRISCVPCRELSType:
relas = append(relas, relas = append(relas,
elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVPCRELHI20, sym: idx, addend: r.Addend}, elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVPCRELHI20, sym: idx, addend: r.Addend},
elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rRISCVPCRELLO12S, sym: idx, addend: 0}, elfRela{off: uint64(fn.Offset + r.Off + 4), typ: rRISCVPCRELLO12S, sym: secSymText, addend: auipc},
) )
case RelRISCVJal: case RelRISCVJal:
relas = append(relas, elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVJAL, sym: idx, addend: r.Addend}) relas = append(relas, elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCVJAL, sym: idx, addend: r.Addend})
case RelPCRelAbs:
relas = append(relas, elfRela{off: uint64(fn.Offset + r.Off), typ: rRISCV32, sym: idx, addend: r.Addend})
default: default:
return nil, fmt.Errorf("relocation kind %v unsupported in ELF emission", r.Kind) return nil, fmt.Errorf("relocation kind %v unsupported in ELF emission", r.Kind)
} }
@@ -196,9 +205,25 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
out = append(out, 0) out = append(out, 0)
} }
} }
dw := appendDWARFSections(&out, img, "gasm.s", symIdx, dwAlign) dw := appendDWARFSections(&out, img, dwarfSourceName(img), symIdx, dwAlign, cfiRISCV64)
dwarfStart := 0 // section index of .debug_abbrev, set when DWARF is present
if dw != nil { if dw != nil {
nSections += 4 // Five DWARF sections: .debug_abbrev, .debug_info, .debug_line,
// .debug_line_str and .debug_frame (the CIE is unconditional, so
// the frame section is always present), plus the relocation
// sections below when they carry entries.
dwarfStart = nSections
nSections += 5
appendDWARFRelas(&out, dw, rRISCVAbs64, dwAlign)
if dw.infoRelaCount > 0 {
nSections++
}
if dw.lineRelaCount > 0 {
nSections++
}
if dw.frameRelaCount > 0 {
nSections++
}
} }
align(8) align(8)
@@ -227,13 +252,37 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24) putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
} }
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0) putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil { if dw != nil {
// secIdx is a running section index: each putSh below emits the
// next header, and the sh_info of a .rela section names the index
// of the section it relocates.
secIdx := dwarfStart
putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0) putSh(".debug_abbrev", shtProgbits, 0, dw.abbrevOff, dw.abbrevSize, 0, 0, 1, 0)
secIdx++
putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0) putSh(".debug_info", shtProgbits, 0, dw.infoOff, dw.infoSize, 0, 0, 1, 0)
secInfoIdx := secIdx
secIdx++
if dw.infoRelaCount > 0 {
putSh(".rela.debug_info", shtRela, 0, dw.infoRelaOff, 24*dw.infoRelaCount, secSymtab, secInfoIdx, 8, 24)
secIdx++
}
putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0) putSh(".debug_line", shtProgbits, 0, dw.lineOff, dw.lineSize, 0, 0, 1, 0)
secLineIdx := secIdx
secIdx++
if dw.lineRelaCount > 0 {
putSh(".rela.debug_line", shtRela, 0, dw.lineRelaOff, 24*dw.lineRelaCount, secSymtab, secLineIdx, 8, 24)
secIdx++
}
putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0) putSh(".debug_line_str", shtProgbits, 0, dw.lineStrOff, dw.lineStrSize, 0, 0, 1, 0)
secIdx++
if dw.frameSize > 0 { if dw.frameSize > 0 {
putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0) putSh(".debug_frame", shtProgbits, 0, dw.frameOff, dw.frameSize, 0, 0, 8, 0)
secFrameIdx := secIdx
secIdx++
if dw.frameRelaCount > 0 {
putSh(".rela.debug_frame", shtRela, 0, dw.frameRelaOff, 24*dw.frameRelaCount, secSymtab, secFrameIdx, 8, 24)
}
} }
} }
@@ -246,7 +295,7 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
le.PutUint64(hdr[24:], 0) le.PutUint64(hdr[24:], 0)
le.PutUint64(hdr[32:], 0) le.PutUint64(hdr[32:], 0)
le.PutUint64(hdr[40:], uint64(shoff)) le.PutUint64(hdr[40:], uint64(shoff))
le.PutUint32(hdr[48:], 0) le.PutUint32(hdr[48:], efRISCVFloatAbiDouble)
le.PutUint16(hdr[52:], 64) le.PutUint16(hdr[52:], 64)
le.PutUint16(hdr[54:], 0) le.PutUint16(hdr[54:], 0)
le.PutUint16(hdr[56:], 0) le.PutUint16(hdr[56:], 0)
+15 -3
View File
@@ -36,10 +36,15 @@ func Encodable(mnemonic string) bool {
} }
// CMOV carries size then condition (CMOVLGT); SET carries the condition // CMOV carries size then condition (CMOVLGT); SET carries the condition
// alone (SETNE). // alone (SETNE). The size letter is checked exactly as encodeCmov does,
// so a spelling like CMOVBGT is not reported encodable when Encode
// would reject it.
if rest, ok := strings.CutPrefix(upper, "CMOV"); ok && len(rest) >= 2 { if rest, ok := strings.CutPrefix(upper, "CMOV"); ok && len(rest) >= 2 {
if _, ok := jccMap[rest[1:]]; ok { switch rest[0] {
return true case 'W', 'L', 'Q':
if _, ok := jccMap[rest[1:]]; ok {
return true
}
} }
} }
if rest, ok := strings.CutPrefix(upper, "SET"); ok { if rest, ok := strings.CutPrefix(upper, "SET"); ok {
@@ -81,9 +86,16 @@ func Encodable(mnemonic string) bool {
"BSWAP", "BSWAP",
"PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2", "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2",
"MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX", "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX",
"MOVBWZX", "MOVBWSX", "MOVBLSX", "MOVBQSX", "MOVWQSX", "MOVLQZX",
"CVTSL2SD", "CVTSQ2SD", "CVTSL2SD", "CVTSQ2SD",
"MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS": "MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
return true return true
} }
// Full-name dispatches the size split would eat (a trailing width
// letter that is part of the mnemonic).
switch upper {
case "PMOVMSKB":
return true
}
return false return false
} }
+30 -8
View File
@@ -40,10 +40,20 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeRet() return e.encodeRet()
case upper == "NOP": case upper == "NOP":
return e.emit(&instr{opcode: []byte{0x90}, modrm: -1, sib: -1}) return e.emit(&instr{opcode: []byte{0x90}, modrm: -1, sib: -1})
case upper == "CALL": case upper == "CALL" || upper == "JMP":
return e.encodeJmpRel(ops, []byte{0xE8}) // Through a register or memory: FF /2 (CALL) or FF /4 (JMP).
case upper == "JMP": // Anything else is a rel32 against a label resolved by the assembler.
return e.encodeJmpRel(ops, []byte{0xE9}) if len(ops) == 1 {
switch ops[0].(type) {
case Reg, Mem:
return e.encodeIndirectBranch(upper, ops)
}
}
opcode := []byte{0xE8}
if upper == "JMP" {
opcode = []byte{0xE9}
}
return e.encodeJmpRel(ops, opcode)
} }
if cc, ok := condCode(upper); ok { if cc, ok := condCode(upper); ok {
return e.encodeJcc(cc, ops) return e.encodeJcc(cc, ops)
@@ -91,6 +101,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
if m, ok := sseBinTable[base]; ok { if m, ok := sseBinTable[base]; ok {
return e.encodeSSEBin(m, ops) return e.encodeSSEBin(m, ops)
} }
// PMOVMSKB ends in a width letter the size split would eat, so it
// dispatches on the full name like the packed binaries above.
if upper == "PMOVMSKB" {
return e.encodePmovmskb(upper, ops)
}
switch base { switch base {
case "MOV": case "MOV":
return e.encodeMov(ops, size) return e.encodeMov(ops, size)
@@ -107,16 +122,17 @@ func (e *enc) encode(mnem string, ops []Operand) error {
case "IMUL", "IMUL3": case "IMUL", "IMUL3":
return e.encodeImul(ops, size) return e.encodeImul(ops, size)
case "PUSH": case "PUSH":
return e.encodePushPop(ops, true) return e.encodePushPop(ops, size, true)
case "POP": case "POP":
return e.encodePushPop(ops, false) return e.encodePushPop(ops, size, false)
case "BSF", "BSR", "LZCNT", "TZCNT", "POPCNT": case "BSF", "BSR", "LZCNT", "TZCNT", "POPCNT":
return e.encodeCount(base, ops, size) return e.encodeCount(base, ops, size)
case "BSWAP": case "BSWAP":
return e.encodeBswap(ops, size) return e.encodeBswap(ops, size)
case "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2": case "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2":
return e.encodePrefetch(base, ops) return e.encodePrefetch(base, ops)
case "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX": case "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX",
"MOVBWZX", "MOVBWSX", "MOVBLSX", "MOVBQSX", "MOVWQSX", "MOVLQZX":
return e.encodeMovExtend(base, ops) return e.encodeMovExtend(base, ops)
case "CVTSL2SD", "CVTSQ2SD": case "CVTSL2SD", "CVTSQ2SD":
return e.encodeCvtsi2sd(base == "CVTSQ2SD", ops) return e.encodeCvtsi2sd(base == "CVTSQ2SD", ops)
@@ -284,7 +300,7 @@ func setRM(i *instr, reg Reg, rm Operand, opSize int) error {
} }
// setRMDigit fills in the ModR/M for an instruction whose reg field is an // setRMDigit fills in the ModR/M for an instruction whose reg field is an
// opcode /digit extension (0–7), which carries none of the register REX rules. // opcode /digit extension (0-7), which carries none of the register REX rules.
func setRMDigit(i *instr, digit int, rm Operand, opSize int) error { func setRMDigit(i *instr, digit int, rm Operand, opSize int) error {
return setRMReg(i, digit, false, false, rm, opSize) return setRMReg(i, digit, false, false, rm, opSize)
} }
@@ -335,6 +351,12 @@ func setMem(i *instr, regField int, m Mem) error {
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths. // a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
func memComponents(regField int, m Mem) (modrm, sib int, disp []byte, xBit, bBit int, err error) { func memComponents(regField int, m Mem) (modrm, sib int, disp []byte, xBit, bBit int, err error) {
sib = -1 sib = -1
// A displacement wider than int32 fits no encoding form; truncating it
// would address a different location, and go tool asm reports "offset
// too large" for the same operand.
if m.Disp < -(1<<31) || m.Disp > (1<<31)-1 {
return 0, -1, nil, 0, 0, fmt.Errorf("displacement %d does not fit in 32 bits", m.Disp)
}
// RIP-relative: neither base nor index. // RIP-relative: neither base nor index.
if !m.HasBase && !m.HasIndex { if !m.HasBase && !m.HasIndex {
return regField<<3 | 0x05, -1, le32(m.Disp), 0, 0, nil // mod=00, rm=101 return regField<<3 | 0x05, -1, le32(m.Disp), 0, 0, nil // mod=00, rm=101
+184 -2
View File
@@ -149,6 +149,45 @@ func TestPushPop(t *testing.T) {
checkSyntax(t, "push rbx", "PUSHQ", BX) checkSyntax(t, "push rbx", "PUSHQ", BX)
checkSyntax(t, "pop r12", "POPQ", Reg{idx: 12, size: 8}) checkSyntax(t, "pop r12", "POPQ", Reg{idx: 12, size: 8})
checkSyntax(t, "push 0x5", "PUSHQ", Imm(5)) checkSyntax(t, "push 0x5", "PUSHQ", Imm(5))
// The W spelling carries the 0x66 operand-size prefix, byte for byte
// with go tool asm; the L and B spellings are illegal in 64-bit mode
// there and rejected here rather than silently widened.
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"PUSHW AX", "PUSHW", []Operand{AX}, "6650"},
{"POPW AX", "POPW", []Operand{AX}, "6658"},
{"PUSHW $5", "PUSHW", []Operand{Imm(5)}, "666a05"},
{"PUSHW (AX)", "PUSHW", []Operand{Ptr(AX, 0, 2)}, "66ff30"},
{"PUSHQ AX", "PUSHQ", []Operand{AX}, "50"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s: bytes %s, want %s", c.name, got, c.want)
}
}
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"PUSHL AX", "PUSHL", []Operand{AX}},
{"PUSHL R8", "PUSHL", []Operand{Reg{idx: 8, size: 8}}},
{"POPL BX", "POPL", []Operand{BX}},
{"PUSHB AX", "PUSHB", []Operand{AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
} }
func TestUnary(t *testing.T) { func TestUnary(t *testing.T) {
@@ -181,6 +220,39 @@ func TestControl(t *testing.T) {
checkOp(t, x86asm.JBE, "JLS", Imm(0)) checkOp(t, x86asm.JBE, "JLS", Imm(0))
} }
// TestIndirectControlFlow pins the indirect JMP/CALL forms: FF /4 for JMP and
// FF /2 for CALL through a register or memory. A REX appears only for the
// extended registers, never REX.W: the branch operand size is fixed at 64
// bits in long mode.
func TestIndirectControlFlow(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"JMP AX", "JMP", []Operand{AX}, "ffe0"},
{"CALL AX", "CALL", []Operand{AX}, "ffd0"},
{"JMP (BX)", "JMP", []Operand{Ptr(BX, 0, 8)}, "ff23"},
{"CALL (BX)", "CALL", []Operand{Ptr(BX, 0, 8)}, "ff13"},
{"JMP 8(BX)", "JMP", []Operand{Ptr(BX, 8, 8)}, "ff6308"},
{"CALL -16(BX)", "CALL", []Operand{Ptr(BX, -16, 8)}, "ff53f0"},
{"JMP R8", "JMP", []Operand{Reg{idx: 8, size: 2}}, "41ffe0"},
{"CALL R9", "CALL", []Operand{Reg{idx: 9, size: 2}}, "41ffd1"},
{"JMP R15", "JMP", []Operand{Reg{idx: 15, size: 2}}, "41ffe7"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestSSEMoveGroundTruth checks the legacy (non-VEX) SSE moves byte for byte // TestSSEMoveGroundTruth checks the legacy (non-VEX) SSE moves byte for byte
// against the Go assembler. wantOp is the decoder's name, which differs from // against the Go assembler. wantOp is the decoder's name, which differs from
// the Plan 9 spelling for the octa moves (MOVOU = MOVDQU, MOVO = MOVDQA). // the Plan 9 spelling for the octa moves (MOVOU = MOVDQU, MOVO = MOVDQA).
@@ -230,7 +302,7 @@ func TestSSEMoveGroundTruth(t *testing.T) {
// TestGoFlacScalarTail encodes the scalar tail of an analyze kernel to confirm // TestGoFlacScalarTail encodes the scalar tail of an analyze kernel to confirm
// the encoder handles a realistic instruction sequence. // the encoder handles a realistic instruction sequence.
func TestGoFlacScalarTail(t *testing.T) { func TestGoFlacScalarTail(t *testing.T) {
// MOVQ swin_base+0(FP), SI — modelled as MOVQ disp(reg), reg. // MOVQ swin_base+0(FP), SI; modelled as MOVQ disp(reg), reg.
checkSyntax(t, "mov rsi, qword ptr [rax+0x10]", "MOVQ", Ptr(AX, 0x10, 8), SI) checkSyntax(t, "mov rsi, qword ptr [rax+0x10]", "MOVQ", Ptr(AX, 0x10, 8), SI)
checkSyntax(t, "lea r9, ptr [rsi+4*rbx]", "LEAQ", Idx(SI, BX, 4, 0, 8), Reg{idx: 9, size: 8}) checkSyntax(t, "lea r9, ptr [rsi+4*rbx]", "LEAQ", Idx(SI, BX, 4, 0, 8), Reg{idx: 9, size: 8})
checkSyntax(t, "and r10, -0x8", "ANDQ", Imm(-8), Reg{idx: 10, size: 8}) checkSyntax(t, "and r10, -0x8", "ANDQ", Imm(-8), Reg{idx: 10, size: 8})
@@ -285,6 +357,18 @@ func TestScalarGroundTruth(t *testing.T) {
{"MOVBQZX AL,R8", "MOVBQZX", []Operand{AL, r8}, "4c0fb6c0", "MOVZX"}, {"MOVBQZX AL,R8", "MOVBQZX", []Operand{AL, r8}, "4c0fb6c0", "MOVZX"},
{"MOVWLZX AX,CX", "MOVWLZX", []Operand{AX, CX}, "0fb7c8", "MOVZX"}, {"MOVWLZX AX,CX", "MOVWLZX", []Operand{AX, CX}, "0fb7c8", "MOVZX"},
{"MOVWQZX AX,R8", "MOVWQZX", []Operand{AX, r8}, "4c0fb7c0", "MOVZX"}, {"MOVWQZX AX,R8", "MOVWQZX", []Operand{AX, r8}, "4c0fb7c0", "MOVZX"},
// The width pairs the toolchain accepts and GOROOT uses; bytes
// pinned from go tool asm (see testdata/verify/widen_amd64.s).
{"MOVBWZX (BX),R11W", "MOVBWZX", []Operand{Ptr(BX, 0, 1), Reg{idx: 11, size: 2}}, "66440fb61b", "MOVZX"},
{"MOVBWSX (BX),R11W", "MOVBWSX", []Operand{Ptr(BX, 0, 1), Reg{idx: 11, size: 2}}, "66440fbe1b", "MOVSX"},
{"MOVBLSX (BX),AX", "MOVBLSX", []Operand{Ptr(BX, 0, 1), AX}, "0fbe03", "MOVSX"},
{"MOVBQSX (BX),R8", "MOVBQSX", []Operand{Ptr(BX, 0, 1), r8}, "4c0fbe03", "MOVSX"},
{"MOVWQSX (BX),R9", "MOVWQSX", []Operand{Ptr(BX, 0, 2), r9}, "4c0fbf0b", "MOVSX"},
// A long to quad zero-extend is a plain 32-bit move.
{"MOVLQZX (BX),DX", "MOVLQZX", []Operand{Ptr(BX, 0, 4), DX}, "8b13", "MOV"},
{"MOVLQZX AX,DX", "MOVLQZX", []Operand{AX, DX}, "8bd0", "MOV"},
{"PMOVMSKB X1,AX", "PMOVMSKB", []Operand{vreg(t, "X1"), AX}, "660fd7c1", "PMOVMSKB"},
{"PMOVMSKB X11,CX", "PMOVMSKB", []Operand{vreg(t, "X11"), CX}, "66410fd7cb", "PMOVMSKB"},
{"CVTSL2SD R8,X13", "CVTSL2SD", []Operand{r8, vreg(t, "X13")}, "f2450f2ae8", "CVTSI2SD"}, {"CVTSL2SD R8,X13", "CVTSL2SD", []Operand{r8, vreg(t, "X13")}, "f2450f2ae8", "CVTSI2SD"},
{"CVTSL2SD AX,X0", "CVTSL2SD", []Operand{AX, vreg(t, "X0")}, "f20f2ac0", "CVTSI2SD"}, {"CVTSL2SD AX,X0", "CVTSL2SD", []Operand{AX, vreg(t, "X0")}, "f20f2ac0", "CVTSI2SD"},
{"CVTSQ2SD R8,X13", "CVTSQ2SD", []Operand{r8, vreg(t, "X13")}, "f24d0f2ae8", "CVTSI2SD"}, {"CVTSQ2SD R8,X13", "CVTSQ2SD", []Operand{r8, vreg(t, "X13")}, "f24d0f2ae8", "CVTSI2SD"},
@@ -364,6 +448,104 @@ func TestScalarErrors(t *testing.T) {
} }
} }
// TestImmediateOutOfRange pins the go-tool-asm parity of the immediate and
// displacement spans: a scalar immediate must fit a signed or unsigned 32-bit
// word (only MOVQ reg, $imm takes the full int64), a scalar shift count must
// be an unsigned byte, and a displacement must fit int32. Every rejected
// shape here is rejected by `go tool asm` too; every accepted one encodes the
// same bytes.
func TestImmediateOutOfRange(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
}{
{"SHLQ count 300", "SHLQ", []Operand{Imm(300), AX}},
{"SHLQ count -1", "SHLQ", []Operand{Imm(-1), AX}},
{"SHLW count 256", "SHLW", []Operand{Imm(256), DX}},
{"SHLB count 300", "SHLB", []Operand{Imm(300), BL}},
{"MOVL imm32+", "MOVL", []Operand{Imm(4294967296), AX}},
{"MOVL imm32-", "MOVL", []Operand{Imm(-2147483649), AX}},
{"MOVW imm32+", "MOVW", []Operand{Imm(4294967296), AX}},
{"MOVB imm32+", "MOVB", []Operand{Imm(4294967296), AL}},
{"ADDB imm32+", "ADDB", []Operand{Imm(4294967296), AL}},
{"ADDL imm32+", "ADDL", []Operand{Imm(4294967296), AX}},
{"ADDQ imm32+", "ADDQ", []Operand{Imm(8589934592), AX}},
{"CMPQ imm32+", "CMPQ", []Operand{AX, Imm(4294967296)}},
{"CMPQ imm32-", "CMPQ", []Operand{AX, Imm(-2147483649)}},
{"TESTL imm32+", "TESTL", []Operand{Imm(4294967296), AX}},
{"IMUL3L imm32+", "IMUL3L", []Operand{Imm(4294967296), CX, DX}},
{"PUSHQ imm32+", "PUSHQ", []Operand{Imm(4294967296)}},
{"MOVQ mem imm32+", "MOVQ", []Operand{Imm(4294967296), Ptr(AX, 0, 8)}},
{"disp32+", "MOVQ", []Operand{Ptr(AX, 4294967296, 8), BX}},
{"disp32+ max", "MOVQ", []Operand{Ptr(AX, 2147483648, 8), BX}},
{"disp32-", "MOVQ", []Operand{Ptr(AX, -2147483649, 8), BX}},
{"VEX disp32+", "VMOVDQU", []Operand{Ptr(AX, 4294967296, 32), vreg(t, "Y1")}},
{"EVEX disp32+", "VMOVDQU32", []Operand{Ptr(AX, 4294967296, 64), vreg(t, "Z1")}},
}
for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
// TestImmediateTruncation pins the toolchain-matching truncations inside the
// accepted 32-bit span: the narrower fields take the low bits silently, byte
// for byte with `go tool asm` (which rejects none of these).
func TestImmediateTruncation(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"ADDB $256,BL", "ADDB", []Operand{Imm(256), BL}, "80c300"},
{"ADDB $1000,BL", "ADDB", []Operand{Imm(1000), BL}, "80c3e8"},
{"MOVB $256,AL", "MOVB", []Operand{Imm(256), AL}, "b000"},
{"MOVB $-129,AL", "MOVB", []Operand{Imm(-129), AL}, "b07f"},
{"MOVW $65536,AX", "MOVW", []Operand{Imm(65536), AX}, "66b80000"},
{"MOVW $65535,AX", "MOVW", []Operand{Imm(65535), AX}, "66b8ffff"},
{"MOVW $-32769,AX", "MOVW", []Operand{Imm(-32769), AX}, "66b8ff7f"},
{"MOVL $4294967295,AX", "MOVL", []Operand{Imm(4294967295), AX}, "b8ffffffff"},
{"ADDQ $4294967295,AX", "ADDQ", []Operand{Imm(4294967295), AX}, "4805ffffffff"},
{"CMPB BL,$255", "CMPB", []Operand{BL, Imm(255)}, "80fbff"},
{"CMPQ AX,$4294967295", "CMPQ", []Operand{AX, Imm(4294967295)}, "483dffffffff"},
{"MOVQ $4294967295,0(AX)", "MOVQ", []Operand{Imm(4294967295), Ptr(AX, 0, 8)}, "48c700ffffffff"},
{"SHLQ $255,AX", "SHLQ", []Operand{Imm(255), AX}, "48c1e0ff"},
{"SHLQ $0,AX", "SHLQ", []Operand{Imm(0), AX}, "48c1e000"},
// The one form beyond the 32-bit span: the imm64 MOVQ register move.
{"MOVQ $4294967296,AX", "MOVQ", []Operand{Imm(4294967296), AX}, "48b80000000001000000"},
{"MOVQ disp32 max", "MOVQ", []Operand{Ptr(AX, 2147483647, 8), BX}, "488b98ffffff7f"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s: bytes %s, want %s", c.name, got, c.want)
}
}
}
// TestEncodableCmovSize pins the linter contract for CMOVcc: Encodable must
// reject the spellings Encode rejects, so a mnemonic like CMOVBGT (no size
// letter) is not reported as encodable.
func TestEncodableCmovSize(t *testing.T) {
for _, m := range []string{"CMOVBGT", "CMOVXEQ", "CMOVB", "CMOV", "CMOVWXX"} {
if Encodable(m) {
t.Errorf("Encodable(%q) = true, want false", m)
}
}
for _, m := range []string{"CMOVLGT", "CMOVQGT", "CMOVWLS", "CMOVLEQ"} {
if !Encodable(m) {
t.Errorf("Encodable(%q) = false, want true", m)
}
}
}
// TestSSEBinGroundTruth checks the legacy packed/scalar binary family // TestSSEBinGroundTruth checks the legacy packed/scalar binary family
// byte for byte (no prefix / 66 / F2 / F3 variants). // byte for byte (no prefix / 66 / F2 / F3 variants).
func TestSSEBinGroundTruth(t *testing.T) { func TestSSEBinGroundTruth(t *testing.T) {
@@ -421,7 +603,7 @@ func TestSSEShuffleGroundTruth(t *testing.T) {
} }
// TestMOVQXMMGroundTruth pins the SSE2 packed-quadword move encodings: // TestMOVQXMMGroundTruth pins the SSE2 packed-quadword move encodings:
// loads and register moves on F3 0F 7E, stores on 66 0F D6 — the forms // loads and register moves on F3 0F 7E, stores on 66 0F D6; the forms
// the GPR-move fallback silently corrupted. // the GPR-move fallback silently corrupted.
func TestMOVQXMMGroundTruth(t *testing.T) { func TestMOVQXMMGroundTruth(t *testing.T) {
cases := []struct { cases := []struct {
+119 -102
View File
@@ -9,10 +9,10 @@ import (
) )
// This file implements EVEX (AVX-512) instruction encoding: the four-byte // This file implements EVEX (AVX-512) instruction encoding: the four-byte
// EVEX prefix with 5-bit vector register fields (Z0–Z31, X/Y 16–31), the // EVEX prefix with 5-bit vector register fields (Z0-Z31, X/Y 16-31), the
// compressed disp8×N displacement, and the operand shapes the go-flac // compressed disp8×N displacement, and the operand shapes the go-flac
// AVX-512 kernels use plus the common floating-point and conversion set. // AVX-512 kernels use plus the common floating-point and conversion set.
// Masking follows the Go assembler's spelling: an explicit K1–K7 operand // Masking follows the Go assembler's spelling: an explicit K1-K7 operand
// anywhere among the operands (merging) plus a ".Z" mnemonic suffix for // anywhere among the operands (merging) plus a ".Z" mnemonic suffix for
// zeroing. K-register operands (mask destinations, KMOVW, KTESTW) are // zeroing. K-register operands (mask destinations, KMOVW, KTESTW) are
// supported too. // supported too.
@@ -36,7 +36,7 @@ type evexSpec struct {
// are taken from the Go assembler's opcode tables, which are authoritative // are taken from the Go assembler's opcode tables, which are authoritative
// for byte-for-byte agreement. // for byte-for-byte agreement.
var evexTable = map[string]evexSpec{ var evexTable = map[string]evexSpec{
// EVEX.128/256/512.66.0F — integer arithmetic / logic, NDS form. // EVEX.128/256/512.66.0F, integer arithmetic / logic, NDS form.
"VPADDD": {1, 0xFE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPADDD": {1, 0xFE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDQ": {1, 0xD4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPADDQ": {1, 0xD4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBD": {1, 0xFA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPSUBD": {1, 0xFA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -48,25 +48,25 @@ var evexTable = map[string]evexSpec{
"VPCMPEQD": {1, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPCMPEQD": {1, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD231PD": {2, 0xB8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VFMADD231PD": {2, 0xB8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1 — packed double arithmetic. // EVEX.128/256/512.66.0F.W1, packed double arithmetic.
"VADDPD": {1, 0x58, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VADDPD": {1, 0x58, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMULPD": {1, 0x59, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VMULPD": {1, 0x59, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VSUBPD": {1, 0x5C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VSUBPD": {1, 0x5C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VDIVPD": {1, 0x5E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VDIVPD": {1, 0x5E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMINPD": {1, 0x5D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VMINPD": {1, 0x5D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMAXPD": {1, 0x5F, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VMAXPD": {1, 0x5F, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.W0 — packed single arithmetic. // EVEX.128/256/512.0F.W0, packed single arithmetic.
"VADDPS": {1, 0x58, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VADDPS": {1, 0x58, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VMULPS": {1, 0x59, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VMULPS": {1, 0x59, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VSUBPS": {1, 0x5C, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VSUBPS": {1, 0x5C, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VDIVPS": {1, 0x5E, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VDIVPS": {1, 0x5E, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VMINPS": {1, 0x5D, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VMINPS": {1, 0x5D, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VMAXPS": {1, 0x5F, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}}, "VMAXPS": {1, 0x5F, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1 — packed double unpack. // EVEX.128/256/512.66.0F.W1, packed double unpack.
"VUNPCKLPD": {1, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VUNPCKLPD": {1, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKHPD": {1, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VUNPCKHPD": {1, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128.F2.0F.W1 — scalar double arithmetic (the packed opcodes with // EVEX.128.F2.0F.W1, scalar double arithmetic (the packed opcodes with
// an F2 pp; the EVEX forms exist for masked and zeroing use). The // an F2 pp; the EVEX forms exist for masked and zeroing use). The
// memory operand is a single double, so disp8×N = 8. // memory operand is a single double, so disp8×N = 8.
"VADDSD": {1, 0x58, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}}, "VADDSD": {1, 0x58, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}},
@@ -76,7 +76,7 @@ var evexTable = map[string]evexSpec{
"VMINSD": {1, 0x5D, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}}, "VMINSD": {1, 0x5D, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}},
"VMAXSD": {1, 0x5F, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}}, "VMAXSD": {1, 0x5F, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}},
// EVEX.128.F3.0F.W0 — scalar single arithmetic (disp8×N = 4). // EVEX.128.F3.0F.W0, scalar single arithmetic (disp8×N = 4).
"VADDSS": {1, 0x58, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VADDSS": {1, 0x58, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
"VSUBSS": {1, 0x5C, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VSUBSS": {1, 0x5C, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
"VMULSS": {1, 0x59, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VMULSS": {1, 0x59, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
@@ -84,38 +84,38 @@ var evexTable = map[string]evexSpec{
"VMINSS": {1, 0x5D, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VMINSS": {1, 0x5D, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
"VMAXSS": {1, 0x5F, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VMAXSS": {1, 0x5F, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
// EVEX.512.66.0F3A — align (NDS + imm8). // EVEX.512.66.0F3A, align (NDS + imm8).
"VALIGND": {3, 0x03, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VALIGND": {3, 0x03, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F — immediate shift (VPSRAD /4). // EVEX.128/256/512.66.0F, immediate shift (VPSRAD /4).
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}}, "VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1 — variable shift with an XMM count (VPSRAQ; // EVEX.128/256/512.66.0F.W1, variable shift with an XMM count (VPSRAQ;
// the W bit distinguishes it from VPSRAD's E2 form). // the W bit distinguishes it from VPSRAD's E2 form).
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.F3.0F.W1 — signed qword to packed double (reg=dst, // EVEX.128/256/512.F3.0F.W1, signed qword to packed double (reg=dst,
// rm=src, no vvvv). // rm=src, no vvvv).
"VCVTQQ2PD": {1, 0xE6, 1, 2, -1, vexRM, [3]int{16, 32, 64}}, "VCVTQQ2PD": {1, 0xE6, 1, 2, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.128/256/512.F2.0F.W1 — duplicate the low double (reg=dst, // EVEX.128/256/512.F2.0F.W1, duplicate the low double (reg=dst,
// rm=src, no vvvv): a 128-bit destination reads a single double from // rm=src, no vvvv): a 128-bit destination reads a single double from
// memory (disp8×8), the wider ones read the full operand. // memory (disp8×8), the wider ones read the full operand.
"VMOVDDUP": {1, 0x12, 1, 3, -1, vexRM, [3]int{8, 32, 64}}, "VMOVDDUP": {1, 0x12, 1, 3, -1, vexRM, [3]int{8, 32, 64}},
// EVEX.128/256/512.0F.W0 — signed dword to packed single (reg=dst, // EVEX.128/256/512.0F.W0, signed dword to packed single (reg=dst,
// rm=src, no vvvv, no mandatory prefix — as in the VEX form). // rm=src, no vvvv, no mandatory prefix, as in the VEX form).
"VCVTDQ2PS": {1, 0x5B, 0, 0, -1, vexRM, [3]int{16, 32, 64}}, "VCVTDQ2PS": {1, 0x5B, 0, 0, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.W0 — packed single to packed double: the // EVEX.128/256/512.0F.W0, packed single to packed double: the
// destination is twice the source width and sets the length; disp8×N // destination is twice the source width and sets the length; disp8×N
// follows the narrow memory source. No F3 prefix: the Go assembler // follows the narrow memory source. No F3 prefix: the Go assembler
// emits this instruction with pp = 00 (Intel's maps would call that // emits this instruction with pp = 00 (Intel's maps would call that
// undefined) and gasm reproduces the Go assembler's bytes — its machine // undefined) and gasm reproduces the Go assembler's bytes, its machine
// code is the oracle, not the manual. // code is the oracle, not the manual.
"VCVTPS2PD": {1, 0x5A, 0, 0, -1, vexRM, [3]int{8, 16, 32}}, "VCVTPS2PD": {1, 0x5A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.128/256/512.F3.0F.W0 — signed dword to packed double (the EVEX // EVEX.128/256/512.F3.0F.W0, signed dword to packed double (the EVEX
// form of the VEX instruction; the destination sets the length, disp8×N // form of the VEX instruction; the destination sets the length, disp8×N
// follows the narrow memory source). // follows the narrow memory source).
"VCVTDQ2PD": {1, 0xE6, 0, 2, -1, vexRM, [3]int{8, 16, 32}}, "VCVTDQ2PD": {1, 0xE6, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
// EVEX packed double → dword conversions: the source is the wide // EVEX packed double → dword conversions: the source is the wide
// operand and the mnemonic fixes the length — the bare names are // operand and the mnemonic fixes the length, the bare names are
// 512-bit only (ZMM source, XMM destination), the X/Y spellings are // 512-bit only (ZMM source, XMM destination), the X/Y spellings are
// EVEX-128/256. Exactly one slot of n is valid; it names the vector // EVEX-128/256. Exactly one slot of n is valid; it names the vector
// length (and the disp8×N multiplier) a register or memory source // length (and the disp8×N multiplier) a register or memory source
@@ -127,7 +127,7 @@ var evexTable = map[string]evexSpec{
"VCVTTPD2DQX": {1, 0xE6, 1, 1, -1, vexRMSrcLen, [3]int{16, 0, 0}}, "VCVTTPD2DQX": {1, 0xE6, 1, 1, -1, vexRMSrcLen, [3]int{16, 0, 0}},
"VCVTTPD2DQY": {1, 0xE6, 1, 1, -1, vexRMSrcLen, [3]int{0, 32, 0}}, "VCVTTPD2DQY": {1, 0xE6, 1, 1, -1, vexRMSrcLen, [3]int{0, 32, 0}},
// EVEX.66.0F3A — ternary logic and lane shuffles (NDS + imm8). // EVEX.66.0F3A, ternary logic and lane shuffles (NDS + imm8).
"VPTERNLOGD": {3, 0x25, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VPTERNLOGD": {3, 0x25, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPTERNLOGQ": {3, 0x25, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VPTERNLOGQ": {3, 0x25, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFI32X4": {3, 0x43, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VSHUFI32X4": {3, 0x43, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
@@ -136,11 +136,11 @@ var evexTable = map[string]evexSpec{
"VSHUFF64X2": {3, 0x23, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VSHUFF64X2": {3, 0x23, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPALIGNR": {3, 0x0F, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VPALIGNR": {3, 0x0F, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F — the EVEX forms of the VEX two-source shuffle. // EVEX.66.0F, the EVEX forms of the VEX two-source shuffle.
"VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F3A — lane insert ($imm, xsrc, zsrc1, zdst). // EVEX.66.0F3A, lane insert ($imm, xsrc, zsrc1, zdst).
"VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}}, "VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
"VINSERTF32X8": {3, 0x1A, 0, 1, -1, vexNDS3Imm, [3]int{0, 0, 32}}, "VINSERTF32X8": {3, 0x1A, 0, 1, -1, vexNDS3Imm, [3]int{0, 0, 32}},
"VINSERTF64X2": {3, 0x18, 1, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}}, "VINSERTF64X2": {3, 0x18, 1, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
@@ -150,7 +150,7 @@ var evexTable = map[string]evexSpec{
"VINSERTI64X2": {3, 0x38, 1, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}}, "VINSERTI64X2": {3, 0x38, 1, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
"VINSERTI64X4": {3, 0x3A, 1, 1, -1, vexNDS3Imm, [3]int{0, 0, 32}}, "VINSERTI64X4": {3, 0x3A, 1, 1, -1, vexNDS3Imm, [3]int{0, 0, 32}},
// EVEX.66.0F3A — lane extract (reg=source, rm=XMM/YMM destination, // EVEX.66.0F3A, lane extract (reg=source, rm=XMM/YMM destination,
// imm8). // imm8).
"VEXTRACTF32X4": {3, 0x19, 0, 1, -1, vexExtract, [3]int{0, 16, 16}}, "VEXTRACTF32X4": {3, 0x19, 0, 1, -1, vexExtract, [3]int{0, 16, 16}},
"VEXTRACTF32X8": {3, 0x1B, 0, 1, -1, vexExtract, [3]int{0, 0, 32}}, "VEXTRACTF32X8": {3, 0x1B, 0, 1, -1, vexExtract, [3]int{0, 0, 32}},
@@ -159,14 +159,14 @@ var evexTable = map[string]evexSpec{
"VEXTRACTI32X8": {3, 0x3B, 0, 1, -1, vexExtract, [3]int{0, 0, 32}}, "VEXTRACTI32X8": {3, 0x3B, 0, 1, -1, vexExtract, [3]int{0, 0, 32}},
"VEXTRACTI64X2": {3, 0x39, 1, 1, -1, vexExtract, [3]int{0, 16, 16}}, "VEXTRACTI64X2": {3, 0x39, 1, 1, -1, vexExtract, [3]int{0, 16, 16}},
// EVEX.66.0F — compare with an opmask destination ($imm, src2, src1, // EVEX.66.0F, compare with an opmask destination ($imm, src2, src1,
// kdst): NDS3Imm with the K register in the reg field. // kdst): NDS3Imm with the K register in the reg field.
"VCMPPD": {1, 0xC2, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VCMPPD": {1, 0xC2, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VCMPPS": {1, 0xC2, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VCMPPS": {1, 0xC2, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VCMPSD": {1, 0xC2, 1, 3, -1, vexNDS3Imm, [3]int{8, 8, 8}}, "VCMPSD": {1, 0xC2, 1, 3, -1, vexNDS3Imm, [3]int{8, 8, 8}},
"VCMPSS": {1, 0xC2, 0, 2, -1, vexNDS3Imm, [3]int{4, 4, 4}}, "VCMPSS": {1, 0xC2, 0, 2, -1, vexNDS3Imm, [3]int{4, 4, 4}},
// EVEX.66.0F3A — integer compares with an opmask destination, the same // EVEX.66.0F3A, integer compares with an opmask destination, the same
// NDS3Imm-with-k-reg shape as the floating-point compares; W selects the // NDS3Imm-with-k-reg shape as the floating-point compares; W selects the
// operand width (byte/word vs dword/qword), the opcode the signedness. // operand width (byte/word vs dword/qword), the opcode the signedness.
// The memory form takes a full vector, so disp8×N is 16/32/64. // The memory form takes a full vector, so disp8×N is 16/32/64.
@@ -179,7 +179,7 @@ var evexTable = map[string]evexSpec{
"VPCMPQ": {3, 0x1F, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VPCMPQ": {3, 0x1F, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPCMPUQ": {3, 0x1E, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}}, "VPCMPUQ": {3, 0x1E, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F38 — permutes (NDS form). // EVEX.66.0F38, permutes (NDS form).
"VPERMB": {2, 0x8D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPERMB": {2, 0x8D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMW": {2, 0x8D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPERMW": {2, 0x8D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2D": {2, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPERMI2D": {2, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -188,7 +188,7 @@ var evexTable = map[string]evexSpec{
"VPERMT2Q": {2, 0x7E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPERMT2Q": {2, 0x7E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2PD": {2, 0x7F, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPERMT2PD": {2, 0x7F, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.66.0F — the wider integer set (NDS form). // EVEX.66.0F, the wider integer set (NDS form).
"VPMADDWD": {1, 0xF5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPMADDWD": {1, 0xF5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -199,34 +199,34 @@ var evexTable = map[string]evexSpec{
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKUSDW": {2, 0x2B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPACKUSDW": {2, 0x2B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.66.0F38 — absolute values and replicating moves (reg=dst, // EVEX.66.0F38, absolute values and replicating moves (reg=dst,
// rm=src). // rm=src).
"VPABSB": {2, 0x1C, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VPABSB": {2, 0x1C, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPABSW": {2, 0x1D, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VPABSW": {2, 0x1D, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPABSD": {2, 0x1E, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VPABSD": {2, 0x1E, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPABSQ": {2, 0x1F, 1, 1, -1, vexRM, [3]int{16, 32, 64}}, "VPABSQ": {2, 0x1F, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.F3.0F — replicate even/odd singles. // EVEX.F3.0F, replicate even/odd singles.
"VMOVSLDUP": {1, 0x12, 0, 2, -1, vexRM, [3]int{16, 32, 64}}, "VMOVSLDUP": {1, 0x12, 0, 2, -1, vexRM, [3]int{16, 32, 64}},
"VMOVSHDUP": {1, 0x16, 0, 2, -1, vexRM, [3]int{16, 32, 64}}, "VMOVSHDUP": {1, 0x16, 0, 2, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.66.0F38 — sign/zero-extending moves; the memory source is the // EVEX.66.0F38, sign/zero-extending moves; the memory source is the
// narrow half (here byte to word). // narrow half (here byte to word).
"VPMOVSXBW": {2, 0x20, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VPMOVSXBW": {2, 0x20, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
"VPMOVZXBW": {2, 0x30, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VPMOVZXBW": {2, 0x30, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F — packed single conversions (reg=dst, rm=src). // EVEX.66.0F, packed single conversions (reg=dst, rm=src).
"VCVTPS2DQ": {1, 0x5B, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VCVTPS2DQ": {1, 0x5B, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VCVTTPS2DQ": {1, 0x5B, 0, 2, -1, vexRM, [3]int{16, 32, 64}}, "VCVTTPS2DQ": {1, 0x5B, 0, 2, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.66.0F38 — broadcast a single/double to all lanes (reg=dst, // EVEX.66.0F38, broadcast a single/double to all lanes (reg=dst,
// rm=scalar memory; disp8×N is the element size). // rm=scalar memory; disp8×N is the element size).
"VBROADCASTSS": {2, 0x18, 0, 1, -1, vexRM, [3]int{4, 4, 4}}, "VBROADCASTSS": {2, 0x18, 0, 1, -1, vexRM, [3]int{4, 4, 4}},
"VBROADCASTSD": {2, 0x19, 1, 1, -1, vexRM, [3]int{0, 8, 8}}, "VBROADCASTSD": {2, 0x19, 1, 1, -1, vexRM, [3]int{0, 8, 8}},
// EVEX.66.0F38 — expand loads (rm → vector register destination). // EVEX.66.0F38, expand loads (rm → vector register destination).
"VEXPANDPD": {2, 0x88, 1, 1, -1, vexRM, [3]int{8, 8, 8}}, "VEXPANDPD": {2, 0x88, 1, 1, -1, vexRM, [3]int{8, 8, 8}},
"VEXPANDPS": {2, 0x88, 0, 1, -1, vexRM, [3]int{4, 4, 4}}, "VEXPANDPS": {2, 0x88, 0, 1, -1, vexRM, [3]int{4, 4, 4}},
"VPEXPANDD": {2, 0x89, 0, 1, -1, vexRM, [3]int{4, 4, 4}}, "VPEXPANDD": {2, 0x89, 0, 1, -1, vexRM, [3]int{4, 4, 4}},
"VPEXPANDQ": {2, 0x89, 1, 1, -1, vexRM, [3]int{8, 8, 8}}, "VPEXPANDQ": {2, 0x89, 1, 1, -1, vexRM, [3]int{8, 8, 8}},
// EVEX.66.0F38 — compress stores (vector register source → rm), and the // EVEX.66.0F38, compress stores (vector register source → rm), and the
// remaining narrowing stores. // remaining narrowing stores.
"VCOMPRESSPD": {2, 0x8A, 1, 1, -1, vexRMRev, [3]int{8, 8, 8}}, "VCOMPRESSPD": {2, 0x8A, 1, 1, -1, vexRMRev, [3]int{8, 8, 8}},
"VCOMPRESSPS": {2, 0x8A, 0, 1, -1, vexRMRev, [3]int{4, 4, 4}}, "VCOMPRESSPS": {2, 0x8A, 0, 1, -1, vexRMRev, [3]int{4, 4, 4}},
@@ -235,7 +235,7 @@ var evexTable = map[string]evexSpec{
"VPMOVWB": {2, 0x30, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}}, "VPMOVWB": {2, 0x30, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
"VPMOVQB": {2, 0x32, 0, 2, -1, vexRMRev, [3]int{2, 4, 8}}, "VPMOVQB": {2, 0x32, 0, 2, -1, vexRMRev, [3]int{2, 4, 8}},
// EVEX.66.0F — rotates (immediate form: /0 right, /1 left). // EVEX.66.0F, rotates (immediate form: /0 right, /1 left).
"VPRORD": {1, 0x72, 0, 1, 0, vexShiftImm, [3]int{16, 32, 64}}, "VPRORD": {1, 0x72, 0, 1, 0, vexShiftImm, [3]int{16, 32, 64}},
"VPRORQ": {1, 0x72, 1, 1, 0, vexShiftImm, [3]int{16, 32, 64}}, "VPRORQ": {1, 0x72, 1, 1, 0, vexShiftImm, [3]int{16, 32, 64}},
"VPROLD": {1, 0x72, 0, 1, 1, vexShiftImm, [3]int{16, 32, 64}}, "VPROLD": {1, 0x72, 0, 1, 1, vexShiftImm, [3]int{16, 32, 64}},
@@ -248,14 +248,14 @@ var evexTable = map[string]evexSpec{
"VPSRLQ": {1, 0x73, 1, 1, 2, vexShiftImm, [3]int{16, 32, 64}}, "VPSRLQ": {1, 0x73, 1, 1, 2, vexShiftImm, [3]int{16, 32, 64}},
"VPSLLQ": {1, 0x73, 1, 1, 6, vexShiftImm, [3]int{16, 32, 64}}, "VPSLLQ": {1, 0x73, 1, 1, 6, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.66.0F38 — floating-point helpers, packed (reg=dst, rm=src). // EVEX.66.0F38, floating-point helpers, packed (reg=dst, rm=src).
"VRCP14PD": {2, 0x4C, 1, 1, -1, vexRM, [3]int{16, 32, 64}}, "VRCP14PD": {2, 0x4C, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VRCP14PS": {2, 0x4C, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VRCP14PS": {2, 0x4C, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VRSQRT14PD": {2, 0x4E, 1, 1, -1, vexRM, [3]int{16, 32, 64}}, "VRSQRT14PD": {2, 0x4E, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VRSQRT14PS": {2, 0x4E, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VRSQRT14PS": {2, 0x4E, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VGETEXPPD": {2, 0x42, 1, 1, -1, vexRM, [3]int{16, 32, 64}}, "VGETEXPPD": {2, 0x42, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VGETEXPPS": {2, 0x42, 0, 1, -1, vexRM, [3]int{16, 32, 64}}, "VGETEXPPS": {2, 0x42, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.66.0F38 — floating-point helpers, scalar (NDS form: src2 is // EVEX.66.0F38, floating-point helpers, scalar (NDS form: src2 is
// rm, src1 is vvvv, the XMM destination is reg). Like the scalar 0F3A // rm, src1 is vvvv, the XMM destination is reg). Like the scalar 0F3A
// forms, these take the 66 prefix; W selects double/single. // forms, these take the 66 prefix; W selects double/single.
"VRCP14SD": {2, 0x4D, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}}, "VRCP14SD": {2, 0x4D, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}},
@@ -264,13 +264,13 @@ var evexTable = map[string]evexSpec{
"VRSQRT14SS": {2, 0x4F, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}}, "VRSQRT14SS": {2, 0x4F, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}},
"VGETEXPSD": {2, 0x43, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}}, "VGETEXPSD": {2, 0x43, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}},
"VGETEXPSS": {2, 0x43, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}}, "VGETEXPSS": {2, 0x43, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}},
// EVEX.66.0F38 — scale by a power of two (NDS form). // EVEX.66.0F38, scale by a power of two (NDS form).
"VSCALEFPD": {2, 0x2C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VSCALEFPD": {2, 0x2C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VSCALEFPS": {2, 0x2C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VSCALEFPS": {2, 0x2C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VSCALEFSD": {2, 0x2D, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}}, "VSCALEFSD": {2, 0x2D, 1, 1, -1, vexNDS3, [3]int{8, 8, 8}},
"VSCALEFSS": {2, 0x2D, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}}, "VSCALEFSS": {2, 0x2D, 0, 1, -1, vexNDS3, [3]int{4, 4, 4}},
// EVEX.66.0F3A — packed round/getmant/reduce ($imm, src, dst: reg=dst, // EVEX.66.0F3A, packed round/getmant/reduce ($imm, src, dst: reg=dst,
// rm=src, imm8). // rm=src, imm8).
"VRNDSCALEPD": {3, 0x09, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VRNDSCALEPD": {3, 0x09, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VRNDSCALEPS": {3, 0x08, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VRNDSCALEPS": {3, 0x08, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
@@ -278,7 +278,7 @@ var evexTable = map[string]evexSpec{
"VGETMANTPS": {3, 0x26, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VGETMANTPS": {3, 0x26, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VREDUCEPD": {3, 0x56, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VREDUCEPD": {3, 0x56, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VREDUCEPS": {3, 0x56, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VREDUCEPS": {3, 0x56, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.66.0F3A — scalar round/getmant/reduce and fixup/range (NDS + // EVEX.66.0F3A, scalar round/getmant/reduce and fixup/range (NDS +
// imm8: $imm, src2, src1, dst). The scalar 0F3A forms all take the 66 // imm8: $imm, src2, src1, dst). The scalar 0F3A forms all take the 66
// prefix; W selects double/single. // prefix; W selects double/single.
"VRNDSCALESD": {3, 0x0B, 1, 1, -1, vexNDS3Imm, [3]int{8, 8, 8}}, "VRNDSCALESD": {3, 0x0B, 1, 1, -1, vexNDS3Imm, [3]int{8, 8, 8}},
@@ -296,7 +296,7 @@ var evexTable = map[string]evexSpec{
"VRANGESD": {3, 0x51, 1, 1, -1, vexNDS3Imm, [3]int{8, 8, 8}}, "VRANGESD": {3, 0x51, 1, 1, -1, vexNDS3Imm, [3]int{8, 8, 8}},
"VRANGESS": {3, 0x51, 0, 1, -1, vexNDS3Imm, [3]int{4, 4, 4}}, "VRANGESS": {3, 0x51, 0, 1, -1, vexNDS3Imm, [3]int{4, 4, 4}},
// EVEX.66.0F3A — floating-point class test ($imm, src, kdst): the // EVEX.66.0F3A, floating-point class test ($imm, src, kdst): the
// reg field carries the opmask destination. The packed forms carry an // reg field carries the opmask destination. The packed forms carry an
// explicit length in the mnemonic (X/Y/Z). // explicit length in the mnemonic (X/Y/Z).
"VFPCLASSPDX": {3, 0x66, 1, 1, -1, vexImmRM, [3]int{16, 0, 0}}, "VFPCLASSPDX": {3, 0x66, 1, 1, -1, vexImmRM, [3]int{16, 0, 0}},
@@ -308,7 +308,7 @@ var evexTable = map[string]evexSpec{
"VFPCLASSSD": {3, 0x67, 1, 1, -1, vexImmRM, [3]int{8, 0, 0}}, "VFPCLASSSD": {3, 0x67, 1, 1, -1, vexImmRM, [3]int{8, 0, 0}},
"VFPCLASSSS": {3, 0x67, 0, 1, -1, vexImmRM, [3]int{4, 0, 0}}, "VFPCLASSSS": {3, 0x67, 0, 1, -1, vexImmRM, [3]int{4, 0, 0}},
// EVEX — the remaining conversions. VCVTQQ2PS narrows (the 512-bit // EVEX, the remaining conversions. VCVTQQ2PS narrows (the 512-bit
// source sets the length); the rest follow the destination. // source sets the length); the rest follow the destination.
"VCVTQQ2PS": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{0, 0, 64}}, "VCVTQQ2PS": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{0, 0, 64}},
"VCVTPD2QQ": {1, 0x7B, 1, 1, -1, vexRM, [3]int{16, 32, 64}}, "VCVTPD2QQ": {1, 0x7B, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
@@ -316,13 +316,13 @@ var evexTable = map[string]evexSpec{
"VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}}, "VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}}, "VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F38 — half-precision convert (half-width source). // EVEX.66.0F38, half-precision convert (half-width source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F3A — half-precision convert back ($imm, src, dst: reg=src, // EVEX.66.0F3A, half-precision convert back ($imm, src, dst: reg=src,
// rm=dst, imm8 — the extract layout). // rm=dst, imm8, the extract layout).
"VCVTPS2PH": {3, 0x1D, 0, 1, -1, vexExtract, [3]int{8, 16, 32}}, "VCVTPS2PH": {3, 0x1D, 0, 1, -1, vexExtract, [3]int{8, 16, 32}},
// EVEX — unsigned and truncating conversions. The PD sources are the // EVEX, unsigned and truncating conversions. The PD sources are the
// wide operand (the bare names are 512-bit only, the X/Y spellings fix // wide operand (the bare names are 512-bit only, the X/Y spellings fix
// the length); the PS/UQQ destinations are wide and follow the // the length); the PS/UQQ destinations are wide and follow the
// destination. // destination.
@@ -349,7 +349,7 @@ var evexTable = map[string]evexSpec{
"VCVTQQ2PSX": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{16, 0, 0}}, "VCVTQQ2PSX": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{16, 0, 0}},
"VCVTQQ2PSY": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{0, 32, 0}}, "VCVTQQ2PSY": {1, 0x5B, 1, 0, -1, vexRMSrcLen, [3]int{0, 32, 0}},
// EVEX.66.0F38 — the remaining sign/zero-extending moves (narrow // EVEX.66.0F38, the remaining sign/zero-extending moves (narrow
// source; disp8×N follows its size). // source; disp8×N follows its size).
"VPMOVSXBD": {2, 0x21, 0, 1, -1, vexRM, [3]int{4, 8, 16}}, "VPMOVSXBD": {2, 0x21, 0, 1, -1, vexRM, [3]int{4, 8, 16}},
"VPMOVSXBQ": {2, 0x22, 0, 1, -1, vexRM, [3]int{2, 4, 8}}, "VPMOVSXBQ": {2, 0x22, 0, 1, -1, vexRM, [3]int{2, 4, 8}},
@@ -361,7 +361,7 @@ var evexTable = map[string]evexSpec{
"VPMOVZXWQ": {2, 0x34, 0, 1, -1, vexRM, [3]int{4, 8, 16}}, "VPMOVZXWQ": {2, 0x34, 0, 1, -1, vexRM, [3]int{4, 8, 16}},
"VPMOVZXDQ": {2, 0x35, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VPMOVZXDQ": {2, 0x35, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.F3.0F38 — the remaining narrowing stores (vector source in reg, // EVEX.F3.0F38, the remaining narrowing stores (vector source in reg,
// narrow destination in r/m): signed, unsigned and the D/Q truncations. // narrow destination in r/m): signed, unsigned and the D/Q truncations.
"VPMOVSDB": {2, 0x21, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}}, "VPMOVSDB": {2, 0x21, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}},
"VPMOVSQB": {2, 0x22, 0, 2, -1, vexRMRev, [3]int{2, 4, 8}}, "VPMOVSQB": {2, 0x22, 0, 2, -1, vexRMRev, [3]int{2, 4, 8}},
@@ -378,7 +378,7 @@ var evexTable = map[string]evexSpec{
"VPMOVDB": {2, 0x31, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}}, "VPMOVDB": {2, 0x31, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}},
"VPMOVQW": {2, 0x34, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}}, "VPMOVQW": {2, 0x34, 0, 2, -1, vexRMRev, [3]int{4, 8, 16}},
// EVEX.F3.0F38 — mask/vector conversions: M2* moves an opmask register // EVEX.F3.0F38, mask/vector conversions: M2* moves an opmask register
// into a vector (rm = K source, reg = vector destination), *2M does the // into a vector (rm = K source, reg = vector destination), *2M does the
// reverse (reg = K destination, rm = vector source, the length follows // reverse (reg = K destination, rm = vector source, the length follows
// the vector). // the vector).
@@ -391,7 +391,7 @@ var evexTable = map[string]evexSpec{
"VPMOVD2M": {2, 0x39, 0, 2, -1, vexRM, [3]int{16, 32, 64}}, "VPMOVD2M": {2, 0x39, 0, 2, -1, vexRM, [3]int{16, 32, 64}},
"VPMOVQ2M": {2, 0x39, 1, 2, -1, vexRM, [3]int{16, 32, 64}}, "VPMOVQ2M": {2, 0x39, 1, 2, -1, vexRM, [3]int{16, 32, 64}},
// EVEX — scalar conversions between vector and general-purpose // EVEX, scalar conversions between vector and general-purpose
// registers. Vector to GPR (two operands: vec/mem source, GPR // registers. Vector to GPR (two operands: vec/mem source, GPR
// destination, vvvv unused): the signed and truncated pair, and the // destination, vvvv unused): the signed and truncated pair, and the
// unsigned forms (EVEX only). // unsigned forms (EVEX only).
@@ -421,22 +421,22 @@ var evexTable = map[string]evexSpec{
"VCVTUSI2SDQ": {1, 0x7B, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}}, "VCVTUSI2SDQ": {1, 0x7B, 1, 3, -1, vexNDS3, [3]int{8, 8, 8}},
"VCVTUSI2SSL": {1, 0x7B, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}}, "VCVTUSI2SSL": {1, 0x7B, 0, 2, -1, vexNDS3, [3]int{4, 4, 4}},
"VCVTUSI2SSQ": {1, 0x7B, 1, 2, -1, vexNDS3, [3]int{8, 8, 8}}, "VCVTUSI2SSQ": {1, 0x7B, 1, 2, -1, vexNDS3, [3]int{8, 8, 8}},
// EVEX.128/256/512.66.0F38.W0 — sign-extend dwords to qwords; the memory // EVEX.128/256/512.66.0F38.W0, sign-extend dwords to qwords; the memory
// operand is the narrow source, so disp8×N follows its size (8/16/32 for // operand is the narrow source, so disp8×N follows its size (8/16/32 for
// the xmm/ymm/zmm destination lengths). // the xmm/ymm/zmm destination lengths).
"VPMOVSXDQ": {2, 0x25, 0, 1, -1, vexRM, [3]int{8, 16, 32}}, "VPMOVSXDQ": {2, 0x25, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.512.66.0F3A.W1 — lane extract (reg=ZMM source, rm=YMM/memory // EVEX.512.66.0F3A.W1, lane extract (reg=ZMM source, rm=YMM/memory
// destination, imm8). // destination, imm8).
"VEXTRACTI64X4": {3, 0x3B, 1, 1, -1, vexExtract, [3]int{0, 0, 32}}, "VEXTRACTI64X4": {3, 0x3B, 1, 1, -1, vexExtract, [3]int{0, 0, 32}},
"VEXTRACTF64X4": {3, 0x1B, 1, 1, -1, vexExtract, [3]int{0, 0, 32}}, "VEXTRACTF64X4": {3, 0x1B, 1, 1, -1, vexExtract, [3]int{0, 0, 32}},
// EVEX.66.0F38 — more integer NDS forms (W distinguishes D/Q). // EVEX.66.0F38, more integer NDS forms (W distinguishes D/Q).
"VPMULLD": {2, 0x40, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPMULLD": {2, 0x40, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULLQ": {2, 0x40, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPMULLQ": {2, 0x40, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMD": {2, 0x36, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}}, "VPERMD": {2, 0x36, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}},
// EVEX.128/256/512 — the wider integer set (AVX-512 F/BW): byte/word // EVEX.128/256/512, the wider integer set (AVX-512 F/BW): byte/word
// arithmetic, the bitwise ops with D/Q suffixes, min/max, averages and // arithmetic, the bitwise ops with D/Q suffixes, min/max, averages and
// variable shifts. All NDS form; W distinguishes element size. // variable shifts. All NDS form; W distinguishes element size.
"VPADDB": {1, 0xFC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPADDB": {1, 0xFC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -474,21 +474,21 @@ var evexTable = map[string]evexSpec{
"VPSRAVQ": {2, 0x46, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPSRAVQ": {2, 0x46, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX forms of instructions that also exist in VEX (selected when a ZMM // EVEX forms of instructions that also exist in VEX (selected when a ZMM
// or K register, or indices 16–31, demand EVEX). // or K register, or indices 16-31, demand EVEX).
"VPSHUFD": {1, 0x70, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}}, "VPSHUFD": {1, 0x70, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPSHUFB": {2, 0x00, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}}, "VPSHUFB": {2, 0x00, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.66.0F — immediate shift (VPSLLD /6). // EVEX.66.0F, immediate shift (VPSLLD /6).
"VPSLLD": {1, 0x72, 0, 1, 6, vexShiftImm, [3]int{16, 32, 64}}, "VPSLLD": {1, 0x72, 0, 1, 6, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.F3.0F38.W0 — narrowing stores: reg = wide source, rm = narrow // EVEX.F3.0F38.W0, narrowing stores: reg = wide source, rm = narrow
// destination (VPMOVDW dword→word, VPMOVQD qword→dword). // destination (VPMOVDW dword→word, VPMOVQD qword→dword).
"VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}}, "VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
"VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}}, "VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
} }
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode // evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
// depends on the source kind — a GPR source uses opReg, a memory source uses // depends on the source kind, a GPR source uses opReg, a memory source uses
// opMem with a disp8×N of n. // opMem with a disp8×N of n.
type evexBcastSpec struct { type evexBcastSpec struct {
mapSel int mapSel int
@@ -499,50 +499,53 @@ type evexBcastSpec struct {
} }
var evexBcastTable = map[string]evexBcastSpec{ var evexBcastTable = map[string]evexBcastSpec{
// EVEX.128/256/512.66.0F38 — broadcast a dword/qword to all lanes. // EVEX.128/256/512.66.0F38, broadcast a dword/qword to all lanes.
"VPBROADCASTD": {2, 0x7C, 0x58, 0, 4}, "VPBROADCASTD": {2, 0x7C, 0x58, 0, 4},
"VPBROADCASTQ": {2, 0x7C, 0x59, 1, 8}, "VPBROADCASTQ": {2, 0x7C, 0x59, 1, 8},
// EVEX.128/256/512.66.0F38 — broadcast a byte/word (GPR or memory // EVEX.128/256/512.66.0F38, broadcast a byte/word (GPR or memory
// source) to all lanes. // source) to all lanes.
"VPBROADCASTB": {2, 0x7A, 0x78, 0, 1}, "VPBROADCASTB": {2, 0x7A, 0x78, 0, 1},
"VPBROADCASTW": {2, 0x7B, 0x79, 0, 2}, "VPBROADCASTW": {2, 0x7B, 0x79, 0, 2},
} }
// evexMoveSpec describes an EVEX move (load and store opcodes, like the VEX // evexMoveSpec describes an EVEX move (load and store opcodes, like the VEX
// move table). // move table). vecOK and xmmOnly mirror the VEX twin's operand rules: a
// scalar move (vecOK false, xmmOnly true) takes XMM↔memory operands only.
type evexMoveSpec struct { type evexMoveSpec struct {
mapSel int mapSel int
pp int pp int
load byte // r/m → vector load byte // r/m → vector
store byte // vector → r/m store byte // vector → r/m
w int w int
n [3]int n [3]int
vecOK bool // the non-memory operand may be a vector register
xmmOnly bool // wider than XMM registers are rejected
} }
// evexMoveTable maps an upper-case EVEX move mnemonic to its encoding. // evexMoveTable maps an upper-case EVEX move mnemonic to its encoding.
var evexMoveTable = map[string]evexMoveSpec{ var evexMoveTable = map[string]evexMoveSpec{
// EVEX.128/256/512.F3.0F.W0 — unaligned integer move. // EVEX.128/256/512.F3.0F.W0, unaligned integer move.
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}}, "VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512.F3.0F.W1 — unaligned qword move. // EVEX.128/256/512.F3.0F.W1, unaligned qword move.
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}}, "VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512.F2.0F.W0 — unaligned byte move (byte/word moves use the // EVEX.128/256/512.F2.0F.W0, unaligned byte move (byte/word moves use the
// F2 prefix, dword/qword moves F3; the element size only changes the tuple // F2 prefix, dword/qword moves F3; the element size only changes the tuple
// semantics). // semantics).
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}}, "VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512.F2.0F.W1 — unaligned word move (shares the qword // EVEX.128/256/512.F2.0F.W1, unaligned word move (shares the qword
// encoding). // encoding).
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}}, "VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512.66.0F.W1 — unaligned packed double move. // EVEX.128/256/512.66.0F.W1, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}}, "VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512 — aligned packed moves. // EVEX.128/256/512, aligned packed moves.
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}}, "VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}}, "VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false},
// EVEX.128/256/512.66.0F — aligned integer moves. // EVEX.128/256/512.66.0F, aligned integer moves.
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}}, "VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}}, "VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
// EVEX.128.F3.0F.W0 — scalar single move, memory operands (the // EVEX.128.F3.0F.W0, scalar single move, memory operands (the
// three-operand register form is not supported). // three-operand register form is not supported).
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}}, "VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true},
} }
// isEvex reports whether the mnemonic has an EVEX encoding we handle. // isEvex reports whether the mnemonic has an EVEX encoding we handle.
@@ -559,7 +562,7 @@ func isEvex(mnemUpper string) bool {
// evexRequired reports whether the operands force the EVEX encoding of a // evexRequired reports whether the operands force the EVEX encoding of a
// mnemonic that also has a VEX form: ZMM and K registers do, and so do // mnemonic that also has a VEX form: ZMM and K registers do, and so do
// register indices 16–31, which only EVEX can represent (X16–Y31 exist // register indices 16-31, which only EVEX can represent (X16-Y31 exist
// solely under AVX-512). // solely under AVX-512).
func evexRequired(upper string, ops []Operand) bool { func evexRequired(upper string, ops []Operand) bool {
_, inVex := vexTable[upper] _, inVex := vexTable[upper]
@@ -578,7 +581,7 @@ func evexRequired(upper string, ops []Operand) bool {
// evexSuffix carries the EVEX mnemonic suffixes the Go assembler accepts: // evexSuffix carries the EVEX mnemonic suffixes the Go assembler accepts:
// zeroing (.Z), a rounding mode (.RN_SAE, .RD_SAE, .RU_SAE, .RZ_SAE), // zeroing (.Z), a rounding mode (.RN_SAE, .RD_SAE, .RU_SAE, .RZ_SAE),
// suppress-all-exceptions (.SAE) and memory broadcast (.BCST). Masking is // suppress-all-exceptions (.SAE) and memory broadcast (.BCST). Masking is
// not a suffix — Go writes it as an explicit K operand. // not a suffix, Go writes it as an explicit K operand.
type evexSuffix struct { type evexSuffix struct {
zeroing bool zeroing bool
sae bool sae bool
@@ -668,7 +671,7 @@ var evexRound = map[string]bool{
} }
// evexBcstN maps an instruction accepting .BCST to the broadcast element // evexBcstN maps an instruction accepting .BCST to the broadcast element
// size — the disp8×N multiplier for its memory operand. // size, the disp8×N multiplier for its memory operand.
var evexBcstN = map[string]int{ var evexBcstN = map[string]int{
"VADDPD": 8, "VSUBPD": 8, "VMULPD": 8, "VDIVPD": 8, "VADDPD": 8, "VSUBPD": 8, "VMULPD": 8, "VDIVPD": 8,
"VMINPD": 8, "VMAXPD": 8, "VMINPD": 8, "VMAXPD": 8,
@@ -689,7 +692,7 @@ var evexBcstN = map[string]int{
"VCVTTPD2QQ": 8, "VCVTTPS2QQ": 4, "VCVTUQQ2PD": 8, "VCVTUQQ2PS": 8, "VCVTTPD2QQ": 8, "VCVTTPS2QQ": 4, "VCVTUQQ2PD": 8, "VCVTUQQ2PS": 8,
} }
// splitMask extracts an explicit mask register (K1–K7) from the operand list, // splitMask extracts an explicit mask register (K1-K7) from the operand list,
// returning the remaining operands and the mask index. K0 is not a usable // returning the remaining operands and the mask index. K0 is not a usable
// mask (aaa = 0 means "no mask"), matching the assembler. // mask (aaa = 0 means "no mask"), matching the assembler.
func splitMask(ops []Operand) ([]Operand, int, error) { func splitMask(ops []Operand) ([]Operand, int, error) {
@@ -712,7 +715,7 @@ func splitMask(ops []Operand) ([]Operand, int, error) {
} }
// encodeEvex encodes an EVEX instruction with operands in Plan 9 order. The // encodeEvex encodes an EVEX instruction with operands in Plan 9 order. The
// mask, when present, is an explicit K1–K7 operand anywhere among the // mask, when present, is an explicit K1-K7 operand anywhere among the
// operands; the mnemonic suffix carries zeroing, rounding/SAE and // operands; the mnemonic suffix carries zeroing, rounding/SAE and
// broadcast. // broadcast.
func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error { func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error {
@@ -1022,6 +1025,12 @@ func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask i
var rm Operand var rm Operand
switch { switch {
case srcIsVec && dstIsVec: case srcIsVec && dstIsVec:
// A store-form reg-reg move, the layout the Go assembler uses; a
// scalar move has no two-register form at all (the register form
// takes three operands), matching the VEX twin's vecOK rule.
if !ms.vecOK {
return fmt.Errorf("%s does not take two vector registers", mnem)
}
reg, rm = srcReg, dst reg, rm = srcReg, dst
case srcIsVec: case srcIsVec:
if !memOperand(dst) { if !memOperand(dst) {
@@ -1037,12 +1046,18 @@ func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask i
default: default:
return fmt.Errorf("%s needs a vector register operand", mnem) return fmt.Errorf("%s needs a vector register operand", mnem)
} }
// The scalar move is 128-bit only, so the register the length follows
// must be an XMM (the VEX twin's xmmOnly rule; EVEX also reaches ZMM,
// hence the inequality rather than a YMM test).
if ms.xmmOnly && reg.size != 16 {
return fmt.Errorf("%s operates on XMM registers only", mnem)
}
spec := evexSpec{mapSel: ms.mapSel, opcode: op, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n} spec := evexSpec{mapSel: ms.mapSel, opcode: op, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n}
return e.emitEvexFields(spec, reg.vecLenBit(), reg.idx, -1, rm, mask, sfx) return e.emitEvexFields(spec, reg.vecLenBit(), reg.idx, -1, rm, mask, sfx)
} }
// encodeEvexRMSrcLen encodes a length-narrowing conversion: OP src, dst with // encodeEvexRMSrcLen encodes a length-narrowing conversion: OP src, dst with
// the destination always XMM and the length fixed by the mnemonic — the // the destination always XMM and the length fixed by the mnemonic, the
// single valid slot of spec.n names the vector length (and the disp8×N // single valid slot of spec.n names the vector length (and the disp8×N
// multiplier) a register or memory source encodes. // multiplier) a register or memory source encodes.
func (e *enc) encodeEvexRMSrcLen(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error { func (e *enc) encodeEvexRMSrcLen(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error {
@@ -1061,7 +1076,7 @@ func (e *enc) encodeEvexRMSrcLen(spec evexSpec, ops []Operand, mask int, sfx eve
return e.emitEvexFields(spec, ll, dstReg.idx, -1, src, mask, sfx) return e.emitEvexFields(spec, ll, dstReg.idx, -1, src, mask, sfx)
} }
// soleLen returns the vector-length index of the single valid slot of n — // soleLen returns the vector-length index of the single valid slot of n
// the length a length-fixed mnemonic (the EVEX conversion spellings) encodes // the length a length-fixed mnemonic (the EVEX conversion spellings) encodes
// regardless of its operands. // regardless of its operands.
func soleLen(n [3]int) (int, error) { func soleLen(n [3]int) (int, error) {
@@ -1131,8 +1146,8 @@ func (e *enc) encodeEvexBcast(bs evexBcastSpec, ops []Operand, mask int, sfx eve
// emitEvexFields emits the EVEX prefix, opcode, ModR/M, SIB and displacement // emitEvexFields emits the EVEX prefix, opcode, ModR/M, SIB and displacement
// (disp8×N compressed) for the given precomputed fields. regIdx is the // (disp8×N compressed) for the given precomputed fields. regIdx is the
// unextended reg-field register index, or a /digit (0–7); vvvvIdx is the // unextended reg-field register index, or a /digit (0-7); vvvvIdx is the
// vvvv register index, or -1 when unused. mask (K1–K7, 0 = unmasked) and // vvvv register index, or -1 when unused. mask (K1-K7, 0 = unmasked) and
// zeroing fill the aaa and z bits of the P2 byte. // zeroing fill the aaa and z bits of the P2 byte.
func (e *enc) emitEvexFields(spec evexSpec, ll, regIdx, vvvvIdx int, rm Operand, mask int, sfx evexSuffix) error { func (e *enc) emitEvexFields(spec evexSpec, ll, regIdx, vvvvIdx int, rm Operand, mask int, sfx evexSuffix) error {
if ll > 2 { if ll > 2 {
@@ -1171,9 +1186,6 @@ func (e *enc) emitEvexFields(spec evexSpec, ll, regIdx, vvvvIdx int, rm Operand,
if r.idx&16 != 0 { if r.idx&16 != 0 {
xBar = 0 xBar = 0
} }
if r.idx&16 != 0 {
xBar = 0
}
case Mem: case Mem:
var err error var err error
modrm, sib, disp, xBar, bBar, err = memComponentsEvex(regIdx&7, r, spec.n[ll]) modrm, sib, disp, xBar, bBar, err = memComponentsEvex(regIdx&7, r, spec.n[ll])
@@ -1232,6 +1244,11 @@ func (e *enc) emitEvexFields(spec evexSpec, ll, regIdx, vvvvIdx int, rm Operand,
func memComponentsEvex(regField int, m Mem, n int) (modrm, sib int, disp []byte, xBar, bBar int, err error) { func memComponentsEvex(regField int, m Mem, n int) (modrm, sib int, disp []byte, xBar, bBar int, err error) {
sib = -1 sib = -1
xBar, bBar = 1, 1 // inverted bits: 1 = no extension xBar, bBar = 1, 1 // inverted bits: 1 = no extension
// The disp32 fallback bounds the displacement by int32, and the
// compressed disp8 form reaches at most ±127×64, well inside it.
if m.Disp < -(1<<31) || m.Disp > (1<<31)-1 {
return 0, -1, nil, 0, 0, fmt.Errorf("displacement %d does not fit in 32 bits", m.Disp)
}
if !m.HasBase && !m.HasIndex { if !m.HasBase && !m.HasIndex {
return regField<<3 | 0x05, -1, le32(m.Disp), 1, 1, nil // RIP-relative return regField<<3 | 0x05, -1, le32(m.Disp), 1, 1, nil // RIP-relative
} }
@@ -1324,7 +1341,7 @@ func isScatter(upper string) bool {
} }
// vsibLen validates a VSIB memory operand (the index must be a vector // vsibLen validates a VSIB memory operand (the index must be a vector
// register) and returns it with the vector length the index selects — the // register) and returns it with the vector length the index selects, the
// EVEX L'L field follows the index register, not the data register. // EVEX L'L field follows the index register, not the data register.
func vsibLen(op Operand, what string) (Mem, int, error) { func vsibLen(op Operand, what string) (Mem, int, error) {
m, ok := op.(Mem) m, ok := op.(Mem)
@@ -1383,7 +1400,7 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
return e.emitVexFields(spec, dst.vecLenBit(), dst.idx&7, rBit, 15-maskReg.idx, vsib) return e.emitVexFields(spec, dst.vecLenBit(), dst.idx&7, rBit, 15-maskReg.idx, vsib)
} }
// encodeScatter encodes a scatter (EVEX only): OP src, K, vsib — reg = src, // encodeScatter encodes a scatter (EVEX only): OP src, K, vsib, reg = src,
// rm = the VSIB memory operand, the K mask in aaa and L following the VSIB // rm = the VSIB memory operand, the K mask in aaa and L following the VSIB
// index. // index.
func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evexSuffix) error { func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evexSuffix) error {
@@ -1411,15 +1428,15 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
// evexKOperand lists the instructions whose K register is a genuine operand // evexKOperand lists the instructions whose K register is a genuine operand
// (the source or destination of a mask/vector conversion) rather than a // (the source or destination of a mask/vector conversion) rather than a
// mask modifier — the M2 and 2M conversions. They take no masking. // mask modifier, the M2 and 2M conversions. They take no masking.
var evexKOperand = map[string]bool{ var evexKOperand = map[string]bool{
"VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true, "VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true,
"VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true, "VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true,
} }
// kmovSpec describes a KMOV width: the opcode depends on the operand // kmovSpec describes a KMOV width: the opcode depends on the operand
// direction — kk (k/mem → K is 90, k → k uses the same), kmem (K → mem), // direction, kk (k/mem → K is 90, k → k uses the same), kmem (K → mem),
// gprk (GPR/mem → K), kgpr (K → GPR) — and the GPR forms carry a mandatory // gprk (GPR/mem → K), kgpr (K → GPR), and the GPR forms carry a mandatory
// prefix and W for the wider widths. // prefix and W for the wider widths.
type kmovSpec struct { type kmovSpec struct {
kk, kmem, gprk, kgpr byte kk, kmem, gprk, kgpr byte
+22 -13
View File
@@ -16,7 +16,7 @@ import (
// kernels use: NDS arithmetic, immediate and variable shifts, shuffles with // kernels use: NDS arithmetic, immediate and variable shifts, shuffles with
// an immediate, lane extracts, narrowing stores, broadcasts from a GPR or // an immediate, lane extracts, narrowing stores, broadcasts from a GPR or
// memory, mask destinations, mask moves, disp8×N compression and the 5-bit // memory, mask destinations, mask moves, disp8×N compression and the 5-bit
// register fields (X/Y 16–31, Z 0–31). // register fields (X/Y 16-31, Z 0-31).
func TestEvexGroundTruth(t *testing.T) { func TestEvexGroundTruth(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
@@ -58,7 +58,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VMOVDQU32 16(SI)(R15*4),Z4", "VMOVDQU32", []Operand{Idx(SI, vreg(t, "R15"), 4, 16, 64), vreg(t, "Z4")}, "62b17e486fa4be10000000"}, {"VMOVDQU32 16(SI)(R15*4),Z4", "VMOVDQU32", []Operand{Idx(SI, vreg(t, "R15"), 4, 16, 64), vreg(t, "Z4")}, "62b17e486fa4be10000000"},
{"VMOVDQU32 Z0,4(SI)(AX*1)", "VMOVDQU32", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f17e487f840604000000"}, {"VMOVDQU32 Z0,4(SI)(AX*1)", "VMOVDQU32", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f17e487f840604000000"},
{"VMOVDQU32 Z3,(DI)(R15*4)", "VMOVDQU32", []Operand{vreg(t, "Z3"), Idx(DI, vreg(t, "R15"), 4, 0, 64)}, "62b17e487f1cbf"}, {"VMOVDQU32 Z3,(DI)(R15*4)", "VMOVDQU32", []Operand{vreg(t, "Z3"), Idx(DI, vreg(t, "R15"), 4, 0, 64)}, "62b17e487f1cbf"},
// VMOVDQU64 — the W1 qword variant. // VMOVDQU64; the W1 qword variant.
{"VMOVDQU64 (SI)(R15*4),Z3", "VMOVDQU64", []Operand{Idx(SI, vreg(t, "R15"), 4, 0, 64), vreg(t, "Z3")}, "62b1fe486f1cbe"}, {"VMOVDQU64 (SI)(R15*4),Z3", "VMOVDQU64", []Operand{Idx(SI, vreg(t, "R15"), 4, 0, 64), vreg(t, "Z3")}, "62b1fe486f1cbe"},
{"VMOVDQU64 Z0,4(SI)(AX*1)", "VMOVDQU64", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f1fe487f840604000000"}, {"VMOVDQU64 Z0,4(SI)(AX*1)", "VMOVDQU64", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f1fe487f840604000000"},
{"VMOVDQU64 Z1,Z2", "VMOVDQU64", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1fe487fca"}, {"VMOVDQU64 Z1,Z2", "VMOVDQU64", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1fe487fca"},
@@ -77,7 +77,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VPSHUFB Z1,Z2,Z3", "VPSHUFB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d4800d9"}, {"VPSHUFB Z1,Z2,Z3", "VPSHUFB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d4800d9"},
{"VMOVDQU8 Z1,Z2", "VMOVDQU8", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f17f487fca"}, {"VMOVDQU8 Z1,Z2", "VMOVDQU8", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f17f487fca"},
{"VMOVDQU16 Z1,Z2", "VMOVDQU16", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff487fca"}, {"VMOVDQU16 Z1,Z2", "VMOVDQU16", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff487fca"},
// Indices 16–31: rm[4] rides in X̄ for register operands. // Indices 16-31: rm[4] rides in X̄ for register operands.
{"VPSHUFD $1,X16,X17", "VPSHUFD", []Operand{Imm(1), vreg(t, "X16"), vreg(t, "X17")}, "62a17d0870c801"}, {"VPSHUFD $1,X16,X17", "VPSHUFD", []Operand{Imm(1), vreg(t, "X16"), vreg(t, "X17")}, "62a17d0870c801"},
{"VMOVUPD (DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 0, 64), vreg(t, "Z14")}, "6271fd481037"}, {"VMOVUPD (DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 0, 64), vreg(t, "Z14")}, "6271fd481037"},
{"VMOVUPD 64(DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 64, 64), vreg(t, "Z14")}, "6271fd48107701"}, {"VMOVUPD 64(DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 64, 64), vreg(t, "Z14")}, "6271fd48107701"},
@@ -96,7 +96,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VPBROADCASTD 4(SI),Z10", "VPBROADCASTD", []Operand{Ptr(SI, 4, 4), vreg(t, "Z10")}, "62727d48585601"}, {"VPBROADCASTD 4(SI),Z10", "VPBROADCASTD", []Operand{Ptr(SI, 4, 4), vreg(t, "Z10")}, "62727d48585601"},
{"VPBROADCASTQ R8,X31", "VPBROADCASTQ", []Operand{vreg(t, "R8"), vreg(t, "X31")}, "6242fd087cf8"}, {"VPBROADCASTQ R8,X31", "VPBROADCASTQ", []Operand{vreg(t, "R8"), vreg(t, "X31")}, "6242fd087cf8"},
{"VPBROADCASTQ AX,Z9", "VPBROADCASTQ", []Operand{AX, vreg(t, "Z9")}, "6272fd487cc8"}, {"VPBROADCASTQ AX,Z9", "VPBROADCASTQ", []Operand{AX, vreg(t, "Z9")}, "6272fd487cc8"},
// Register indices 16–31 exist only in EVEX encodings. // Register indices 16-31 exist only in EVEX encodings.
{"VPBROADCASTD AX,Y30", "VPBROADCASTD", []Operand{AX, vreg(t, "Y30")}, "62627d287cf0"}, {"VPBROADCASTD AX,Y30", "VPBROADCASTD", []Operand{AX, vreg(t, "Y30")}, "62627d287cf0"},
// Packed double arithmetic / unpack (EVEX forms carry W=1). // Packed double arithmetic / unpack (EVEX forms carry W=1).
{"VSUBPD Z1,Z2,Z3", "VSUBPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed485cd9"}, {"VSUBPD Z1,Z2,Z3", "VSUBPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed485cd9"},
@@ -107,7 +107,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VUNPCKHPD Z1,Z2,Z3", "VUNPCKHPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed4815d9"}, {"VUNPCKHPD Z1,Z2,Z3", "VUNPCKHPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed4815d9"},
{"VSUBPD 64(AX),Z1,Z2", "VSUBPD", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5485c5001"}, {"VSUBPD 64(AX),Z1,Z2", "VSUBPD", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5485c5001"},
{"VSUBPD Z17,Z18,Z19", "VSUBPD", []Operand{vreg(t, "Z17"), vreg(t, "Z18"), vreg(t, "Z19")}, "62a1ed405cd9"}, {"VSUBPD Z17,Z18,Z19", "VSUBPD", []Operand{vreg(t, "Z17"), vreg(t, "Z18"), vreg(t, "Z19")}, "62a1ed405cd9"},
// VMOVDDUP — duplicate the low double; disp8×N = 64 at 512 bits, and // VMOVDDUP; duplicate the low double; disp8×N = 64 at 512 bits, and
// X16/X17 force EVEX (the mod=11 rm[4] extension rides in X̄). // X16/X17 force EVEX (the mod=11 rm[4] extension rides in X̄).
{"VMOVDDUP Z1,Z2", "VMOVDDUP", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff4812d1"}, {"VMOVDDUP Z1,Z2", "VMOVDDUP", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff4812d1"},
{"VMOVDDUP 64(AX),Z1", "VMOVDDUP", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1")}, "62f1ff48124801"}, {"VMOVDDUP 64(AX),Z1", "VMOVDDUP", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1")}, "62f1ff48124801"},
@@ -150,7 +150,7 @@ func TestEvexGroundTruth(t *testing.T) {
} }
} }
// TestEvexMasking checks the AVX-512 mask operand (K1–K7, placed freely among // TestEvexMasking checks the AVX-512 mask operand (K1-K7, placed freely among
// the operands) and the .Z zeroing suffix, byte for byte against the Go // the operands) and the .Z zeroing suffix, byte for byte against the Go
// assembler. // assembler.
func TestEvexMasking(t *testing.T) { func TestEvexMasking(t *testing.T) {
@@ -241,11 +241,11 @@ func TestEvexMasking(t *testing.T) {
} }
} }
// TestEvexExtendedGroundTruth covers the wider EVEX/AVX-512 set — ternary // TestEvexExtendedGroundTruth covers the wider EVEX/AVX-512 set; ternary
// logic, lane shuffles/inserts/extracts, compares with a K destination, // logic, lane shuffles/inserts/extracts, compares with a K destination,
// permutes, the wider integer families, expand/compress, broadcasts, // permutes, the wider integer families, expand/compress, broadcasts,
// rotates and word shifts, the opmask instructions, the EVEX suffixes // rotates and word shifts, the opmask instructions, the EVEX suffixes
// (rounding/SAE/broadcast) and the aligned/scalar moves — byte for byte // (rounding/SAE/broadcast) and the aligned/scalar moves; byte for byte
// against the Go assembler. // against the Go assembler.
func TestEvexExtendedGroundTruth(t *testing.T) { func TestEvexExtendedGroundTruth(t *testing.T) {
mem64 := func(base Reg) Operand { return Ptr(base, 0, 64) } mem64 := func(base Reg) Operand { return Ptr(base, 0, 64) }
@@ -275,7 +275,7 @@ func TestEvexExtendedGroundTruth(t *testing.T) {
{"VMULPD.RZ_SAE.Z", "VMULPD.RZ_SAE.Z", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f1edf959d9"}, {"VMULPD.RZ_SAE.Z", "VMULPD.RZ_SAE.Z", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f1edf959d9"},
{"VMAXPD.SAE", "VMAXPD.SAE", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed585fd9"}, {"VMAXPD.SAE", "VMAXPD.SAE", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed585fd9"},
{"VADDPD.BCST", "VADDPD.BCST", []Operand{mem64(AX), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5585810"}, {"VADDPD.BCST", "VADDPD.BCST", []Operand{mem64(AX), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5585810"},
// Packed single arithmetic (same opcodes, no mandatory prefix) — // Packed single arithmetic (same opcodes, no mandatory prefix);
// ZMM, YMM and XMM widths, rounding and broadcast. // ZMM, YMM and XMM widths, rounding and broadcast.
{"VADDPS", "VADDPS", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f16c4858d9"}, {"VADDPS", "VADDPS", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f16c4858d9"},
{"VMULPS", "VMULPS", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ec59d9"}, {"VMULPS", "VMULPS", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ec59d9"},
@@ -399,8 +399,8 @@ func TestEvexExtendedGroundTruth(t *testing.T) {
} }
// TestEvexHelperGroundTruth covers the floating-point helper and conversion // TestEvexHelperGroundTruth covers the floating-point helper and conversion
// tail of the EVEX set — reciprocals, rsqrt, getexp/getmant, scalef, // tail of the EVEX set; reciprocals, rsqrt, getexp/getmant, scalef,
// rndscale, reduce, fixupimm, range, fpclass, the remaining conversions — // rndscale, reduce, fixupimm, range, fpclass, the remaining conversions;
// plus gather/scatter with VSIB addressing, byte for byte against the Go // plus gather/scatter with VSIB addressing, byte for byte against the Go
// assembler. // assembler.
func TestEvexHelperGroundTruth(t *testing.T) { func TestEvexHelperGroundTruth(t *testing.T) {
@@ -502,9 +502,9 @@ func TestEvexHelperGroundTruth(t *testing.T) {
} }
// TestEvexGprGroundTruth covers the scalar conversions between vector and // TestEvexGprGroundTruth covers the scalar conversions between vector and
// general-purpose registers — the signed and truncated VCVT{,T}S{D,S}2SI // general-purpose registers; the signed and truncated VCVT{,T}S{D,S}2SI
// forms (VEX and EVEX), the unsigned EVEX-only forms, and the GPR-to-vector // forms (VEX and EVEX), the unsigned EVEX-only forms, and the GPR-to-vector
// VCVTSI2*/VCVTUSI2* forms with the preserved vector source in vvvv — byte // VCVTSI2*/VCVTUSI2* forms with the preserved vector source in vvvv; byte
// for byte against the Go assembler, including memory sources and extended // for byte against the Go assembler, including memory sources and extended
// GPRs. // GPRs.
func TestEvexGprGroundTruth(t *testing.T) { func TestEvexGprGroundTruth(t *testing.T) {
@@ -675,6 +675,15 @@ func TestEvexErrors(t *testing.T) {
{"align arity", "VALIGND", []Operand{Imm(1), vreg(t, "Z0"), vreg(t, "Z1")}}, {"align arity", "VALIGND", []Operand{Imm(1), vreg(t, "Z0"), vreg(t, "Z1")}},
// VEX-only mnemonics reject registers only EVEX can encode. // VEX-only mnemonics reject registers only EVEX can encode.
{"VMOVMSKPS X16", "VMOVMSKPS", []Operand{vreg(t, "X16"), AX}}, {"VMOVMSKPS X16", "VMOVMSKPS", []Operand{vreg(t, "X16"), AX}},
// The scalar EVEX move matches its VEX twin and the Go assembler:
// XMM↔memory only, never reg-reg and never a wider register (the
// toolchain rejects every one of these shapes).
{"VMOVSS X1,X2", "VMOVSS", []Operand{vreg(t, "X1"), vreg(t, "X2")}},
{"VMOVSS X16,X2", "VMOVSS", []Operand{vreg(t, "X16"), vreg(t, "X2")}},
{"VMOVSS Y1,(AX)", "VMOVSS", []Operand{vreg(t, "Y1"), Ptr(AX, 0, 4)}},
{"VMOVSS Z1,Z2", "VMOVSS", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}},
{"VMOVSS Z1,(AX)", "VMOVSS", []Operand{vreg(t, "Z1"), Ptr(AX, 0, 4)}},
{"VMOVSS (AX),Z2", "VMOVSS", []Operand{Ptr(AX, 0, 4), vreg(t, "Z2")}},
} }
for _, c := range cases { for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil { if _, err := Encode(c.mnem, c.ops...); err == nil {
+104 -21
View File
@@ -14,8 +14,8 @@ import (
"sync" "sync"
) )
// This file emits GOOBJ — the Go toolchain's object format, which cmd/link // This file emits GOOBJ, the Go toolchain's object format, which cmd/link
// consumes directly — so gasm-assembled functions drop into a go build // consumes directly, so gasm-assembled functions drop into a go build
// without the Go assembler. The layout follows cmd/internal/goobj: a // without the Go assembler. The layout follows cmd/internal/goobj: a
// toolchain preamble ("go object ...\n!\n"), the go120ld header with its // toolchain preamble ("go object ...\n!\n"), the go120ld header with its
// block offsets, a string table, symbol definitions, the relocation / // block offsets, a string table, symbol definitions, the relocation /
@@ -68,11 +68,12 @@ const (
kindSDWARFLINES = 20 kindSDWARFLINES = 20
) )
// Symbol flags (cmd/internal/goobj). // Symbol flags (cmd/internal/goobj). The linkname flag is set only for
// //go:linkname symbols (and main.main); ordinary assembly symbols carry
// none, matching cmd/asm's output.
const ( const (
symFlagDupok = 0x01 symFlagDupok = 0x01
symFlagNoSplit = 0x10 symFlagNoSplit = 0x10
symFlag2Link = 0x10 // asm objects flag every named symbol as linkname
symABIStatic = 0xffff symABIStatic = 0xffff
) )
@@ -94,10 +95,12 @@ const (
) )
// Relocation types (cmd/internal/objabi). // Relocation types (cmd/internal/objabi).
// R_PCREL and R_ADDR are stable across Go versions. // R_ADDR, R_CALL, R_PCREL and R_TLS_LE are stable across Go versions.
const ( const (
relocPCRel = 14 // R_PCREL
relocAddr = 1 // R_ADDR relocAddr = 1 // R_ADDR
relocCall = 7 // R_CALL
relocPCRel = 14 // R_PCREL
relocTLSLE = 15 // R_TLS_LE
) )
// relocDWTXTADDRU4 returns the R_DWTXTADDR_U4 relocation type for the // relocDWTXTADDRU4 returns the R_DWTXTADDR_U4 relocation type for the
@@ -140,10 +143,30 @@ func isGo127OrLater() bool {
// Special package indices for symbol references. // Special package indices for symbol references.
const ( const (
pkgIdxNone = 0x7fffffff pkgIdxNone = 0x7fffffff
pkgIdxSelf = 0x7ffffffb pkgIdxSelf = 0x7ffffffb
pkgIdxBuiltin = 0x7ffffffc
) )
// goobjBuiltinMorestackNoctxt is the index of runtime.morestack_noctxt in
// cmd/internal/goobj/builtinlist.go of the toolchain the object targets
// (246 since Go 1.25; the list is append-only).
const goobjBuiltinMorestackNoctxt = 246
// goobjBuiltinMorestack is the builtin reference the toolchain emits for the
// stack-guard call.
var goobjBuiltinMorestack = "runtime\u00b7morestack_noctxt"
// isCallReloc reports whether k is one of the per-arch call relocations a
// direct branch to a TEXT symbol carries.
func isCallReloc(k RelocKind) bool {
switch k {
case RelCall, RelRISCVJal, RelArm64Branch, RelLoong64Branch:
return true
}
return false
}
const goobjMagic = "\x00go120ld" const goobjMagic = "\x00go120ld"
// goSym is one symbol definition under construction. // goSym is one symbol definition under construction.
@@ -178,14 +201,24 @@ type dwarfRelocSet struct {
// does with its -p flag). srcPath names the source file recorded in the // does with its -p flag). srcPath names the source file recorded in the
// object's file table and line tables. The toolchain's object preamble is // object's file table and line tables. The toolchain's object preamble is
// captured from the installed go tool asm, so the output links with the // captured from the installed go tool asm, so the output links with the
// toolchain it was produced on — exactly like a real assembly object. // toolchain it was produced on, exactly like a real assembly object.
func (img *Image) GOObject(pkgPath, srcPath string) ([]byte, error) { func (img *Image) GOObject(pkgPath, srcPath string) ([]byte, error) {
pre, err := toolchainObjectPreamble() pre, err := toolchainObjectPreamble()
if err != nil { if err != nil {
return nil, err return nil, err
} }
// amd64: MinLC 1, R_PCREL for the code relocations. // amd64: MinLC 1, R_PCREL for displacements, R_CALL for calls and
return img.emitGOObject(pkgPath, srcPath, pre, 1, func(Reloc) (uint16, uint8) { return relocPCRel, 4 }) // R_TLS_LE for the stack-guard TLS load.
return img.emitGOObject(pkgPath, srcPath, pre, 1, func(r Reloc) (uint16, uint8) {
switch r.Kind {
case RelCall:
return relocCall, 4
case RelTLSLE:
return relocTLSLE, 4
default:
return relocPCRel, 4
}
})
} }
// emitGOObject assembles the GOOBJ payload for any architecture. pre is // emitGOObject assembles the GOOBJ payload for any architecture. pre is
@@ -198,7 +231,7 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
return nil, fmt.Errorf("GOOBJ emission requires a package path (-p)") return nil, fmt.Errorf("GOOBJ emission requires a package path (-p)")
} }
// The non-package definitions first — the DWARF symbols reference the // The non-package definitions first, the DWARF symbols reference the
// functions by these indices: per function the four pc-value tables // functions by these indices: per function the four pc-value tables
// and the function itself, as cmd/asm lays them out. // and the function itself, as cmd/asm lays them out.
type npSym struct { type npSym struct {
@@ -247,7 +280,7 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
// their relocations cover whole AUIPC/pcalau12i pairs, so // their relocations cover whole AUIPC/pcalau12i pairs, so
// zeroing r.Off would erase the opcode/register bits the linker // zeroing r.Off would erase the opcode/register bits the linker
// preserves when it patches only the immediate. // preserves when it patches only the immediate.
if r.Kind != RelPCRel32 { if r.Kind != RelPCRel32 && r.Kind != RelCall {
continue continue
} }
if r.Off >= 0 && r.Off+4 <= len(code) { if r.Off >= 0 && r.Off+4 <= len(code) {
@@ -255,7 +288,7 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
} }
} }
nps = append(nps, npSym{ nps = append(nps, npSym{
sym: goSym{name: name, abi: abi, typ: kindSTEXT, flag: flag, flag2: symFlag2Link, size: uint32(fn.Size)}, sym: goSym{name: name, abi: abi, typ: kindSTEXT, flag: flag, size: uint32(fn.Size)},
data: code, data: code,
}) })
} }
@@ -285,7 +318,7 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
abi = symABIStatic abi = symABIStatic
} }
defIdx[d.Name] = len(defs) defIdx[d.Name] = len(defs)
defs = append(defs, goSym{name: name, abi: abi, typ: typ, flag: flag, flag2: symFlag2Link, size: uint32(d.Size)}) defs = append(defs, goSym{name: name, abi: abi, typ: typ, flag: flag, size: uint32(d.Size)})
defData = append(defData, img.Data[d.Offset:d.Offset+d.Size]) defData = append(defData, img.Data[d.Offset:d.Offset+d.Size])
} }
fnFiIdx := make([]int, len(img.Funcs)) fnFiIdx := make([]int, len(img.Funcs))
@@ -320,6 +353,13 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
) )
} }
// Index the non-package TEXT definitions by short name for the internal
// call references.
textNpIdx := map[string]int{}
for i, fn := range img.Funcs {
textNpIdx[fn.Name] = fnNpIdx[i]
}
// Resolve external symbol references (cross-package). Build the // Resolve external symbol references (cross-package). Build the
// package index table and determine each external symbol's SymIdx // package index table and determine each external symbol's SymIdx
// by reading the target package's export data. // by reading the target package's export data.
@@ -327,10 +367,20 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
var extPkgIdx map[string]int var extPkgIdx map[string]int
var extSymIdx map[string]int var extSymIdx map[string]int
if len(img.Externals) > 0 { if len(img.Externals) > 0 {
var err error // The morestack call is a builtin reference, not a resolved external.
extPkgTable, extPkgIdx, extSymIdx, err = resolveExternalSymbols(img.Externals) var need []string
if err != nil { for _, n := range img.Externals {
return nil, fmt.Errorf("GOOBJ emission: resolving external symbols: %w", err) if n == goobjBuiltinMorestack {
continue
}
need = append(need, n)
}
if len(need) > 0 {
var err error
extPkgTable, extPkgIdx, extSymIdx, err = resolveExternalSymbols(need)
if err != nil {
return nil, fmt.Errorf("GOOBJ emission: resolving external symbols: %w", err)
}
} }
} }
@@ -342,6 +392,31 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
si := len(defs) + fnNpIdx[i] si := len(defs) + fnNpIdx[i]
for _, r := range fn.Relocs { for _, r := range fn.Relocs {
typ, size := relocField(r) typ, size := relocField(r)
if r.Kind == RelTLSLE {
// The TLS load has no symbol: {0, 0} is the nil ref.
var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = size
binary.LittleEndian.PutUint16(rec[5:], typ)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
binary.LittleEndian.PutUint32(rec[15:], 0)
binary.LittleEndian.PutUint32(rec[19:], 0)
symRelocs[si] = append(symRelocs[si], rec[:]...)
continue
}
if r.External && r.Name == goobjBuiltinMorestack {
// The stack-guard morestack call uses the toolchain's
// builtin reference.
var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = size
binary.LittleEndian.PutUint16(rec[5:], typ)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
binary.LittleEndian.PutUint32(rec[15:], pkgIdxBuiltin)
binary.LittleEndian.PutUint32(rec[19:], goobjBuiltinMorestackNoctxt)
symRelocs[si] = append(symRelocs[si], rec[:]...)
continue
}
if r.External { if r.External {
// Split package-qualified name: "runtime·morestack" → runtime, morestack. // Split package-qualified name: "runtime·morestack" → runtime, morestack.
pkg, name := splitQualified(r.Name) pkg, name := splitQualified(r.Name)
@@ -366,16 +441,24 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
symRelocs[si] = append(symRelocs[si], rec[:]...) symRelocs[si] = append(symRelocs[si], rec[:]...)
continue continue
} }
pkg := uint32(pkgIdxSelf)
di, ok := defIdx[r.Name] di, ok := defIdx[r.Name]
if !ok { if !ok {
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name) // A call to a TEXT function of the same file references the
// non-package definition table.
ni, isText := textNpIdx[r.Name]
if !isText || !isCallReloc(r.Kind) {
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name)
}
pkg = pkgIdxNone
di = ni
} }
var rec [23]byte var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off))) binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = size // field width rec[4] = size // field width
binary.LittleEndian.PutUint16(rec[5:], typ) binary.LittleEndian.PutUint16(rec[5:], typ)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend)) binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
binary.LittleEndian.PutUint32(rec[15:], pkgIdxSelf) binary.LittleEndian.PutUint32(rec[15:], pkg)
binary.LittleEndian.PutUint32(rec[19:], uint32(di)) binary.LittleEndian.PutUint32(rec[19:], uint32(di))
symRelocs[si] = append(symRelocs[si], rec[:]...) symRelocs[si] = append(symRelocs[si], rec[:]...)
} }
+73 -33
View File
@@ -40,8 +40,11 @@ func exportPath(importPath string) (string, error) {
// //
// refs maps package import paths to the symbol names referenced from that // refs maps package import paths to the symbol names referenced from that
// package. The returned pkgIdx maps each import path to its position in // package. The returned pkgIdx maps each import path to its position in
// the blkPkgIdx table (0-based), and symIdx gives each symbol's index within // the blkPkgIdx table, which reserves index 0 for the dummy invalid
// its package. // package (cmd/internal/obj/sym.go: "0 is invalid index"; the loader's
// reader loop starts at 1), so package i sits at block index i+1 and its
// relocations carry i+1. symIdx gives each symbol's index within its
// package.
func resolveExternalGOOBJ(refs map[string][]string) (pkgIdx map[string]int, symIdx map[string]int, err error) { func resolveExternalGOOBJ(refs map[string][]string) (pkgIdx map[string]int, symIdx map[string]int, err error) {
pkgIdx = make(map[string]int, len(refs)) pkgIdx = make(map[string]int, len(refs))
symIdx = make(map[string]int) symIdx = make(map[string]int)
@@ -50,7 +53,9 @@ func resolveExternalGOOBJ(refs map[string][]string) (pkgIdx map[string]int, symI
packages := sortedPkgRefs(refs) packages := sortedPkgRefs(refs)
for i, pkg := range packages { for i, pkg := range packages {
pkgIdx[pkg.path] = i // Block index 0 is the dummy invalid package; the first real
// package starts at 1.
pkgIdx[pkg.path] = i + 1
exp, err := exportPath(pkg.path) exp, err := exportPath(pkg.path)
if err != nil { if err != nil {
return nil, nil, err return nil, nil, err
@@ -84,7 +89,7 @@ func sortedPkgRefs(refs map[string][]string) []pkgRef {
for pkg, syms := range refs { for pkg, syms := range refs {
pkgs = append(pkgs, pkgRef{pkg, syms}) pkgs = append(pkgs, pkgRef{pkg, syms})
} }
// Simple insertion sort — the list is tiny (usually 1–3 packages). // Simple insertion sort, the list is tiny (usually 1-3 packages).
for i := 1; i < len(pkgs); i++ { for i := 1; i < len(pkgs); i++ {
for j := i; j > 0 && pkgs[j-1].path > pkgs[j].path; j-- { for j := i; j > 0 && pkgs[j-1].path > pkgs[j].path; j-- {
pkgs[j-1], pkgs[j] = pkgs[j], pkgs[j-1] pkgs[j-1], pkgs[j] = pkgs[j], pkgs[j-1]
@@ -145,40 +150,65 @@ func parseArDecimal(b []byte) int {
} }
// goobjFile is a parsed GOOBJ file: the string table and the symbol-definition // goobjFile is a parsed GOOBJ file: the string table and the symbol-definition
// block. // blocks. The hashed blocks are kept raw: their symbols carry no names, only
// the loader needs their counts.
type goobjFile struct { type goobjFile struct {
strTab []byte // string table, at headerSize + n strTab []byte // string table, at headerSize + n
symdef []byte // blkSymdef raw block symdef []byte // blkSymdef raw block
npdef []byte // blkNonpkgdef raw block hashed64 []byte // blkHashed64def raw block
hashed []byte // blkHasheddef raw block
npdef []byte // blkNonpkgdef raw block
} }
// symbols returns all symbol names in definition order by scanning the // loaderIndexBase returns the index the first nonpkgdef symbol occupies in the
// symdef and nonpkgdef blocks and resolving each name through the string // loader's per-object symbol array. cmd/link lays the definition blocks out as
// table. Package definitions (blkSymdef) use fully-qualified names like // symdef, hashed64def, hasheddef, nonpkgdef, nonpkgref (loader.go: preloadSyms
// "runtime.morestack"; non-package definitions (blkNonpkgdef) use bare // fills r.syms in exactly that order, and resolve() indexes PkgIdxNone and
// names like "morestack". This combined list matches the index the // cross-package SymIdx into it), so a symbol found in blkNonpkgdef carries the
// linker expects for cross-package references. // three leading blocks' symbol counts as its base.
func (f *goobjFile) loaderIndexBase() int {
return len(f.symdef)/recSymSize + len(f.hashed64)/recSymSize + len(f.hashed)/recSymSize
}
// symbols returns the names of the symdef and nonpkgdef blocks in
// definition order. Package definitions (blkSymdef) use fully-qualified
// names like "runtime.morestack"; non-package definitions (blkNonpkgdef)
// use bare names like "morestack". For lookups by index prefer
// findSymbol: it adds the hashed blocks' count the loader's array
// interleaves between the two.
func (f *goobjFile) symbols() []string { func (f *goobjFile) symbols() []string {
return append(f.defNames(), f.npdefNames()...) return append(f.defNames(), f.npdefNames()...)
} }
// findSymbol returns the index of a symbol within the combined symbol list, // findSymbol returns the index of a symbol within the loader's per-object
// or -1 if not found. It first tries the fully-qualified name (pkg.name), // symbol array, or -1 if not found. It first tries the fully-qualified
// then the bare name. // name (pkg.name), then the bare name (assembly objects store dotless
// names, e.g. runtime's "gogo", for symbols other packages reach through
// a linkname).
func (f *goobjFile) findSymbol(pkg, name string) int { func (f *goobjFile) findSymbol(pkg, name string) int {
base := f.loaderIndexBase()
qualified := pkg + "." + name qualified := pkg + "." + name
syms := f.symbols() for i, s := range f.defNames() {
for i, s := range syms {
if s == qualified { if s == qualified {
return i return i
} }
} }
// Try bare name (for non-package definitions). for i, s := range f.npdefNames() {
for i, s := range syms { if s == qualified {
return base + i
}
}
// Try bare name (for dotless assembly definitions).
for i, s := range f.defNames() {
if s == name { if s == name {
return i return i
} }
} }
for i, s := range f.npdefNames() {
if s == name {
return base + i
}
}
return -1 return -1
} }
@@ -192,12 +222,16 @@ func (f *goobjFile) npdefNames() []string {
return f.readSymNames(f.npdef) return f.readSymNames(f.npdef)
} }
// recSymSize is the size of one Sym record in the definition blocks
// (goobj.SymSize: stringRefSize + 2 + 1 + 1 + 1 + 4 + 4).
const recSymSize = 21
// readSymNames reads symbol names from a symdef/nonpkgdef block. Each record // readSymNames reads symbol names from a symdef/nonpkgdef block. Each record
// is 21 bytes: nameLen (u32), nameOff (u32), abi (u16), typ, flag, flag2, // is 21 bytes: nameLen (u32), nameOff (u32), abi (u16), typ, flag, flag2,
// size (u32), align (u32). nameOff is an absolute offset into the string // size (u32), align (u32). nameOff is an absolute offset into the string
// table. // table.
func (f *goobjFile) readSymNames(block []byte) []string { func (f *goobjFile) readSymNames(block []byte) []string {
const recSize = 21 const recSize = recSymSize
if len(block) < recSize { if len(block) < recSize {
return nil return nil
} }
@@ -247,16 +281,18 @@ func parseGOOBJ(data []byte) (*goobjFile, error) {
// [16:20] flags // [16:20] flags
// [20:96] 19 × uint32 offsets // [20:96] 19 × uint32 offsets
var offs [blkEnd + 1]uint32 var offs [blkEnd + 1]uint32
for i := 0; i <= blkEnd; i++ { for i := range blkEnd + 1 {
offs[i] = binary.LittleEndian.Uint32(payload[20+4*i:]) offs[i] = binary.LittleEndian.Uint32(payload[20+4*i:])
} }
// The string table lives at headerSize. // The string table lives at headerSize.
strTabStart := uint32(goobjHeaderSize) strTabStart := uint32(goobjHeaderSize)
f := &goobjFile{ f := &goobjFile{
strTab: payload[strTabStart:offs[0]], strTab: payload[strTabStart:offs[0]],
symdef: blockSlice(payload, offs, blkSymdef, blkSymdef+1), symdef: blockSlice(payload, offs, blkSymdef, blkSymdef+1),
npdef: blockSlice(payload, offs, blkNonpkgdef, blkNonpkgdef+1), hashed64: blockSlice(payload, offs, blkHashed64def, blkHashed64def+1),
hashed: blockSlice(payload, offs, blkHasheddef, blkHasheddef+1),
npdef: blockSlice(payload, offs, blkNonpkgdef, blkNonpkgdef+1),
} }
return f, nil return f, nil
} }
@@ -306,22 +342,26 @@ func resolveExternalSymbols(externals []string) (pkgTable []string, pkgIdxMap ma
return nil, nil, nil, err return nil, nil, nil, err
} }
// Build the package table in pkgIdx order. // Build the package table in pkgIdx order. The indices are 1-based
// (0 is the dummy invalid package, written by the emitter itself), so
// the table without the dummy is indexed one below.
pkgTable = make([]string, len(pkgIdx1)) pkgTable = make([]string, len(pkgIdx1))
for pkg, idx := range pkgIdx1 { for pkg, idx := range pkgIdx1 {
pkgTable[idx] = pkg pkgTable[idx-1] = pkg
} }
return pkgTable, pkgIdx1, symIdx1, nil return pkgTable, pkgIdx1, symIdx1, nil
} }
// splitQualified splits a qualified Go symbol name (pkgpath·name) into its // splitQualified splits a qualified Go symbol name (pkgpath·name) into its
// package path and local name. The separator is the middle dot (U+00B7). // package path and local name. The separator is the middle dot (U+00B7),
// If no separator is found, the symbol is assumed to be in the current // whose UTF-8 encoding is two bytes, so the search must be string-based:
// package (empty pkg). // IndexByte would match only the second byte and leave the lead byte on
// the package path. If no separator is found, the symbol is assumed to be
// in the current package (empty pkg).
func splitQualified(full string) (pkg, name string) { func splitQualified(full string) (pkg, name string) {
if idx := strings.IndexByte(full, '\u00b7'); idx >= 0 { if before, after, ok := strings.Cut(full, "\u00b7"); ok {
return full[:idx], full[idx+len("\u00b7"):] return before, after
} }
if before, after, ok := strings.Cut(full, "."); ok { if before, after, ok := strings.Cut(full, "."); ok {
return before, after return before, after
+6 -2
View File
@@ -56,8 +56,12 @@ func TestResolveExternalSymbols(t *testing.T) {
if err != nil { if err != nil {
t.Fatalf("resolveExternalGOOBJ: %v", err) t.Fatalf("resolveExternalGOOBJ: %v", err)
} }
if len(pkgIdx) != 1 || pkgIdx["runtime"] != 0 { if len(pkgIdx) != 1 || pkgIdx["runtime"] != 1 {
t.Errorf("pkgIdx = %v, want runtime→0", pkgIdx) // Index 0 is the dummy invalid package in the blkPkgIdx table;
// the loader's reader loop starts at 1 (cmd/link/internal/
// loader/loader.go: "PkgIdx 0 is a dummy invalid package"), so
// the first real package must carry index 1.
t.Errorf("pkgIdx = %v, want runtime→1", pkgIdx)
} }
if _, ok := symIdx["runtime·g0"]; !ok { if _, ok := symIdx["runtime·g0"]; !ok {
t.Errorf("symIdx missing runtime·g0, got %v", symIdx) t.Errorf("symIdx missing runtime·g0, got %v", symIdx)
+192 -5
View File
@@ -117,7 +117,9 @@ DATA mask<>+8(SB)/8, $0x800f0e0d0c0b0a09
if len(defs) != 7 { if len(defs) != 7 {
t.Fatalf("symdefs = %d, want 7", len(defs)) t.Fatalf("symdefs = %d, want 7", len(defs))
} }
if defs[0].name != "mask" || defs[0].abi != 0xffff || defs[0].typ != kindSRODATA || defs[0].size != 16 || defs[0].flag2 != symFlag2Link { // The linkname flag stays clear: the toolchain sets it only for
// //go:linkname symbols, and an ordinary static GLOBL is not one.
if defs[0].name != "mask" || defs[0].abi != 0xffff || defs[0].typ != kindSRODATA || defs[0].size != 16 || defs[0].flag2 != 0 {
t.Errorf("mask symbol = %+v", defs[0]) t.Errorf("mask symbol = %+v", defs[0])
} }
if defs[1].name != "" || defs[1].typ != kindSDATA || defs[1].size != 28 { if defs[1].name != "" || defs[1].typ != kindSDATA || defs[1].size != 28 {
@@ -158,8 +160,8 @@ DATA mask<>+8(SB)/8, $0x800f0e0d0c0b0a09
t.Errorf("funcinfo bytes %x", fi) t.Errorf("funcinfo bytes %x", fi)
} }
// The pc-value tables of addq (non-package indices 0–3, so global // The pc-value tables of addq (non-package indices 0-3, so global
// indices 7–10): pcsp a flat zero over the whole function, pcinline a // indices 7-10): pcsp a flat zero over the whole function, pcinline a
// flat -1, both with the pc delta in MinLC (1) units. // flat -1, both with the pc delta in MinLC (1) units.
pcsp := data[le.Uint32(didx[4*7:]):] pcsp := data[le.Uint32(didx[4*7:]):]
if got := pcsp[:3]; !bytes.Equal(got, []byte{0x02, 19, 0x00}) { if got := pcsp[:3]; !bytes.Equal(got, []byte{0x02, 19, 0x00}) {
@@ -294,7 +296,7 @@ TEXT ·framed(SB), NOSPLIT, $8-0
} }
for i := range wantPCs { for i := range wantPCs {
if pcs[i] != wantPCs[i] || vals[i] != wantVals[i] { if pcs[i] != wantPCs[i] || vals[i] != wantVals[i] {
t.Errorf("pcsp[%d] = (%d,%d), want (%d,%d) — all: %v %v", i, pcs[i], vals[i], wantPCs[i], wantVals[i], pcs, vals) t.Errorf("pcsp[%d] = (%d,%d), want (%d,%d); all: %v %v", i, pcs[i], vals[i], wantPCs[i], wantVals[i], pcs, vals)
} }
} }
// The last two steps unwind the epilogue to zero. // The last two steps unwind the epilogue to zero.
@@ -333,7 +335,7 @@ TEXT ·useext(SB), NOSPLIT, $0-8
// TestGOObjectLinkAndRun is the end-to-end check: assemble the test // TestGOObjectLinkAndRun is the end-to-end check: assemble the test
// functions to a GOOBJ, swap it into a go build in place of the toolchain's // functions to a GOOBJ, swap it into a go build in place of the toolchain's
// assembly object, link, and run — the output must match the baseline // assembly object, link, and run; the output must match the baseline
// binary the Go assembler produced. Skipped when no Go toolchain is // binary the Go assembler produced. Skipped when no Go toolchain is
// available. // available.
func TestGOObjectLinkAndRun(t *testing.T) { func TestGOObjectLinkAndRun(t *testing.T) {
@@ -511,3 +513,188 @@ func fieldAfter(line, flag string) string {
} }
return "" return ""
} }
// TestGOObjectExternalPackageLink is the cross-package end-to-end check: a
// GOOBJ whose code references a real external package symbol (runtime's
// morestack, a plain reference rather than the builtin noctxt form) must
// carry a package index that points past the blkPkgIdx table's dummy entry
// 0, and the object must link against the real runtime. Pre-fix, the
// relocations carried block index 0, which the loader never fills, so the
// reference resolved against whatever object was loaded first and the link
// failed. The binary is not run: morestack returns to the call site's
// stack check, which a hand-written caller has none of.
func TestGOObjectExternalPackageLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
const asmSrc = `
#include "textflag.h"
TEXT ·fn(SB), NOSPLIT, $0-0
CALL ·helper(SB)
RET
TEXT ·helper(SB), NOSPLIT, $0-0
RET
`
const mainSrc = `package main
func fn()
func helper()
func main() {
fn()
helper()
}
`
if err := os.WriteFile(filepath.Join(dir, "main_amd64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module extlink\n\ngo 1.27\n"), 0o644); err != nil {
t.Fatal(err)
}
// Capture the build the toolchain performs and re-run only its link
// step with our object swapped into the package archive, mirroring
// TestGOObjectLinkAndRun.
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj, pkgArch string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_amd64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "pack r") && strings.Contains(line, "_pkg_.a"):
pkgArch = strings.TrimSpace(strings.SplitN(line, "pack r", 2)[1])
pkgArch = strings.Fields(strings.SplitN(pkgArch, "#", 2)[0])[0]
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || pkgArch == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q)", work, asmObj)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
pkgArch = strings.ReplaceAll(pkgArch, "$WORK", work)
// Assemble the source with gasm, then retarget fn's internal call at
// a real external package symbol: the reloc's qualified name drives
// the export-data resolution the way a source-level runtime·sym(SB)
// reference would.
f, errs := parser.Parse("main_amd64.s", asmSrc)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
fn := &img.Funcs[0]
for i := range fn.Relocs {
fn.Relocs[i].Name = "runtime\u00b7morestack"
fn.Relocs[i].External = true
}
img.Externals = []string{"runtime\u00b7morestack"}
obj, err := img.GOObject("main", "main_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
// Structural check: the blkPkgIdx block reserves entry 0 for the
// dummy invalid package and places runtime at entry 1, and fn's call
// relocation carries PkgIdx 1.
v := openGoobj(t, obj)
pkgBlk := v.blk(blkPkgIdx)
if len(pkgBlk) != 2*8 {
t.Fatalf("blkPkgIdx = %d bytes, want two entries", len(pkgBlk))
}
le := binary.LittleEndian
strEntry := func(i int) string {
e := pkgBlk[i*8 : (i+1)*8]
return v.str(le.Uint32(e[4:]), le.Uint32(e[0:]))
}
if s := strEntry(0); s != "" {
t.Errorf("blkPkgIdx[0] = %q, want the dummy empty package", s)
}
if s := strEntry(1); s != "runtime" {
t.Errorf("blkPkgIdx[1] = %q, want runtime", s)
}
relocs := v.blk(blkReloc)
// fn is the last non-package symbol (two functions, four pc tables
// each); its one reloc is the final record.
fnRec := relocs[len(relocs)-23:]
if pIdx := le.Uint32(fnRec[15:]); pIdx != 1 {
t.Errorf("external reloc PkgIdx = %d, want 1 (runtime)", pIdx)
}
// Swap the object into the package archive and link with cmd/link;
// the link line consumes the archive, not the loose object file.
membersDir := filepath.Join(dir, "members")
if err := os.MkdirAll(membersDir, 0o755); err != nil {
t.Fatal(err)
}
extract := exec.Command(goBin, "tool", "pack", "x", pkgArch)
extract.Dir = membersDir
if out, err := extract.CombinedOutput(); err != nil {
t.Fatalf("pack x: %v\n%s", err, out)
}
member := filepath.Join(membersDir, filepath.Base(asmObj))
if err := os.Chmod(member, 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(member, obj, 0o644); err != nil {
t.Fatal(err)
}
listCmd := exec.Command(goBin, "tool", "pack", "t", pkgArch)
listOut, err := listCmd.CombinedOutput()
if err != nil {
t.Fatalf("pack t: %v\n%s", err, listOut)
}
newArch := filepath.Join(dir, "pkg.a")
args := []string{"tool", "pack", "c", newArch}
seen := map[string]bool{}
for m := range strings.FieldsSeq(string(listOut)) {
if seen[m] {
continue
}
seen[m] = true
if err := os.Chmod(filepath.Join(membersDir, m), 0o644); err != nil {
t.Fatal(err)
}
args = append(args, filepath.Join(membersDir, m))
}
pack := exec.Command(goBin, args...)
pack.Dir = membersDir
if out, err := pack.CombinedOutput(); err != nil {
t.Fatalf("pack c: %v\n%s", err, out)
}
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, pkgArch, newArch)
linkLine = strings.ReplaceAll(linkLine, filepath.Join(work, "b001", "exe", "a.out"), filepath.Join(dir, "prog2"))
linkCmd := exec.Command("sh", "-c", "cd "+dir+" && "+linkLine)
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
// The call must have resolved to the real runtime symbol.
dump, err := exec.Command(goBin, "tool", "objdump", "-s", "main.fn", filepath.Join(dir, "prog2")).CombinedOutput()
if err != nil {
t.Fatalf("objdump main.fn: %v\n%s", err, dump)
}
if !bytes.Contains(dump, []byte("runtime.morestack")) {
t.Errorf("main.fn does not call runtime.morestack:\n%s", dump)
}
}
+35 -9
View File
@@ -13,28 +13,54 @@ import (
) )
// GOObjectAARCH64 emits a GOOBJ object file for AArch64. The layout is // GOObjectAARCH64 emits a GOOBJ object file for AArch64. The layout is
// the shared one in goobj.go — the toolchain preamble, the go120ld header // the shared one in goobj.go, the toolchain preamble, the go120ld header
// with its block offsets, the string table, the symbol definitions and the // with its block offsets, the string table, the symbol definitions and the
// reloc/aux/data index arrays — with the arm64 preamble, the MinLC of 4 // reloc/aux/data index arrays, with the arm64 preamble, the MinLC of 4
// for the pc-value deltas, and R_ADDRARM64 relocation types for the // for the pc-value deltas, and the arm64 relocation types for the ADRP
// ADRP+ADD/LDR/STR address pairs. // pairs and BL calls.
//
// The toolchain records one relocation per ADRP pair: a single R_ADDRARM64
// or R_ARM64_PCREL_LDST64 of Siz 8 at the ADRP word, from which the linker
// patches both instructions of the pair (cmd/internal/obj/arm64/asm7.go,
// the ADRP cases: one AddRel with Off at the pair's pc and Siz 8). gasm's
// assembler records the ADRP+ADD form as two word relocs, so the second
// word's twin is dropped here before emission.
func (img *Image) GOObjectAARCH64(pkgPath, srcPath string) ([]byte, error) { func (img *Image) GOObjectAARCH64(pkgPath, srcPath string) ([]byte, error) {
pre, err := toolchainObjectPreambleAARCH64() pre, err := toolchainObjectPreambleAARCH64()
if err != nil { if err != nil {
return nil, err return nil, err
} }
return img.emitGOObject(pkgPath, srcPath, pre, 4, func(r Reloc) (uint16, uint8) { coalesced := *img
if r.Kind == RelArm64Branch { coalesced.Funcs = append([]FuncLayout(nil), img.Funcs...)
for i := range coalesced.Funcs {
rs := coalesced.Funcs[i].Relocs
var keep []Reloc
for j := 0; j < len(rs); j++ {
keep = append(keep, rs[j])
if rs[j].Kind == RelArm64Addr && j+1 < len(rs) &&
rs[j+1].Kind == RelArm64Addr && rs[j+1].Off == rs[j].Off+4 {
j++ // the ADD word's twin: the Siz-8 pair reloc covers it
}
}
coalesced.Funcs[i].Relocs = keep
}
return coalesced.emitGOObject(pkgPath, srcPath, pre, 4, func(r Reloc) (uint16, uint8) {
switch r.Kind {
case RelArm64Branch:
return relocArm64Branch, 4 return relocArm64Branch, 4
case RelArm64LDST64:
return relocArm64LDST64, 8
default:
return relocArm64Addr, 8
} }
return relocArm64Addr, 4
}) })
} }
// arm64 relocation types (cmd/internal/objabi). // arm64 relocation types (cmd/internal/objabi).
const ( const (
relocArm64Addr = 3 // R_ADDRARM64 — ADRP+ADD/LDR/STR pair relocArm64Addr = 3 // R_ADDRARM64, ADRP+ADD pair
relocArm64Branch = 9 // R_CALLARM64 — BL instruction relocArm64Branch = 9 // R_CALLARM64, BL instruction
relocArm64LDST64 = 40 // R_ARM64_PCREL_LDST64, ADRP+LDR/STR pair
) )
// toolchainObjectPreambleAARCH64 returns the "go object ...\n!\n" header // toolchainObjectPreambleAARCH64 returns the "go object ...\n!\n" header
+11 -5
View File
@@ -13,9 +13,9 @@ import (
) )
// GOObjectLOONG64 emits a GOOBJ object file for LoongArch. The layout is // GOObjectLOONG64 emits a GOOBJ object file for LoongArch. The layout is
// the shared one in goobj.go — the toolchain preamble, the go120ld header // the shared one in goobj.go, the toolchain preamble, the go120ld header
// with its block offsets, the string table, the symbol definitions and the // with its block offsets, the string table, the symbol definitions and the
// reloc/aux/data index arrays — with the loong64 preamble, the MinLC of 4 // reloc/aux/data index arrays, with the loong64 preamble, the MinLC of 4
// for the pc-value deltas, and R_LOONG64_ADDR_HI/LO relocation types for // for the pc-value deltas, and R_LOONG64_ADDR_HI/LO relocation types for
// the pcalau12i+addi.d address pairs. // the pcalau12i+addi.d address pairs.
func (img *Image) GOObjectLOONG64(pkgPath, srcPath string) ([]byte, error) { func (img *Image) GOObjectLOONG64(pkgPath, srcPath string) ([]byte, error) {
@@ -25,11 +25,16 @@ func (img *Image) GOObjectLOONG64(pkgPath, srcPath string) ([]byte, error) {
} }
return img.emitGOObject(pkgPath, srcPath, pre, 4, func(r Reloc) (uint16, uint8) { return img.emitGOObject(pkgPath, srcPath, pre, 4, func(r Reloc) (uint16, uint8) {
// A pcalau12i+addi.d pair: the high part carries // A pcalau12i+addi.d pair: the high part carries
// R_LOONG64_ADDR_HI, the low part R_LOONG64_ADDR_LO. // R_LOONG64_ADDR_HI, the low part R_LOONG64_ADDR_LO; the guard's
if r.Kind == RelLoong64AddrLo { // morestack call carries R_CALLLOONG64.
switch {
case r.Kind == RelLoong64AddrLo:
return relocLoong64AddrLo, 4 return relocLoong64AddrLo, 4
case r.Kind == RelLoong64Branch:
return relocCallLoong64, 4
default:
return relocLoong64AddrHi, 4
} }
return relocLoong64AddrHi, 4
}) })
} }
@@ -39,6 +44,7 @@ func (img *Image) GOObjectLOONG64(pkgPath, srcPath string) ([]byte, error) {
const ( const (
relocLoong64AddrHi = 77 // R_LOONG64_ADDR_HI relocLoong64AddrHi = 77 // R_LOONG64_ADDR_HI
relocLoong64AddrLo = 78 // R_LOONG64_ADDR_LO relocLoong64AddrLo = 78 // R_LOONG64_ADDR_LO
relocCallLoong64 = 84 // R_CALLLOONG64
) )
// toolchainObjectPreambleLOONG64 returns the "go object ...\n!\n" header // toolchainObjectPreambleLOONG64 returns the "go object ...\n!\n" header
+2 -2
View File
@@ -13,9 +13,9 @@ import (
) )
// GOObjectRISCV emits a GOOBJ object file for RISC-V. The layout is the // GOObjectRISCV emits a GOOBJ object file for RISC-V. The layout is the
// shared one in goobj.go — the toolchain preamble, the go120ld header with // shared one in goobj.go, the toolchain preamble, the go120ld header with
// its block offsets, the string table, the symbol definitions and the // its block offsets, the string table, the symbol definitions and the
// reloc/aux/data index arrays — with the RISC-V preamble, the MinLC of 2 for // reloc/aux/data index arrays, with the RISC-V preamble, the MinLC of 2 for
// the pc-value deltas, and the single R_RISCV_PCREL_ITYPE/STYPE relocation // the pc-value deltas, and the single R_RISCV_PCREL_ITYPE/STYPE relocation
// per AUIPC pair, matching `go tool asm`'s model (each pair is one 8-byte // per AUIPC pair, matching `go tool asm`'s model (each pair is one 8-byte
// relocation, not the ELF HI20/LO12 pair). // relocation, not the ELF HI20/LO12 pair).
+329
View File
@@ -0,0 +1,329 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"encoding/hex"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// The expected bytes are pinned from `go tool asm` output (Go 1.27, amd64,
// verified with go tool objdump): the stack-split guard classes, the morestack
// block and the auto-NOSPLIT leaf behaviour.
func TestStackGuardBytes(t *testing.T) {
for _, tt := range []struct {
name string
src string
want string
}{
{"leafsmall", "TEXT \u00b7leafsmall(SB), $16-0\n\tRET\n",
"554889e54883ec104883c4105dc3"},
{"leafmed", "TEXT \u00b7leafmed(SB), $256-0\n\tRET\n",
"644c8b3425000000004c8da42478ffffff4d3b66107614554889e54881ec000100004881c4000100005dc3e800000000ebce"},
{"leafbig", "TEXT \u00b7leafbig(SB), $8192-0\n\tRET\n",
"644c8b3425000000004989e44981ec881f0000721a4d3b66107614554889e54881ec002000004881c4002000005dc3e800000000ebca"},
// Class 2 with a body long enough that the underflow JB relaxes to
// rel32: its displacement must span the real 6-byte JB, else the
// branch lands 4 bytes past the morestack block, inside the CALL
// displacement field.
{"leafbiglong", "TEXT \u00b7leafbiglong(SB), $8192-0\n" + strings.Repeat("\tMOVQ AX, BX\n", 40) + "\tRET\n",
"644c8b3425000000004989e44981ec881f00000f82960000004d3b66100f868c000000554889e54881ec00200000" + strings.Repeat("4889c3", 40) + "4881c4002000005dc3e800000000e947ffffff"},
{"callsmall", "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n",
"644c8b342500000000493b66107613554889e54883ec10e8000000004883c4105dc3e800000000ebd7"},
{"nosplit", "TEXT \u00b7nosplit(SB), NOSPLIT, $16-0\n\tRET\n",
"554889e54883ec104883c4105dc3"},
} {
f, errs := parser.Parse("g_amd64.s", tt.src)
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.name, errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.name, err)
}
fn := img.Funcs[0]
// The toolchain's object leaves every relocation field zero for the
// linker, while the gasm image resolves file-internal references, so
// the comparison masks the patch sites the way verify's ground truth
// does.
code := append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
got := hex.EncodeToString(code)
if got != tt.want {
t.Errorf("%s:\n got %s\n want %s", tt.name, got, tt.want)
}
}
}
// TestStackGuardRelocs checks the guard's patch sites: the TLS slot and the
// morestack call.
func TestStackGuardRelocs(t *testing.T) {
f, errs := parser.Parse("g_amd64.s", "TEXT \u00b7f(SB), $256-0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
relocs := img.Funcs[0].Relocs
if len(relocs) != 2 {
t.Fatalf("relocs = %d, want 2", len(relocs))
}
tls, call := relocs[0], relocs[1]
if tls.Kind != RelTLSLE || tls.Off != 5 || tls.Name != "" || tls.External {
t.Errorf("tls reloc = %+v, want RelTLSLE at 5 with no symbol", tls)
}
if call.Kind != RelCall || call.Name != "runtime\u00b7morestack_noctxt" || !call.External {
t.Errorf("call reloc = %+v, want RelCall to runtime.morestack_noctxt", call)
}
}
// TestStackGuardGOObj emissions succeed with the guard's TLS and builtin
// references in play.
func TestStackGuardGOObj(t *testing.T) {
f, errs := parser.Parse("g_amd64.s", "TEXT \u00b7f(SB), $256-0\n\tCALL \u00b7helper(SB)\n\tRET\nTEXT \u00b7helper(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.GOObject("testpkg", "g_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
if !bytes.Contains(obj, []byte("go120ld")) {
t.Fatal("object lacks the GOOBJ magic")
}
}
// The arm64 stack-split guard, pinned from `go tool asm` (Go 1.27, arm64):
// the guard classes, the auto-NOSPLIT leaf behaviour and the morestack
// block. Relocation fields are masked: the toolchain's object leaves them
// zero for the linker, the gasm image resolves file-internal references.
func TestStackGuardBytesARM64(t *testing.T) {
for _, tt := range []struct {
name string
src string
want string
}{
{"leafsmall", "TEXT \u00b7leafsmall(SB), $16-0\n\tRET\n",
"fe0f1ef8fd831ff8fd2300d1fd630091ff830091c0035fd6"},
{"leafmed", "TEXT \u00b7leafmed(SB), $256-0\n\tRET\n",
"900b40f9f14302d13f0210eb09010054f44304d19dfa3fa99f020091fd2300d1fd230491ff430491c0035fd6e3031eaa00000000f3ffff17"},
{"leafbig", "TEXT \u00b7leafbig(SB), $8192-0\n\tRET\n",
"900b40f91bf283d2f1633beba30100543f0210eb690100541b0284d2f4633bcb9dfa3fa99f020091fd2300d11b0184d2fd633b8b1b0284d2ff633b8bc0035fd6e3031eaa00000000eeffff17"},
{"callsmall", "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n",
"900b40f9ff6330eb09010054fe0f1ef8fd831ff8fd2300d100000000fd835ff8fe0742f8c0035fd6e3031eaa00000000f4ffff17"},
{"nosplit", "TEXT \u00b7nosplit(SB), NOSPLIT, $16-0\n\tRET\n",
"fe0f1ef8fd831ff8fd2300d1fd630091ff830091c0035fd6"},
} {
f, errs := parser.Parse("g_arm64.s", tt.src)
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.name, errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.name, err)
}
fn := img.Funcs[0]
code := append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
got := hex.EncodeToString(code)
if got != tt.want {
t.Errorf("%s:\n got %s\n want %s", tt.name, got, tt.want)
}
}
}
// TestStackGuardBranchTargetsARM64 checks the class-2 guard's branch
// positions for a frame whose guard constant needs two MOV words: the
// displacements must be computed from byte offsets (8+4*ml and 16+4*ml), so
// both branches land on the morestack block rather than inside the body.
// The frame size makes the toolchain switch its own prologue decomposition,
// so the assertion is on the branch targets, not pinned bytes.
func TestStackGuardBranchTargetsARM64(t *testing.T) {
f, errs := parser.Parse("g_arm64.s", "TEXT \u00b7f(SB), $65664-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
fn := img.Funcs[0]
code := img.Code[fn.Offset : fn.Offset+fn.Size]
if len(code)%4 != 0 {
t.Fatalf("function size %d is not a word multiple", len(code))
}
// autosize = 65680, so the guard materialises 65552 = MOVZ+MOVK: ml = 2
// and the branches sit at bytes 16 and 24 of the guard prefix.
const morestackBlock = 12 // MOVD R30, R3; BL; B back
blockStart := len(code) - morestackBlock
check := func(name string, off int) {
t.Helper()
w := leWord(code[off:])
imm19 := int32(w>>5) & 0x7FFFF
if imm19&(1<<18) != 0 {
imm19 -= 1 << 19
}
if target := off + int(imm19)*4; target != blockStart {
t.Errorf("%s at byte %d targets byte %d, want the morestack block at %d", name, off, target, blockStart)
}
}
check("B.LO", 16)
check("B.LS", 24)
}
// The riscv64 stack-split guard, pinned from `go tool asm` (Go 1.27,
// riscv64): the morestack call sits between the guard and the body, and the
// guard branches forward over it. Relocation fields are masked.
func TestStackGuardBytesRISCV64(t *testing.T) {
for _, tt := range []struct {
name string
src string
want string
}{
{"leafsmall", "TEXT \u00b7leafsmall(SB), $16-0\n\tRET\n",
"03b30d0163662300000000006ff05fff233411fe211106e08260610167800000"},
{"leafmed", "TEXT \u00b7leafmed(SB), $256-0\n\tRET\n",
"03b30d01930381f763667300000000006ff01fff233c11ee130181ef06e082601301811067800000"},
{"leafbig", "TEXT \u00b7leafbig(SB), $8192-0\n\tRET\n",
"03b30d0189639b8383f863697100f97f9b8f8f07b303f10163667300000000006ff01ffef97f8a9f23bc1ffef97fe13f7e9106e08260896fa12f7e9167800000"},
{"frameless", "TEXT \u00b7frameless(SB), $0-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n",
"03b30d0163662300000000006ff05fff233c11fe611106e0000000008260210167800000"},
{"nosplit", "TEXT \u00b7nosplit(SB), NOSPLIT, $16-0\n\tRET\n",
"233411fe211106e08260610167800000"},
} {
f, errs := parser.Parse("g_riscv64.s", tt.src)
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.name, errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.name, err)
}
fn := img.Funcs[0]
code := append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
got := hex.EncodeToString(code)
if got != tt.want {
t.Errorf("%s:\n got %s\n want %s", tt.name, got, tt.want)
}
}
}
// The loong64 stack-split guard, pinned from `go tool asm` (Go 1.27,
// loong64): every guard class (including the medium class with the
// materialised constant and the big class with the ORI-less constants), the
// auto-NOSPLIT leaf behaviour, the large-frame R30 prologue/epilogue forms
// and the morestack block. Relocation fields are masked.
func TestStackGuardBytesLOONG64(t *testing.T) {
for _, tt := range []struct {
name string
src string
want string
}{
{"leafsmall", "TEXT \u00b7leafsmall(SB), $16-0\n\tRET\n",
"61a0ff2963a0ff026100c0296360c0022000004c"},
{"leafmed", "TEXT \u00b7leafmed(SB), $256-0\n\tRET\n",
"d442c02878e0fd0294e21200801a004061e0fb2963e0fb026100c0296320c4022000004c3f00150000000000ffd7ff53"},
{"nosplit", "TEXT \u00b7nosplit(SB), NOSPLIT, $16-0\n\tRET\n",
"61a0ff2963a0ff026100c0296360c0022000004c"},
// The LR store leaves the 12-bit store-offset range while the SP
// adjust immediate still fits, and the epilogue adjusts through a
// single ORI.
{"fit2048", "TEXT \u00b7fit2048(SB), $2040-0\n\tRET\n",
"d442c0287800e20294e21200802600401e000014de8f1000c103e0296300e0026100c0291e00a00363f810002000004c3f00150000000000ffcbff53"},
// Medium class at the materialisation boundary (off = 2048 still
// immediate, 2049+ goes through R30).
{"med2048off", "TEXT \u00b7med2048off(SB), $2168-0\n\tRET\n",
"d442c0287800e00294e21200802e0040feffff15de8f1000c103de29feffff15de039e0363f810006100c0291e00a20363f810002000004c3f00150000000000ffc3ff53"},
{"medmat", "TEXT \u00b7medmat(SB), $2176-0\n\tRET\n",
"d442c028feffff15dee39f0378f8100094e21200802e0040feffff15de8f1000c1e3dd29feffff15dee39d0363f810006100c0291e20a20363f810002000004c3f00150000000000ffbbff53"},
// Big class with the rounding-split store and the floor-split adjust.
{"leafbig", "TEXT \u00b7leafbig(SB), $8192-0\n\tRET\n",
"d442c0283e000014de23be0378f8120000470044deffff15dee3810378f8100094e2120080320040deffff15de8f1000c1e3ff29beffff15dee3bf0363f810006100c0295e000014de23800363f810002000004c3f00150000000000ffa7ff53"},
// Zero low 12 bits drop the ORI from the store, the adjust and the
// epilogue materialisation.
{"bigzero", "TEXT \u00b7bigzero(SB), $4088-0\n\tRET\n",
"d442c028feffff15de03820378f8100094e21200802a0040feffff15de8f1000c103c029feffff1563f810006100c0293e00001463f810002000004c3f00150000000000ffbfff53"},
// Big class whose first constant has a zero high part: a single ORI.
{"big3976", "TEXT \u00b7big3976(SB), $4096-0\n\tRET\n",
"d442c0281e20be0378f8120000470044feffff15dee3810378f8100094e2120080320040feffff15de8f1000c1e3ff29deffff15dee3bf0363f810006100c0293e000014de23800363f810002000004c3f00150000000000ffabff53"},
// Big class at a multiple of 4096: both guard constants lose their
// ORI word.
{"giantlo0", "TEXT \u00b7giantlo0(SB), $4216-0\n\tRET\n",
"d442c0283e00001478f8120000430044feffff1578f8100094e2120080320040feffff15de8f1000c103fe29deffff15de03be0363f810006100c0293e000014de03820363f810002000004c3f00150000000000ffafff53"},
// Non-leaf big frame: the body call plus the LR restore epilogue.
{"callbig", "TEXT \u00b7callbig(SB), $8192-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n",
"d442c0283e000014de23be0378f81200004f0044deffff15dee3810378f8100094e21200803a0040deffff15de8f1000c1e3ff29beffff15dee3bf0363f810006100c029000000006100c0285e000014de23800363f810002000004c3f00150000000000ff9fff53"},
} {
f, errs := parser.Parse("g_loong64.s", tt.src)
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.name, errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.name, err)
}
fn := img.Funcs[0]
code := append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
got := hex.EncodeToString(code)
if got != tt.want {
t.Errorf("%s:\n got %s\n want %s", tt.name, got, tt.want)
}
}
}
// TestStackGuardGOObjInternalCall checks that GOOBJ emission succeeds when a
// guarded function calls a TEXT symbol of the same file, for every arch's
// call relocation kind.
func TestStackGuardGOObjInternalCall(t *testing.T) {
for _, tt := range []struct {
src string
assemble func(*ast.File) (*Image, error)
}{
{"g_amd64.s", AssembleFile},
{"g_arm64.s", AssembleFileARM64},
{"g_riscv64.s", AssembleFileRISCV},
{"g_loong64.s", AssembleFileLOONG64},
} {
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.src, errs)
}
img, err := tt.assemble(f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.src, err)
}
if _, err := img.GOObject("testpkg", tt.src); err != nil {
t.Errorf("%s: GOObject: %v", tt.src, err)
}
}
}
+158 -48
View File
@@ -21,7 +21,7 @@ var aluOp = map[string]struct {
} }
// unaryOp maps INC/DEC/NEG/NOT to their /digit and base opcode. INC/DEC use // unaryOp maps INC/DEC/NEG/NOT to their /digit and base opcode. INC/DEC use
// the 0xFE/0xFF group (the short 0x40–0x4F forms are REX prefixes in 64-bit // the 0xFE/0xFF group (the short 0x40-0x4F forms are REX prefixes in 64-bit
// mode); NEG/NOT use the 0xF6/0xF7 group. // mode); NEG/NOT use the 0xF6/0xF7 group.
var unaryOp = map[string]struct { var unaryOp = map[string]struct {
digit int digit int
@@ -33,7 +33,7 @@ var unaryOp = map[string]struct {
"NEG": {3, 0xF7}, "NEG": {3, 0xF7},
} }
// shiftOp maps SHL/SHR/SAR to their /digit in the 0xC0/0xC1/0xD0–0xD3 group. // shiftOp maps SHL/SHR/SAR to their /digit in the 0xC0/0xC1/0xD0-0xD3 group.
var shiftOp = map[string]int{ var shiftOp = map[string]int{
"SHL": 4, "SHL": 4,
"SHR": 5, "SHR": 5,
@@ -50,7 +50,7 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
// Integer scalar XMM moves: MOVQ with an XMM operand is the SSE2 // Integer scalar XMM moves: MOVQ with an XMM operand is the SSE2
// packed-quadword move, NOT a GPR move: mem→xmm encodes as F3 0F 7E // packed-quadword move, NOT a GPR move: mem→xmm encodes as F3 0F 7E
// (reg = dst, no REX.W — the Go assembler's form), xmm→mem as // (reg = dst, no REX.W, the Go assembler's form), xmm→mem as
// 66 0F D6 (rm = xmm). Register forms against a GPR use the MOVD // 66 0F D6 (rm = xmm). Register forms against a GPR use the MOVD
// opcodes with REX.W instead: 66 REX.W 0F 6E (gpr→xmm) and // opcodes with REX.W instead: 66 REX.W 0F 6E (gpr→xmm) and
// 66 REX.W 0F 7E (xmm→gpr); the memory opcodes with a register r/m // 66 REX.W 0F 7E (xmm→gpr); the memory opcodes with a register r/m
@@ -103,7 +103,7 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
switch src := src.(type) { switch src := src.(type) {
case Reg: case Reg:
if dstIsReg { if dstIsReg {
// MOV r/m, r: 0x88/0x89, reg=src, rm=dst — the form the Go // MOV r/m, r: 0x88/0x89, reg=src, rm=dst, the form the Go
// assembler emits for register-to-register moves. // assembler emits for register-to-register moves.
i := newInstr(size, []byte{movRM(size)}) i := newInstr(size, []byte{movRM(size)})
if err := setRM(i, src, dst, size); err != nil { if err := setRM(i, src, dst, size); err != nil {
@@ -147,7 +147,7 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
// a signed int32, choosing per sign: // a signed int32, choosing per sign:
// v >= 0: B8+rd imm32 without REX.W (zero-extended by the // v >= 0: B8+rd imm32 without REX.W (zero-extended by the
// hardware, REX.B still emitted for R8-R15); // hardware, REX.B still emitted for R8-R15);
// v < 0: REX.W C7 /0 imm32 (sign-extended — the plain B8+rd // v < 0: REX.W C7 /0 imm32 (sign-extended, the plain B8+rd
// form would zero-extend and corrupt the value). // form would zero-extend and corrupt the value).
// Out-of-range immediates keep the B8+rd imm64 form. // Out-of-range immediates keep the B8+rd imm64 form.
if size == 8 && v >= 0 && v <= (1<<31)-1 { if size == 8 && v >= 0 && v <= (1<<31)-1 {
@@ -173,7 +173,11 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
if dstReg.needsREX(size) { if dstReg.needsREX(size) {
i.rexForced = true i.rexForced = true
} }
i.imm = immediate(v, size, true) imm, err := immediate(v, size, true)
if err != nil {
return err
}
i.imm = imm
return e.emit(i) return e.emit(i)
} }
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. // MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
@@ -185,7 +189,11 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
if err := setRMDigit(i, 0, dst, size); err != nil { if err := setRMDigit(i, 0, dst, size); err != nil {
return err return err
} }
i.imm = immediate(int64(src), size, false) imm, err := immediate(int64(src), size, false)
if err != nil {
return err
}
i.imm = imm
return e.emit(i) return e.emit(i)
} }
return fmt.Errorf("MOV: invalid operands") return fmt.Errorf("MOV: invalid operands")
@@ -226,8 +234,8 @@ func (e *enc) encodeALU(op struct {
return e.encodeALUImm(op.digit, dst, int64(imm), size) return e.encodeALUImm(op.digit, dst, int64(imm), size)
} }
// CMP accepts the immediate in the second position too — CMPL CX, $31 is // CMP accepts the immediate in the second position too, CMPL CX, $31 is
// the form the Go assembler itself accepts — and encodes it identically // the form the Go assembler itself accepts, and encodes it identically
// (CMP r/m, imm sets the flags as first − second). No other ALU op takes // (CMP r/m, imm sets the flags as first − second). No other ALU op takes
// an immediate destination. // an immediate destination.
if imm, ok := dst.(Imm); ok { if imm, ok := dst.(Imm); ok {
@@ -297,11 +305,15 @@ func (e *enc) encodeALU(op struct {
func (e *enc) encodeALUImm(digit int, dst Operand, imm int64, size int) error { func (e *enc) encodeALUImm(digit int, dst Operand, imm int64, size int) error {
if size == 1 { if size == 1 {
immBytes, err := immediate(imm, 1, false)
if err != nil {
return err
}
i := newInstr(1, []byte{0x80}) i := newInstr(1, []byte{0x80})
if err := setRMDigit(i, digit, dst, 1); err != nil { if err := setRMDigit(i, digit, dst, 1); err != nil {
return err return err
} }
i.imm = []byte{byte(int8(imm))} i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
if fits8(imm) { if fits8(imm) {
@@ -313,13 +325,17 @@ func (e *enc) encodeALUImm(digit int, dst Operand, imm int64, size int) error {
i.imm = []byte{byte(int8(imm))} i.imm = []byte{byte(int8(imm))}
return e.emit(i) return e.emit(i)
} }
// 0x81 /digit, imm16/imm32 — or the Go assembler's accumulator short // 0x81 /digit, imm16/imm32, or the Go assembler's accumulator short
// form (opcode+5, no ModR/M) when the destination is AX/AL, which it // form (opcode+5, no ModR/M) when the destination is AX/AL, which it
// prefers over the generic form exactly here. // prefers over the generic form exactly here.
if r, ok := dst.(Reg); ok && r.idx == 0 { if r, ok := dst.(Reg); ok && r.idx == 0 {
accOp := map[int]byte{0: 0x05, 1: 0x0D, 2: 0x15, 3: 0x1D, 4: 0x25, 5: 0x2D, 6: 0x35, 7: 0x3D}[digit] accOp := map[int]byte{0: 0x05, 1: 0x0D, 2: 0x15, 3: 0x1D, 4: 0x25, 5: 0x2D, 6: 0x35, 7: 0x3D}[digit]
i := newInstr(size, []byte{accOp}) i := newInstr(size, []byte{accOp})
i.imm = immediate(imm, size, false) immBytes, err := immediate(imm, size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
// 0x81 /digit, imm16/imm32. // 0x81 /digit, imm16/imm32.
@@ -327,7 +343,11 @@ func (e *enc) encodeALUImm(digit int, dst Operand, imm int64, size int) error {
if err := setRMDigit(i, digit, dst, size); err != nil { if err := setRMDigit(i, digit, dst, size); err != nil {
return err return err
} }
i.imm = immediate(imm, size, false) immBytes, err := immediate(imm, size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
@@ -339,7 +359,7 @@ func (e *enc) encodeTest(ops []Operand, size int) error {
} }
src, dst := ops[0], ops[1] src, dst := ops[0], ops[1]
if imm, ok := src.(Imm); ok { if imm, ok := src.(Imm); ok {
// TEST r/m, imm: 0xF6 (8-bit) / 0xF7 /0 — but the Go assembler // TEST r/m, imm: 0xF6 (8-bit) / 0xF7 /0, but the Go assembler
// always uses the accumulator forms (A8/A9, no ModR/M) when the // always uses the accumulator forms (A8/A9, no ModR/M) when the
// register operand is AL/AX, whatever the immediate's width. // register operand is AL/AX, whatever the immediate's width.
if r, ok := dst.(Reg); ok && r.idx == 0 { if r, ok := dst.(Reg); ok && r.idx == 0 {
@@ -348,7 +368,11 @@ func (e *enc) encodeTest(ops []Operand, size int) error {
op = 0xA8 op = 0xA8
} }
i := newInstr(size, []byte{op}) i := newInstr(size, []byte{op})
i.imm = immediate(int64(imm), size, false) immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
op := byte(0xF7) op := byte(0xF7)
@@ -359,7 +383,11 @@ func (e *enc) encodeTest(ops []Operand, size int) error {
if err := setRMDigit(i, 0, dst, size); err != nil { if err := setRMDigit(i, 0, dst, size); err != nil {
return err return err
} }
i.imm = immediate(int64(imm), size, false) immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
srcReg, ok := src.(Reg) srcReg, ok := src.(Reg)
@@ -457,7 +485,13 @@ func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
} }
return e.emit(i) return e.emit(i)
} }
// 0xC0 (8-bit) / 0xC1, imm8. // 0xC0 (8-bit) / 0xC1, imm8. The count is an unsigned byte: go tool asm
// rejects negative and ≥256 counts, and the hardware masks the count, so
// a silent truncation ($300 encoding 44) would shift by a different
// amount than the source states.
if imm < 0 || imm > 255 {
return fmt.Errorf("shift count $%d is out of the 0..255 range", int64(imm))
}
op := byte(0xC1) op := byte(0xC1)
if size == 1 { if size == 1 {
op = 0xC0 op = 0xC0
@@ -466,7 +500,7 @@ func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
if err := setRMDigit(i, digit, dst, size); err != nil { if err := setRMDigit(i, digit, dst, size); err != nil {
return err return err
} }
i.imm = []byte{byte(int8(imm))} i.imm = []byte{byte(imm)}
return e.emit(i) return e.emit(i)
} }
@@ -508,7 +542,11 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
if err := setRM(i, dstReg, ops[1], size); err != nil { if err := setRM(i, dstReg, ops[1], size); err != nil {
return err return err
} }
i.imm = immediate(int64(imm), size, false) immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i) return e.emit(i)
} }
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops)) return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
@@ -516,10 +554,21 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
// --- PUSH / POP ------------------------------------------------------------- // --- PUSH / POP -------------------------------------------------------------
func (e *enc) encodePushPop(ops []Operand, push bool) error { func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
if len(ops) != 1 { if len(ops) != 1 {
return fmt.Errorf("PUSH/POP expects 1 operand, got %d", len(ops)) return fmt.Errorf("PUSH/POP expects 1 operand, got %d", len(ops))
} }
// In 64-bit mode go tool asm knows the 64-bit push (the default, with or
// without the Q suffix) and the 16-bit W form with its 0x66 operand-size
// prefix, and rejects the B and L spellings outright ("illegal in 64-bit
// mode"); silently widening those would push a different width than the
// source states.
switch size {
case 0, 8, 2:
default:
return fmt.Errorf("PUSH/POP size suffix is illegal in 64-bit mode")
}
w16 := size == 2
switch op := ops[0].(type) { switch op := ops[0].(type) {
case Reg: case Reg:
base := byte(0x50) // PUSH r; POP is 0x58 base := byte(0x50) // PUSH r; POP is 0x58
@@ -527,7 +576,7 @@ func (e *enc) encodePushPop(ops []Operand, push bool) error {
base = 0x58 base = 0x58
} }
// PUSH/POP default to 64-bit in 64-bit mode; no REX.W needed. // PUSH/POP default to 64-bit in 64-bit mode; no REX.W needed.
i := &instr{opcode: []byte{base + byte(op.idx&7)}, modrm: -1, sib: -1} i := &instr{opSize16: w16, opcode: []byte{base + byte(op.idx&7)}, modrm: -1, sib: -1}
i.rexB = op.idx >= 8 i.rexB = op.idx >= 8
return e.emit(i) return e.emit(i)
case Mem: case Mem:
@@ -537,7 +586,7 @@ func (e *enc) encodePushPop(ops []Operand, push bool) error {
opc = 0x8F // POP r/m: /0 opc = 0x8F // POP r/m: /0
digit = 0 digit = 0
} }
i := &instr{opcode: []byte{opc}, modrm: -1, sib: -1} i := &instr{opSize16: w16, opcode: []byte{opc}, modrm: -1, sib: -1}
if err := setRMDigit(i, digit, ops[0], 8); err != nil { if err := setRMDigit(i, digit, ops[0], 8); err != nil {
return err return err
} }
@@ -547,10 +596,17 @@ func (e *enc) encodePushPop(ops []Operand, push bool) error {
return fmt.Errorf("POP does not take an immediate") return fmt.Errorf("POP does not take an immediate")
} }
if fits8(int64(op)) { if fits8(int64(op)) {
i := &instr{opcode: []byte{0x6A}, modrm: -1, sib: -1, imm: []byte{byte(int8(op))}} i := &instr{opSize16: w16, opcode: []byte{0x6A}, modrm: -1, sib: -1, imm: []byte{byte(int8(op))}}
return e.emit(i) return e.emit(i)
} }
i := &instr{opSize16: false, opcode: []byte{0x68}, modrm: -1, sib: -1, imm: le32(int64(op))} // PUSH imm32, sign-extended to 64 bits; go tool asm bounds the
// immediate by the same signed/unsigned 32-bit span as every other
// scalar immediate.
immBytes, err := immediate(int64(op), 8, false)
if err != nil {
return err
}
i := &instr{opSize16: w16, opcode: []byte{0x68}, modrm: -1, sib: -1, imm: immBytes}
return e.emit(i) return e.emit(i)
} }
return fmt.Errorf("PUSH/POP: invalid operand") return fmt.Errorf("PUSH/POP: invalid operand")
@@ -575,6 +631,24 @@ func (e *enc) encodeJmpRel(ops []Operand, opcode []byte) error {
return e.emit(&instr{opcode: opcode, modrm: -1, sib: -1, imm: le32(int64(imm))}) return e.emit(&instr{opcode: opcode, modrm: -1, sib: -1, imm: le32(int64(imm))})
} }
// encodeIndirectBranch encodes JMP/CALL through a register or memory operand:
// FF /4 for JMP, FF /2 for CALL. The operand size is fixed at 64 bits in
// 64-bit mode, so no REX.W is emitted; a REX appears only for R8-R15 bases.
func (e *enc) encodeIndirectBranch(mnem string, ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 1 operand, got %d", mnem, len(ops))
}
digit := 4 // JMP r/m64
if mnem == "CALL" {
digit = 2 // CALL r/m64
}
i := &instr{opcode: []byte{0xFF}, modrm: -1, sib: -1}
if err := setRMDigit(i, digit, ops[0], 8); err != nil {
return err
}
return e.emit(i)
}
// condCode maps a Plan 9 conditional-jump mnemonic to its x86 condition code. // condCode maps a Plan 9 conditional-jump mnemonic to its x86 condition code.
func condCode(upper string) (int, bool) { func condCode(upper string) (int, bool) {
if len(upper) < 2 || upper[0] != 'J' || upper == "JMP" { if len(upper) < 2 || upper[0] != 'J' || upper == "JMP" {
@@ -621,19 +695,28 @@ func (e *enc) encodeJcc(cc int, ops []Operand) error {
// immediate encodes an immediate of the given operand size. full64 selects the // immediate encodes an immediate of the given operand size. full64 selects the
// 64-bit immediate form (only valid for MOV r64, imm64); otherwise a 32-bit // 64-bit immediate form (only valid for MOV r64, imm64); otherwise a 32-bit
// sign-extended immediate is used for 64-bit operands. // sign-extended immediate is used for 64-bit operands.
func immediate(v int64, size int, full64 bool) []byte { //
// The span mirrors go tool asm: every scalar immediate must fit a signed or
// unsigned 32-bit word, and the narrower fields then take the low bits
// silently (ADDB $256, AL encodes imm8 0, MOVW $65536, AX imm16 0). Only the
// imm64 form may exceed the span; anything wider elsewhere is an error rather
// than a truncation the source never asked for.
func immediate(v int64, size int, full64 bool) ([]byte, error) {
if !(size == 8 && full64) && (v < -(1<<31) || v > (1<<32)-1) {
return nil, fmt.Errorf("immediate $%d does not fit in 32 bits", v)
}
switch size { switch size {
case 1: case 1:
return []byte{byte(int8(v))} return []byte{byte(int8(v))}, nil
case 2: case 2:
return le16(v) return le16(v), nil
case 4: case 4:
return le32(v) return le32(v), nil
default: // 8 default: // 8
if full64 { if full64 {
return le64(v) return le64(v), nil
} }
return le32(v) // sign-extended imm32 return le32(v), nil // sign-extended imm32
} }
} }
@@ -678,7 +761,7 @@ func (e *enc) encodeCmov(upper string, ops []Operand) error {
} }
// encodeSet encodes a conditional byte set: SET + condition (SETNE, SETEQ, …), // encodeSet encodes a conditional byte set: SET + condition (SETNE, SETEQ, …),
// always a byte write — 0F 90+cc /0 into a register or memory operand. // always a byte write, 0F 90+cc /0 into a register or memory operand.
func (e *enc) encodeSet(upper string, ops []Operand) error { func (e *enc) encodeSet(upper string, ops []Operand) error {
if len(ops) != 1 { if len(ops) != 1 {
return fmt.Errorf("SETcc expects 1 operand, got %d", len(ops)) return fmt.Errorf("SETcc expects 1 operand, got %d", len(ops))
@@ -711,8 +794,8 @@ var countOp = map[string]struct {
"POPCNT": {0xB8, 0xF3}, "POPCNT": {0xB8, 0xF3},
} }
// encodeCount encodes the bit-scan and bit-count family — BSF (0F BC), // encodeCount encodes the bit-scan and bit-count family, BSF (0F BC),
// BSR (0F BD), TZCNT (F3 0F BC), LZCNT (F3 0F BD) and POPCNT (F3 0F B8) — // BSR (0F BD), TZCNT (F3 0F BC), LZCNT (F3 0F BD) and POPCNT (F3 0F B8)
// with reg = dst and rm = src. The size suffix selects the operand width // with reg = dst and rm = src. The size suffix selects the operand width
// (BSFQ, TZCNTL, …). Note BSF/BSR leave the destination undefined when the // (BSFQ, TZCNTL, …). Note BSF/BSR leave the destination undefined when the
// source is zero (unlike their F3-prefixed counterparts); callers must // source is zero (unlike their F3-prefixed counterparts); callers must
@@ -755,15 +838,24 @@ func (e *enc) encodeBswap(ops []Operand, size int) error {
// width. The source is narrower than the destination, so the plain size-suffix // width. The source is narrower than the destination, so the plain size-suffix
// convention does not apply to these names. // convention does not apply to these names.
var movExtendOp = map[string]struct { var movExtendOp = map[string]struct {
op []byte op []byte
dst64 bool dstSize int
}{ }{
"MOVBLZX": {[]byte{0x0F, 0xB6}, false}, // byte → long, zero-extend "MOVBLZX": {[]byte{0x0F, 0xB6}, 4}, // byte → long, zero-extend
"MOVBQZX": {[]byte{0x0F, 0xB6}, true}, // byte → quad, zero-extend "MOVBQZX": {[]byte{0x0F, 0xB6}, 8}, // byte → quad, zero-extend
"MOVWLZX": {[]byte{0x0F, 0xB7}, false}, // word → long, zero-extend "MOVWLZX": {[]byte{0x0F, 0xB7}, 4}, // word → long, zero-extend
"MOVWQZX": {[]byte{0x0F, 0xB7}, true}, // word → quad, zero-extend "MOVWQZX": {[]byte{0x0F, 0xB7}, 8}, // word → quad, zero-extend
"MOVWLSX": {[]byte{0x0F, 0xBF}, false}, // word → long, sign-extend "MOVWLSX": {[]byte{0x0F, 0xBF}, 4}, // word → long, sign-extend
"MOVLQSX": {[]byte{0x63}, true}, // long → quad, sign-extend (MOVSXD) "MOVLQSX": {[]byte{0x63}, 8}, // long → quad, sign-extend (MOVSXD)
"MOVBWZX": {[]byte{0x0F, 0xB6}, 2}, // byte → word, zero-extend
"MOVBWSX": {[]byte{0x0F, 0xBE}, 2}, // byte → word, sign-extend
"MOVBLSX": {[]byte{0x0F, 0xBE}, 4}, // byte → long, sign-extend
"MOVBQSX": {[]byte{0x0F, 0xBE}, 8}, // byte → quad, sign-extend
"MOVWQSX": {[]byte{0x0F, 0xBF}, 8}, // word → quad, sign-extend
// A long → quad zero-extend is a plain 32-bit move: every 32-bit
// operation zero-extends its result into the full register, so the
// toolchain lowers MOVLQZX to the plain MOVL encoding.
"MOVLQZX": {[]byte{0x8B}, 4},
} }
// encodeMovExtend encodes a mixed-width extending move: reg = dst (the wider // encodeMovExtend encodes a mixed-width extending move: reg = dst (the wider
@@ -777,12 +869,30 @@ func (e *enc) encodeMovExtend(base string, ops []Operand) error {
if !ok { if !ok {
return fmt.Errorf("%s destination must be a register", base) return fmt.Errorf("%s destination must be a register", base)
} }
size := 4 i := newInstr(spec.dstSize, spec.op)
if spec.dst64 { if err := setRM(i, dstReg, ops[0], spec.dstSize); err != nil {
size = 8 return err
} }
i := newInstr(size, spec.op) return e.emit(i)
if err := setRM(i, dstReg, ops[0], size); err != nil { }
// encodePmovmskb encodes PMOVMSKB, the legacy SSE2 byte mask extract: the
// XMM source's sign bytes pack into a GP destination, 66 0F D7 /r.
func (e *enc) encodePmovmskb(base string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", base, len(ops))
}
srcReg, srcVec := vecReg(ops[0])
if !srcVec {
return fmt.Errorf("%s source must be an XMM register", base)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("%s destination must be a register", base)
}
i := newInstr(4, []byte{0x0F, 0xD7})
i.prefix = 0x66
if err := setRM(i, dstReg, srcReg, 4); err != nil {
return err return err
} }
return e.emit(i) return e.emit(i)
@@ -801,8 +911,8 @@ type sseMove struct {
} }
var sseMoveTable = map[string]sseMove{ var sseMoveTable = map[string]sseMove{
"MOVOU": {0xF3, 0x6F, 0x7F}, // MOVDQU — unaligned octa "MOVOU": {0xF3, 0x6F, 0x7F}, // MOVDQU, unaligned octa
"MOVO": {0x66, 0x6F, 0x7F}, // MOVDQA — aligned octa "MOVO": {0x66, 0x6F, 0x7F}, // MOVDQA, aligned octa
"MOVUPS": {0x00, 0x10, 0x11}, // unaligned packed single "MOVUPS": {0x00, 0x10, 0x11}, // unaligned packed single
"MOVAPS": {0x00, 0x28, 0x29}, // aligned packed single "MOVAPS": {0x00, 0x28, 0x29}, // aligned packed single
"MOVUPD": {0x66, 0x10, 0x11}, // unaligned packed double "MOVUPD": {0x66, 0x10, 0x11}, // unaligned packed double
+3 -3
View File
@@ -19,8 +19,8 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
// TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel — // TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel;
// all functions plus the file-local mask24 constant — and checks that every // all functions plus the file-local mask24 constant; and checks that every
// static-symbol load resolves to the right bytes in the image. // static-symbol load resolves to the right bytes in the image.
func TestAssembleGoFlacAVX2Kernel(t *testing.T) { func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx2_amd64.s" path := "../../go-libraries/go-flac/avx2_amd64.s"
@@ -81,7 +81,7 @@ func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
} }
// TestAssembleGoFlacAVX512Kernel assembles the whole production AVX-512 // TestAssembleGoFlacAVX512Kernel assembles the whole production AVX-512
// kernel — all functions plus the file-global idx16 constant — and checks // kernel, all functions plus the file-global idx16 constant, and checks
// that the static-symbol load resolves to the right bytes in the image. // that the static-symbol load resolves to the right bytes in the image.
func TestAssembleGoFlacAVX512Kernel(t *testing.T) { func TestAssembleGoFlacAVX512Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx512_amd64.s" path := "../../go-libraries/go-flac/avx512_amd64.s"
+2 -2
View File
@@ -94,9 +94,9 @@ DATA ·table<>+0(SB)/8, $0x1122334455667788
} }
// The debug_line program: LNE_set_address (the R_ADDR relocation // The debug_line program: LNE_set_address (the R_ADDR relocation
// carries the function address), then one row per line change — the // carries the function address), then one row per line change; the
// TEXT is on line 4 (a leading blank line precedes the include), the // TEXT is on line 4 (a leading blank line precedes the include), the
// instructions on lines 5–9 — an advance to the 20-byte end and an // instructions on lines 5-9; an advance to the 20-byte end and an
// end-of-sequence. // end-of-sequence.
linesOff := le.Uint32(dataIdx[4*2:]) linesOff := le.Uint32(dataIdx[4*2:])
lines := dataBlk[linesOff : linesOff+21] lines := dataBlk[linesOff : linesOff+21]
+120 -80
View File
@@ -6,6 +6,7 @@ package asm
import ( import (
"fmt" "fmt"
"sort" "sort"
"strconv"
"sourcedock.dev/petrbalvin/gasm-devkit/ast" "sourcedock.dev/petrbalvin/gasm-devkit/ast"
) )
@@ -15,7 +16,7 @@ import (
// file-local static symbols are encoded RIP-relative and resolved within the // file-local static symbols are encoded RIP-relative and resolved within the
// image, so the raw bytes are self-consistent and executable at any base // image, so the raw bytes are self-consistent and executable at any base
// address; references to external symbols are recorded as relocations // address; references to external symbols are recorded as relocations
// (Funcs[i].Relocs, Externals) and left unresolved — the object-file // (Funcs[i].Relocs, Externals) and left unresolved, the object-file
// emitters turn them into linker relocations. // emitters turn them into linker relocations.
type Image struct { type Image struct {
Code []byte // concatenated function bodies Code []byte // concatenated function bodies
@@ -24,6 +25,10 @@ type Image struct {
Symbols map[string]int // static symbol → byte offset within the image Symbols map[string]int // static symbol → byte offset within the image
DataSyms []DataSymbol // GLOBL symbols, in layout order DataSyms []DataSymbol // GLOBL symbols, in layout order
Externals []string // referenced but undefined symbols, sorted Externals []string // referenced but undefined symbols, sorted
// SourcePath is the assembled file's path, recorded in the DWARF
// sections in place of a placeholder name. Empty when the image was
// not built from a named file.
SourcePath string
} }
// FuncLayout describes one assembled function within an Image. // FuncLayout describes one assembled function within an Image.
@@ -80,27 +85,34 @@ func (fl *FuncLayout) LineAt(offset int) int {
return 0 return 0
} }
// RelocKind Reloc is one static-symbol reference within a function body: the disp32 // RelocKind discriminates the relocation a static-symbol reference needs;
// field at Off (function-relative) must reach the symbol plus Addend, // the encoders record one per SB reference, and the object-file emitters map
// measured from After, the address just past the instruction. An External // it to their format's relocation type.
// relocation names a symbol no GLOBL in the file defines; the object-file
// emitters carry it into the output's relocation table.
// RelocKind discriminates the type of relocation needed.
type RelocKind int type RelocKind int
const ( const (
RelPCRel32 RelocKind = iota // 32-bit PC-relative (amd64) RelPCRel32 RelocKind = iota // 32-bit PC-relative (amd64)
RelCall // R_CALL: CALL to a function symbol (amd64)
RelTLSLE // R_TLS_LE: local-exec TLS load, no symbol (amd64 guard)
RelRISCVPCRELIType // R_RISCV_PCREL_ITYPE (AUIPC + I-type pair) RelRISCVPCRELIType // R_RISCV_PCREL_ITYPE (AUIPC + I-type pair)
RelRISCVPCRELSType // R_RISCV_PCREL_STYPE (AUIPC + S-type pair) RelRISCVPCRELSType // R_RISCV_PCREL_STYPE (AUIPC + S-type pair)
RelRISCVJal // R_RISCV_JAL (J-type call) RelRISCVJal // R_RISCV_JAL (J-type call)
RelPCRelAbs // 32-bit absolute (R_RISCV_32)
RelLoong64AddrHi // R_LOONG64_ADDR_HI (pcalau12i) RelLoong64AddrHi // R_LOONG64_ADDR_HI (pcalau12i)
RelLoong64AddrLo // R_LOONG64_ADDR_LO (addi.d/ld/st) RelLoong64AddrLo // R_LOONG64_ADDR_LO (addi.d/ld/st)
RelArm64Addr // R_ADDRARM64 (ADRP + ADD/LDR/STR pair) RelArm64Addr // R_ADDRARM64 (ADRP + ADD pair)
RelArm64Branch // R_CALLARM64 (BL instruction) RelArm64Branch // R_CALLARM64 (BL instruction)
RelArm64LDST64 // R_ARM64_PCREL_LDST64 (ADRP + 64-bit LDR/STR pair)
RelLoong64Branch // R_CALLLOONG64 (BL instruction)
) )
type Reloc struct { type Reloc struct {
// Off is the function-relative offset of the field the linker patches
// and After the address just past the instruction, the base the
// assembler measures PC-relative displacements from. Name plus
// Addend select the target: the symbol plus the byte offset. An
// External relocation names a symbol no GLOBL in the file defines;
// the object-file emitters carry it into the output's relocation
// table.
Off int Off int
After int After int
Name string Name string
@@ -132,7 +144,7 @@ func (img *Image) Bytes() []byte {
// reference to a file-local static symbol becomes a RIP-relative load whose // reference to a file-local static symbol becomes a RIP-relative load whose
// displacement is resolved against that layout; a reference to a symbol no // displacement is resolved against that layout; a reference to a symbol no
// GLOBL defines is recorded as an external relocation (Externals) with its // GLOBL defines is recorded as an external relocation (Externals) with its
// displacement left zero — the object-file emitters resolve it at link // displacement left zero, the object-file emitters resolve it at link
// time, while the raw image (Bytes) cannot represent it. // time, while the raw image (Bytes) cannot represent it.
func AssembleFile(f *ast.File) (*Image, error) { func AssembleFile(f *ast.File) (*Image, error) {
dataSyms, err := collectData(f) dataSyms, err := collectData(f)
@@ -145,7 +157,8 @@ func AssembleFile(f *ast.File) (*Image, error) {
} }
link := &linkInfo{symbols: known, allowExternal: true} link := &linkInfo{symbols: known, allowExternal: true}
img := &Image{Symbols: map[string]int{}} img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
textOff := map[string]int{}
type asmFunc struct { type asmFunc struct {
name string name string
patches []sbPatch patches []sbPatch
@@ -183,6 +196,7 @@ func AssembleFile(f *ast.File) (*Image, error) {
for _, s := range steps { for _, s := range steps {
fl.Spadj = append(fl.Spadj, SpadjStep{PC: s.pc, Value: s.value}) fl.Spadj = append(fl.Spadj, SpadjStep{PC: s.pc, Value: s.value})
} }
textOff[t.Name.Name] = len(img.Code)
img.Funcs = append(img.Funcs, fl) img.Funcs = append(img.Funcs, fl)
img.Code = append(img.Code, code...) img.Code = append(img.Code, code...)
funcs = append(funcs, asmFunc{name: t.Name.Name, patches: patches}) funcs = append(funcs, asmFunc{name: t.Name.Name, patches: patches})
@@ -215,13 +229,27 @@ func AssembleFile(f *ast.File) (*Image, error) {
base := img.Funcs[i].Offset base := img.Funcs[i].Offset
code := img.Code[base : base+img.Funcs[i].Size] code := img.Code[base : base+img.Funcs[i].Size]
for _, p := range fn.patches { for _, p := range fn.patches {
reloc := Reloc{Off: p.off, After: p.after, Name: p.name, Addend: p.addend} reloc := Reloc{Off: p.off, After: p.after, Name: p.name, Addend: p.addend, Kind: p.kind}
if p.kind == RelTLSLE {
// The TLS slot has no symbol: the linker fills the offset
// from the runtime's TLS layout.
img.Funcs[i].Relocs = append(img.Funcs[i].Relocs, reloc)
continue
}
if imgOff, ok := img.Symbols[p.name]; ok { if imgOff, ok := img.Symbols[p.name]; ok {
rel := int64(imgOff) + p.addend - int64(base+p.after) rel := int64(imgOff) + p.addend - int64(base+p.after)
if rel < -1<<31 || rel >= 1<<31 { if rel < -1<<31 || rel >= 1<<31 {
return nil, fmt.Errorf("%s: displacement to %q out of rel32 range", fn.name, p.name) return nil, fmt.Errorf("%s: displacement to %q out of rel32 range", fn.name, p.name)
} }
copy(code[p.off:p.off+4], le32(rel)) copy(code[p.off:p.off+4], le32(rel))
} else if imgOff, ok := textOff[p.name]; ok {
// A CALL to a TEXT function of the same file: resolve the
// displacement against the function's layout position.
rel := int64(imgOff) + p.addend - int64(base+p.after)
if rel < -1<<31 || rel >= 1<<31 {
return nil, fmt.Errorf("%s: displacement to %q out of rel32 range", fn.name, p.name)
}
copy(code[p.off:p.off+4], le32(rel))
} else { } else {
reloc.External = true reloc.External = true
externals[p.name] = true externals[p.name] = true
@@ -246,7 +274,7 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
return nil, err return nil, err
} }
img := &Image{Symbols: map[string]int{}} img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls { for _, d := range f.Decls {
t, ok := d.(*ast.Text) t, ok := d.(*ast.Text)
if !ok { if !ok {
@@ -318,7 +346,7 @@ func AssembleFileLOONG64(f *ast.File) (*Image, error) {
return nil, err return nil, err
} }
img := &Image{Symbols: map[string]int{}} img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls { for _, d := range f.Decls {
t, ok := d.(*ast.Text) t, ok := d.(*ast.Text)
if !ok { if !ok {
@@ -419,78 +447,90 @@ type dataSym struct {
} }
// collectData gathers the file's static symbols (GLOBL) and their initial // collectData gathers the file's static symbols (GLOBL) and their initial
// contents (DATA) into byte buffers, in declaration order. // contents (DATA) into byte buffers. Two passes: the Plan 9 convention puts
// every DATA line before its symbol's GLOBL, so the symbols are registered
// before the initialisers are applied.
func collectData(f *ast.File) ([]dataSym, error) { func collectData(f *ast.File) ([]dataSym, error) {
index := map[string]int{} index := map[string]int{}
var syms []dataSym var syms []dataSym
for _, d := range f.Decls { for _, d := range f.Decls {
switch dd := d.(type) { gd, ok := d.(*ast.Globl)
case *ast.Globl: if !ok {
if dd.Name == nil || dd.Name.Pseudo != "SB" { continue
continue }
} if gd.Name == nil || gd.Name.Pseudo != "SB" {
name := dd.Name.Name continue
if _, dup := index[name]; dup { }
return nil, fmt.Errorf("duplicate GLOBL %q", name) name := gd.Name.Name
} if _, dup := index[name]; dup {
size := 0 return nil, fmt.Errorf("duplicate GLOBL %q", name)
if dd.Size != nil && dd.Size.Imm.HasVal { }
size = int(dd.Size.Imm.Val) size := 0
} if gd.Size != nil && gd.Size.Imm.HasVal {
index[name] = len(syms) size = int(gd.Size.Imm.Val)
ds := dataSym{ }
name: name, index[name] = len(syms)
pkg: dd.Name.Pkg, ds := dataSym{
buf: make([]byte, size), name: name,
size: size, pkg: gd.Name.Pkg,
static: dd.Name.Static, buf: make([]byte, size),
} size: size,
for _, f := range dd.Flags { static: gd.Name.Static,
switch f { }
case "RODATA": for _, f := range gd.Flags {
ds.rodata = true switch f {
case "DUPOK": case "RODATA":
ds.dupok = true ds.rodata = true
case "1": case "DUPOK":
ds.dupok = true ds.dupok = true
case "8": default:
ds.rodata = true // Legacy numeric flag constants (runtime/textflag.h):
case "9": // DUPOK is 2, RODATA is 8; combinations arrive as one
ds.dupok = true // number (e.g. 10 = RODATA|DUPOK).
ds.rodata = true if n, err := strconv.Atoi(f); err == nil {
if n&2 != 0 {
ds.dupok = true
}
if n&8 != 0 {
ds.rodata = true
}
} }
} }
syms = append(syms, ds) }
syms = append(syms, ds)
case *ast.Data: }
if dd.Name == nil || dd.Name.Pseudo != "SB" { for _, d := range f.Decls {
continue dd, ok := d.(*ast.Data)
} if !ok {
i, ok := index[dd.Name.Name] continue
if !ok { }
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name) if dd.Name == nil || dd.Name.Pseudo != "SB" {
} continue
if dd.Value == nil || !dd.Value.Imm.HasVal { }
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name) i, ok := index[dd.Name.Name]
} if !ok {
w := dd.Width return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
switch w { }
case 1, 2, 4, 8: if dd.Value == nil || !dd.Value.Imm.HasVal {
default: return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w) }
} w := dd.Width
off := dd.Name.Offset switch w {
buf := syms[i].buf case 1, 2, 4, 8:
if off < 0 || off+int64(w) > int64(len(buf)) { default:
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf)) return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
} }
v := dd.Value.Imm.Val off := dd.Name.Offset
if dd.Value.Imm.Neg { buf := syms[i].buf
v = -v if off < 0 || off+int64(w) > int64(len(buf)) {
} return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
for j := range w { }
buf[off+int64(j)] = byte(v >> (8 * j)) v := dd.Value.Imm.Val
} if dd.Value.Imm.Neg {
v = -v
}
for j := range w {
buf[off+int64(j)] = byte(v >> (8 * j))
} }
} }
return syms, nil return syms, nil
+40 -2
View File
@@ -10,8 +10,8 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
// TestAssembleFileStaticData checks the whole-image layout — code, padding // TestAssembleFileStaticData checks the whole-image layout; code, padding
// and the data section — and that the RIP-relative displacements of static // and the data section; and that the RIP-relative displacements of static
// symbol loads resolve to the right bytes. // symbol loads resolve to the right bytes.
func TestAssembleFileStaticData(t *testing.T) { func TestAssembleFileStaticData(t *testing.T) {
f, errs := parser.Parse("d_amd64.s", ` f, errs := parser.Parse("d_amd64.s", `
@@ -128,3 +128,41 @@ DATA x<>+0(SB)/4, $1
t.Errorf("single-function SB: error %v, want a file-level-assembly error", err) t.Errorf("single-function SB: error %v, want a file-level-assembly error", err)
} }
} }
// TestCollectDataNumericFlags pins the numeric GLOBL flag constants from
// runtime/textflag.h: DUPOK is 2, RODATA is 8, and combinations arrive as
// one number (9 = NOPROF|RODATA, 10 = RODATA|DUPOK).
func TestCollectDataNumericFlags(t *testing.T) {
tests := []struct {
flags string
rodata bool
dupok bool
}{
{"2", false, true},
{"8", true, false},
{"9", true, false}, // NOPROF|RODATA, not DUPOK
{"10", true, true}, // RODATA|DUPOK
{"RODATA", true, false},
{"DUPOK", false, true},
{"RODATA|DUPOK", true, true},
}
for _, tt := range tests {
src := "TEXT \u00b7f(SB), NOSPLIT, $0\n\tRET\nGLOBL sym(SB), " + tt.flags + ", $8\n"
f, errs := parser.Parse("f_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse %q: %v", tt.flags, errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble %q: %v", tt.flags, err)
}
if len(img.DataSyms) != 1 {
t.Fatalf("%q: data syms = %d, want 1", tt.flags, len(img.DataSyms))
}
d := img.DataSyms[0]
if d.Rodata != tt.rodata || d.Dupok != tt.dupok {
t.Errorf("flags %q: rodata=%v dupok=%v, want rodata=%v dupok=%v",
tt.flags, d.Rodata, d.Dupok, tt.rodata, tt.dupok)
}
}
}
+154 -45
View File
@@ -6,6 +6,7 @@ package asm
import ( import (
"fmt" "fmt"
"math/bits" "math/bits"
"strconv"
"strings" "strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast" "sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -13,17 +14,19 @@ import (
// assembleLOONG64 assembles a LoongArch (loong64) TEXT function body into // assembleLOONG64 assembles a LoongArch (loong64) TEXT function body into
// machine code. Every instruction is 4 bytes; the MOV pseudo-instruction and // machine code. Every instruction is 4 bytes; the MOV pseudo-instruction and
// the immediate-arithmetic forms expand to 2–5 instructions when the // the immediate-arithmetic forms expand to 2-5 instructions when the
// immediate does not fit, so the layout is computed in two passes (sizes, // immediate does not fit, so the layout is computed in two passes (sizes,
// then encoding with resolved branch targets). // then encoding with resolved branch targets).
// //
// The emitted bytes match the Go toolchain's loong64 assembler, which is the // The emitted bytes match the Go toolchain's loong64 assembler, which is the
// ground-truth oracle: prologue/epilogue, FP/SP frame mapping, branch // ground-truth oracle: prologue/epilogue (including the large-frame R30
// encodings and the MOV immediate expansions all follow cmd/internal/obj/ // materialisations), FP/SP frame mapping, the stack-split guard classes, and
// loong64's asmout cases. // branch encodings all follow cmd/internal/obj/loong64. The morestack block
// at the end of split functions carries the runtime.morestack_noctxt call.
func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) { func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
fi := loong64ComputeFrame(t) fi := loong64ComputeFrame(t)
prologue := loong64Prologue(fi) prologue := loong64Prologue(fi)
guardLen := loong64GuardLen(fi)
chain := loong64JumpChain(t) chain := loong64JumpChain(t)
resolve := func(name string) string { resolve := func(name string) string {
if r, ok := chain[name]; ok { if r, ok := chain[name]; ok {
@@ -35,16 +38,17 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
var relocs []Reloc var relocs []Reloc
var spadj []SpadjStep var spadj []SpadjStep
// The prologue (3 instructions when a frame is present) raises the SP // The prologue raises the SP delta by autosize; the boundary is reported
// delta by autosize; the boundary is reported at the third instruction's // after the SP adjust instruction, exactly as the toolchain's pctospadj
// pc, exactly as the toolchain's pctospadj does. // does. The prologue (3 instructions when a frame is present) may
// materialise its store or adjust through R30, which widens it.
if fi.autosize != 0 { if fi.autosize != 0 {
spadj = append(spadj, SpadjStep{PC: 8, Value: fi.autosize}) spadj = append(spadj, SpadjStep{PC: guardLen + (loong64StoreWords(fi.autosize)+loong64AdjustWords(-int64(fi.autosize)))*4, Value: fi.autosize})
} }
// Pass 1: label offsets from the instruction sizes. // Pass 1: label offsets from the instruction sizes.
offsets := map[string]int{} offsets := map[string]int{}
pos := len(prologue) pos := guardLen + len(prologue)
for _, stmt := range t.Body { for _, stmt := range t.Body {
switch s := stmt.(type) { switch s := stmt.(type) {
case *ast.Label: case *ast.Label:
@@ -54,9 +58,25 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
} }
} }
// Pass 2: encode. Relocation offsets are recorded function-relative. // Pass 2: encode. The guard prefix precedes the prologue; its branches
out := append([]byte(nil), prologue...) // target the morestack block at the end of the function, which the first
pc := len(prologue) // pass has sized.
bodyLen := 0
{
p := guardLen + len(prologue)
for _, stmt := range t.Body {
if in, ok := stmt.(*ast.Instr); ok {
p += loong64InstrSize(in, fi)
}
}
bodyLen = p - (guardLen + len(prologue))
}
var out []byte
if fi.needSplit {
out = append(out, loong64GuardBytes(fi, guardLen+len(prologue)+bodyLen)...)
}
out = append(out, prologue...)
pc := guardLen + len(prologue)
preCount := len(relocs) preCount := len(relocs)
var lines []LineEntry var lines []LineEntry
for _, stmt := range t.Body { for _, stmt := range t.Body {
@@ -69,23 +89,29 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err) return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err)
} }
for j := preCount; j < len(relocs); j++ { for j := preCount; j < len(relocs); j++ {
relocs[j].Off += pc - len(prologue) // Make the relocation offsets function-relative: each instruction
// records its reloc offset relative to its own start, and pc is
// that instruction's offset from the function start (prologue
// included). After shifts by the same amount.
relocs[j].Off += pc
relocs[j].After += pc
} }
preCount = len(relocs) preCount = len(relocs)
lines = append(lines, LineEntry{Offset: pc, Line: in.Pos().Line}) lines = append(lines, LineEntry{Offset: pc, Line: in.Pos().Line})
// The RET's epilogue closes the frame: the SP delta returns to zero // The RET's epilogue closes the frame: the SP delta returns to zero
// after the addi.d (one instruction for a leaf, two for a non-leaf // after the frame-deallocating ADDV.
// with the LR restore).
if strings.ToUpper(in.Mnemonic.Text) == "RET" && fi.autosize != 0 { if strings.ToUpper(in.Mnemonic.Text) == "RET" && fi.autosize != 0 {
epi := 4 spadj = append(spadj, SpadjStep{PC: pc + loong64EpilogueWords(fi)*4, Value: 0})
if !fi.leaf {
epi = 8
}
spadj = append(spadj, SpadjStep{PC: pc + epi, Value: 0})
} }
out = append(out, code...) out = append(out, code...)
pc += len(code) pc += len(code)
} }
if fi.needSplit {
block, blReloc := loong64MoreStackBlock(pc)
out = append(out, block...)
relocs = append(relocs, blReloc)
pc += len(block)
}
return out, offsets, relocs, lines, spadj, nil return out, offsets, relocs, lines, spadj, nil
} }
@@ -224,15 +250,18 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
} }
return l64wordLE(uint32(immFromOperand(ops[0]))), nil return l64wordLE(uint32(immFromOperand(ops[0]))), nil
case "JMP", "B": case "JMP", "B":
return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve) return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve, relocs)
case "JAL", "CALL", "BL": case "JAL", "CALL", "BL":
return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve) return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve, relocs)
case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD": case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD":
return encodeLOONG64Mov(instr, mnem, fi, relocs) return encodeLOONG64Mov(instr, mnem, fi, relocs)
} }
// 16-bit branches (BEQ/BNE/BLT/BGE/BLTU/BGEU) and JIRL. // 16-bit branches (BEQ/BNE/BLT/BGE/BLTU/BGEU) and JIRL.
if op, ok := l64branchTable[mnem]; ok { if op, ok := l64branchTable[mnem]; ok {
if mnem == "JIRL" {
return encodeLOONG64Jirl(op, ops)
}
return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve) return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve)
} }
// Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ, // Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ,
@@ -413,11 +442,20 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
if rj < 0 || rd < 0 { if rj < 0 || rd < 0 {
return nil, fmt.Errorf("invalid register operand") return nil, fmt.Errorf("invalid register operand")
} }
// The toolchain validates the bit numbers ("illegal bit number"):
// 0..31 for the .w forms, 0..63 for the .d forms, lsb <= msb.
b := 64
if strings.HasSuffix(mnem, "W") {
b = 32
}
if msb < 0 || msb >= b || lsb < 0 || lsb >= b || lsb > msb {
return nil, fmt.Errorf("%s: illegal bit number (msb %d, lsb %d)", mnem, msb, lsb)
}
return l64wordLE(l64irir(enc.op, msb, rj, lsb, rd)), nil return l64wordLE(l64irir(enc.op, msb, rj, lsb, rd)), nil
case l64Firrr: case l64Firrr:
// ALSL: INSTR $sa, rj, rk, rd (the toolchain's optab places rj in // ALSL: INSTR $sa, rj, rk, rd (the toolchain's optab places rj in
// the second register position); the source amount is 1–4, encoded // the second register position); the source amount is 1-4, encoded
// as sa-1. // as sa-1.
if len(ops) != 4 { if len(ops) != 4 {
return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
@@ -485,7 +523,7 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
// //
// JMP/B label → b label JMP/B (rj) → jirl r0, rj, 0 // JMP/B label → b label JMP/B (rj) → jirl r0, rj, 0
// JAL/CALL/BL label → bl label JAL/CALL/BL (rj) → jirl r1, rj, 0 // JAL/CALL/BL label → bl label JAL/CALL/BL (rj) → jirl r1, rj, 0
func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string) ([]byte, error) { func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string, relocs *[]Reloc) ([]byte, error) {
if len(instr.Operands) != 1 { if len(instr.Operands) != 1 {
return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(instr.Operands)) return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(instr.Operands))
} }
@@ -502,6 +540,19 @@ func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[stri
} }
return l64wordLE(l64irr16(l64branchTable["JIRL"], 0, rj, rd)), nil return l64wordLE(l64irr16(l64branchTable["JIRL"], 0, rj, rd)), nil
} }
// Direct symbol: sym+off(SB) → b/bl with an R_CALLLOONG64 relocation
// (the linker fills the offset), as the toolchain does for CALL/BL/JAL
// and for tail-calling JMP.
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
opc := l64jumpTable["B"]
if link {
opc = l64jumpTable["BL"]
}
if relocs != nil {
*relocs = append(*relocs, Reloc{Off: 0, After: 4, Name: op.Addr.Sym.Name, Kind: RelLoong64Branch, Addend: op.Addr.Sym.Offset})
}
return l64wordLE(l64bbl(opc, 0)), nil
}
// Direct: label → b/bl. // Direct: label → b/bl.
target := resolve(l64Label(op)) target := resolve(l64Label(op))
targetOff, ok := offsets[target] targetOff, ok := offsets[target]
@@ -516,6 +567,46 @@ func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[stri
return l64wordLE(l64bbl(opc, v)), nil return l64wordLE(l64bbl(opc, v)), nil
} }
// encodeLOONG64Jirl encodes the raw JIRL spelling, JIRL rd, rj, offset, the
// form the verify trampolines use. The (rj) indirect form without an offset
// is handled by encodeLOONG64Branch.
func encodeLOONG64Jirl(op uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 3 {
return nil, fmt.Errorf("JIRL expects 3 operands, got %d", len(ops))
}
rd := l64Reg(ops[0])
rj := l64Reg(ops[1])
if rd < 0 || rj < 0 {
return nil, fmt.Errorf("invalid register operand")
}
off, ok := l64offsetOperand(ops[2])
if !ok {
return nil, fmt.Errorf("JIRL expects an immediate offset, got %q", ops[2].Raw)
}
if (int64(off)<<16)>>16 != int64(off) {
return nil, fmt.Errorf("JIRL offset %d out of the 16-bit range", off)
}
return l64wordLE(l64irr16(op, int(off), rj, rd)), nil
}
// l64offsetOperand reads a bare numeric branch offset: an immediate ($n) or a
// plain number, which parses as an empty address carrying the digits in Raw.
func l64offsetOperand(op *ast.Operand) (int32, bool) {
if op.Imm.HasVal {
v := op.Imm.Val
if op.Imm.Neg {
v = -v
}
return int32(v), true
}
if op.Kind == ast.OpAddr && op.Addr.Sym == nil && op.Addr.Base == "" && op.Addr.Index == "" {
if v, err := strconv.ParseInt(op.Raw, 0, 64); err == nil {
return int32(v), true
}
}
return 0, false
}
// encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU): // encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU):
// INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain // INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain
// turns into the 21-bit BEQZ/BNEZ form when the register is the only operand. // turns into the 21-bit BEQZ/BNEZ form when the register is the only operand.
@@ -536,6 +627,15 @@ func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, o
if rj < 0 { if rj < 0 {
return nil, fmt.Errorf("invalid register operand") return nil, fmt.Errorf("invalid register operand")
} }
if mnem == "BLTU" || mnem == "BGEU" {
// The unsigned compares have no single-register pseudo: the
// toolchain keeps the register-register form with rd = R0
// (bltu rj, r0 is never taken), not a sometimes-taken beqz.
if (v<<16)>>16 != v {
return nil, fmt.Errorf("branch to %q too far (16-bit range)", target)
}
return l64wordLE(l64irr16(op, v, rj, 0)), nil
}
if (v<<11)>>11 != v { if (v<<11)>>11 != v {
return nil, fmt.Errorf("branch to %q too far (21-bit range)", target) return nil, fmt.Errorf("branch to %q too far (21-bit range)", target)
} }
@@ -578,8 +678,8 @@ func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, o
// encodeLOONG64Branch21 encodes a single-register branch: BLTZ/BGEZ and // encodeLOONG64Branch21 encodes a single-register branch: BLTZ/BGEZ and
// BFPT/BFPF use the 21-bit offset form (register in the rj field), while // BFPT/BFPF use the 21-bit offset form (register in the rj field), while
// BGTZ/BLEZ — which the toolchain encodes with the register in the rd field // BGTZ/BLEZ, which the toolchain encodes with the register in the rd field
// and a 16-bit offset — are handled separately. // and a 16-bit offset, are handled separately.
func encodeLOONG64Branch21(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) { func encodeLOONG64Branch21(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
if len(ops) != 2 { if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
@@ -692,7 +792,7 @@ func encodeLOONG64ImmArith(mnem string, de l64DualEnc, ops []*ast.Operand) ([]by
} }
// isLoong64ShiftD reports whether a shift-immediate opcode constant is one of // isLoong64ShiftD reports whether a shift-immediate opcode constant is one of
// the 6-bit (.d) variants — the toolchain distinguishes them by the bit // the 6-bit (.d) variants, the toolchain distinguishes them by the bit
// position of the opcode field (bits [25:16]). // position of the opcode field (bits [25:16]).
func isLoong64ShiftD(op uint32) bool { func isLoong64ShiftD(op uint32) bool {
return op&0x03ff0000 != 0 && op>>25 == 0 return op&0x03ff0000 != 0 && op>>25 == 0
@@ -740,7 +840,7 @@ func l64MemOperands(ops []*ast.Operand, fi loong64FrameInfo) (rd, rj int, off in
// ---- the MOV pseudo-instruction ---- // ---- the MOV pseudo-instruction ----
// encodeLOONG64Mov encodes the MOV family — the load/store/immediate // encodeLOONG64Mov encodes the MOV family, the load/store/immediate
// workhorse of Go's loong64 assembly. MOV is an alias of MOVV (the width // workhorse of Go's loong64 assembly. MOV is an alias of MOVV (the width
// mnemonics MOVB/MOVH/MOVW/MOVV/MOVBU/MOVHU/MOVWU/MOVF/MOVD select the // mnemonics MOVB/MOVH/MOVW/MOVV/MOVBU/MOVHU/MOVWU/MOVF/MOVD select the
// access width). The forms, mirroring the toolchain: // access width). The forms, mirroring the toolchain:
@@ -775,9 +875,16 @@ func encodeLOONG64Mov(instr *ast.Instr, mnem string, fi loong64FrameInfo, relocs
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("%s $imm: invalid destination register", mnem) return nil, fmt.Errorf("%s $imm: invalid destination register", mnem)
} }
// MOVF/MOVD $imm, Fd → materialise in R30, then movgr2fr.{w,d}. // MOVW $imm, Fd is the only immediate-to-F form the toolchain's optab
if (mnem == "MOVF" || mnem == "MOVD") && loong64RegClass(operandRegName(dst)) == l64ClsFP { // accepts (AMOVW's C_12CON against C_FREG): it materialises the
return encodeLOONG64ImmToFp(rd, l64Imm64(src), mnem), nil // constant in R30 and moves it across with movgr2fr.w. MOVV/MOVF/
// MOVD are illegal combinations there, and are diagnosed here rather
// than silently written into the GPR of the register's number.
if loong64RegClass(operandRegName(dst)) == l64ClsFP {
if mnem != "MOVW" {
return nil, fmt.Errorf("%s $imm: illegal combination with an F register destination (only MOVW $c, Fd is supported)", mnem)
}
return encodeLOONG64ImmToFp(rd, l64Imm64(src))
} }
return encodeLOONG64LoadImm(rd, l64Imm64(src), mnem), nil return encodeLOONG64LoadImm(rd, l64Imm64(src), mnem), nil
} }
@@ -858,8 +965,8 @@ func loong64MovSize(mnem string, ops []*ast.Operand, fi loong64FrameInfo) int {
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" { if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" {
return 8 // pcalau12i + addi.d return 8 // pcalau12i + addi.d
} }
if (mnem == "MOVF" || mnem == "MOVD") && loong64RegClass(operandRegName(dst)) == l64ClsFP { if loong64RegClass(operandRegName(dst)) == l64ClsFP {
return 8 // addi/ori r30 + movgr2fr return 8 // ori/addi.w r30 + movgr2fr.w (an encode-time diagnostic when invalid)
} }
v := l64Imm64(src) v := l64Imm64(src)
if v == 0 { if v == 0 {
@@ -900,22 +1007,24 @@ func loong64MovSize(mnem string, ops []*ast.Operand, fi loong64FrameInfo) int {
} }
} }
// encodeLOONG64ImmToFp materialises a 12-bit immediate in R30 and moves it to // encodeLOONG64ImmToFp materialises a 12-bit immediate in R30 and moves it
// an F register (the toolchain's case 34: movgr2fr.w/movgr2fr.d). // to an F register, the toolchain's expansion of MOVW $c, Fd: ori (which
func encodeLOONG64ImmToFp(fd int, v int64, mnem string) []byte { // zero-extends) for the positive span, addi.w for zero and the negative
// ori for positive constants, addi.d for zero/negative. // span, then movgr2fr.w. The toolchain's optab accepts no wider constant on
op := uint32(0x00b << 22) // this path (it never materialises one fully first), so values outside
if v > 0 { // [-2048, 4095] are diagnosed rather than masked into si12.
op = 0x00e << 22 func encodeLOONG64ImmToFp(fd int, v int64) ([]byte, error) {
if v < -2048 || v > 4095 {
return nil, fmt.Errorf("MOVW $%d: immediate out of the [-2048, 4095] range for an F register destination", v)
} }
mov := uint32(0x452a << 10) // movgr2fr.d op := uint32(0x00a << 22) // addi.w r30, r0, v (sign-extends)
if mnem == "MOVF" { if v > 0 {
mov = 0x4529 << 10 // movgr2fr.w op = 0x00e << 22 // ori r30, r0, v (zero-extends)
} }
return l64WordsLE( return l64WordsLE(
l64irr(op, int(v), 0, 30), l64irr(op, int(v), 0, 30),
l64rr(mov, 30, fd), l64rr(0x4529<<10, 30, fd), // movgr2fr.w fd, r30
) ), nil
} }
// ---- 64-bit immediate classification ---- // ---- 64-bit immediate classification ----
+16 -14
View File
@@ -9,7 +9,7 @@ package asm
// an opcode constant, and the format selects the bit layout. The opcode // an opcode constant, and the format selects the bit layout. The opcode
// constants and formats are transcribed from the Go toolchain's own loong64 // constants and formats are transcribed from the Go toolchain's own loong64
// backend (cmd/internal/obj/loong64), so the emitted bytes match `go tool asm` // backend (cmd/internal/obj/loong64), so the emitted bytes match `go tool asm`
// exactly — the ground-truth oracle for the verify suite. // exactly, the ground-truth oracle for the verify suite.
// //
// All LoongArch instructions are 32 bits, little-endian. The formats used // All LoongArch instructions are 32 bits, little-endian. The formats used
// here (per the LoongArch Volume I specification): // here (per the LoongArch Volume I specification):
@@ -33,8 +33,8 @@ package asm
import "maps" import "maps"
// loong64RegNum returns the 5-bit register number for a LoongArch register // loong64RegNum returns the 5-bit register number for a LoongArch register
// name: R0–R31 (integer), F0–F31 (floating point), FCC0–FCC7 (condition // name: R0-R31 (integer), F0-F31 (floating point), FCC0-FCC7 (condition
// flags), FCSR0–FCSR31 (control/status) and the ABI aliases the runtime's // flags), FCSR0-FCSR31 (control/status) and the ABI aliases the runtime's
// assembly uses. Returns -1 for an unrecognised name. // assembly uses. Returns -1 for an unrecognised name.
func loong64RegNum(name string) int { func loong64RegNum(name string) int {
switch name { switch name {
@@ -103,7 +103,7 @@ func loong64RegNum(name string) int {
case "R31", "S8": case "R31", "S8":
return 31 return 31
} }
// F0–F31, FCC0–FCC7, FCSR0–FCSR31. // F0-F31, FCC0-FCC7, FCSR0-FCSR31.
if len(name) >= 4 && name[:4] == "FCSR" { if len(name) >= 4 && name[:4] == "FCSR" {
return loong64RegSpecial(name[4:], 31) return loong64RegSpecial(name[4:], 31)
} }
@@ -199,7 +199,9 @@ func l64rrrr(op uint32, r1, r2, r3, r4 int) uint32 {
} }
// l64irir encodes a BSTRINS/BSTRPICK instruction: op | msb<<16 | rj<<5 | lsb<<10 | rd. // l64irir encodes a BSTRINS/BSTRPICK instruction: op | msb<<16 | rj<<5 | lsb<<10 | rd.
// The msb/lsb fields are 6 bits wide (0–63) and are validated by the caller. // The msb/lsb fields are 6 bits wide and are inserted unmasked: the caller
// must have validated them (0..31 for the .w forms, 0..63 for the .d forms,
// lsb <= msb), the same rule the toolchain enforces as "illegal bit number".
func l64irir(op uint32, msb, rj, lsb, rd int) uint32 { func l64irir(op uint32, msb, rj, lsb, rd int) uint32 {
return op | uint32(msb)<<16 | uint32(rj&0x1f)<<5 | uint32(lsb)<<10 | uint32(rd&0x1f) return op | uint32(msb)<<16 | uint32(rj&0x1f)<<5 | uint32(lsb)<<10 | uint32(rd&0x1f)
} }
@@ -280,7 +282,7 @@ var l64DualTable = map[string]l64DualEnc{}
var l64InstrTable = map[string]l64Enc{} var l64InstrTable = map[string]l64Enc{}
func init() { func init() {
// 3R — integer. // 3R, integer.
rrr := map[string]uint32{ rrr := map[string]uint32{
"ADD": 0x20 << 15, "ADDW": 0x20 << 15, "ADDV": 0x21 << 15, "ADDVU": 0x21 << 15, "ADD": 0x20 << 15, "ADDW": 0x20 << 15, "ADDV": 0x21 << 15, "ADDVU": 0x21 << 15,
"SUB": 0x22 << 15, "SUBW": 0x22 << 15, "SUBV": 0x23 << 15, "SUBVU": 0x23 << 15, "SUB": 0x22 << 15, "SUBW": 0x22 << 15, "SUBV": 0x23 << 15, "SUBVU": 0x23 << 15,
@@ -300,7 +302,7 @@ func init() {
"CRCWBW": 0x48 << 15, "CRCWHW": 0x49 << 15, "CRCWWW": 0x4a << 15, "CRCWVW": 0x4b << 15, "CRCWBW": 0x48 << 15, "CRCWHW": 0x49 << 15, "CRCWWW": 0x4a << 15, "CRCWVW": 0x4b << 15,
"CRCCWBW": 0x4c << 15, "CRCCWHW": 0x4d << 15, "CRCCWWW": 0x4e << 15, "CRCCWVW": 0x4f << 15, "CRCCWBW": 0x4c << 15, "CRCCWHW": 0x4d << 15, "CRCCWWW": 0x4e << 15, "CRCCWVW": 0x4f << 15,
} }
// 3R — floating point. // 3R, floating point.
rrr["MULF"] = 0x209 << 15 rrr["MULF"] = 0x209 << 15
rrr["MULD"] = 0x20a << 15 rrr["MULD"] = 0x20a << 15
rrr["DIVF"] = 0x20d << 15 rrr["DIVF"] = 0x20d << 15
@@ -390,12 +392,12 @@ func init() {
"ROTRV": {rrr: 0x37 << 15, imm: 0x004d << 16, shift: true}, "ROTRV": {rrr: 0x37 << 15, imm: 0x004d << 16, shift: true},
}) })
// 2RI12 — pure immediate arithmetic (LU52ID has no register form). // 2RI12, pure immediate arithmetic (LU52ID has no register form).
l64InstrTable["LU52ID"] = l64Enc{format: l64Firr, op: 0x00c << 22} l64InstrTable["LU52ID"] = l64Enc{format: l64Firr, op: 0x00c << 22}
// ADDV16 (addu16i.d): 2RI16 with the immediate shifted right by 16. // ADDV16 (addu16i.d): 2RI16 with the immediate shifted right by 16.
l64InstrTable["ADDV16"] = l64Enc{format: l64Firr16, op: 0x4 << 26} l64InstrTable["ADDV16"] = l64Enc{format: l64Firr16, op: 0x4 << 26}
// 2RI14 — LL/SC are aliased by the Go assembler to the pointer loads and // 2RI14, LL/SC are aliased by the Go assembler to the pointer loads and
// stores (ldptr/stptr), with the offset scaled by 4. // stores (ldptr/stptr), with the offset scaled by 4.
l64InstrTable["MOVWP"] = l64Enc{format: l64Firr14, op: 0x25 << 24} // stptr.w l64InstrTable["MOVWP"] = l64Enc{format: l64Firr14, op: 0x25 << 24} // stptr.w
l64InstrTable["MOVVP"] = l64Enc{format: l64Firr14, op: 0x27 << 24} // stptr.d l64InstrTable["MOVVP"] = l64Enc{format: l64Firr14, op: 0x27 << 24} // stptr.d
@@ -414,7 +416,7 @@ func init() {
// LUI is the Plan 9 spelling of lu12i.w. // LUI is the Plan 9 spelling of lu12i.w.
l64InstrTable["LUI"] = l64Enc{format: l64Fir20, op: 0x0a << 25} l64InstrTable["LUI"] = l64Enc{format: l64Fir20, op: 0x0a << 25}
// 4R — fused multiply-add. // 4R, fused multiply-add.
rrrr := map[string]uint32{ rrrr := map[string]uint32{
"FMADDF": 0x81 << 20, "FMADDD": 0x82 << 20, "FMADDF": 0x81 << 20, "FMADDD": 0x82 << 20,
"FMSUBF": 0x85 << 20, "FMSUBD": 0x86 << 20, "FMSUBF": 0x85 << 20, "FMSUBD": 0x86 << 20,
@@ -425,7 +427,7 @@ func init() {
l64InstrTable[m] = l64Enc{format: l64Frrrr, op: op} l64InstrTable[m] = l64Enc{format: l64Frrrr, op: op}
} }
// IRIR — bit-field insert/extract. // IRIR, bit-field insert/extract.
irir := map[string]uint32{ irir := map[string]uint32{
"BSTRINSW": 0x3<<21 | 0x0<<15, "BSTRINSW": 0x3<<21 | 0x0<<15,
"BSTRINSV": 0x2 << 22, "BSTRINSV": 0x2 << 22,
@@ -436,7 +438,7 @@ func init() {
l64InstrTable[m] = l64Enc{format: l64Firir, op: op} l64InstrTable[m] = l64Enc{format: l64Firir, op: op}
} }
// 3RI2 — ALSL. // 3RI2, ALSL.
irrr := map[string]uint32{ irrr := map[string]uint32{
"ALSLW": 0x2 << 17, "ALSLWU": 0x3 << 17, "ALSLV": 0x16 << 17, "ALSLW": 0x2 << 17, "ALSLWU": 0x3 << 17, "ALSLV": 0x16 << 17,
} }
@@ -452,7 +454,7 @@ func init() {
// PRELD. // PRELD.
l64InstrTable["PRELD"] = l64Enc{format: l64Fpreld, op: 0x0ab << 22} l64InstrTable["PRELD"] = l64Enc{format: l64Fpreld, op: 0x0ab << 22}
// Atomics — 3R with the AM field order (rk=value, rj=address, rd=result). // Atomics, 3R with the AM field order (rk=value, rj=address, rd=result).
am := map[string]uint32{ am := map[string]uint32{
"AMSWAPB": 0x070B8 << 15, "AMSWAPH": 0x070B9 << 15, "AMSWAPB": 0x070B8 << 15, "AMSWAPH": 0x070B9 << 15,
"AMSWAPW": 0x070C0 << 15, "AMSWAPV": 0x070C1 << 15, "AMSWAPW": 0x070C0 << 15, "AMSWAPV": 0x070C1 << 15,
@@ -477,7 +479,7 @@ func init() {
} }
// l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the // l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the
// register move between the integer and floating-point register banks — the // register move between the integer and floating-point register banks, the
// MOVW/MOVV specials the Go assembler accepts. // MOVW/MOVV specials the Go assembler accepts.
var l64FpMovTable = map[string]uint32{ var l64FpMovTable = map[string]uint32{
"MOVV.R.F": 0x452a << 10, // movgr2fr.d "MOVV.R.F": 0x452a << 10, // movgr2fr.d
+37
View File
@@ -291,3 +291,40 @@ done:
t.Errorf("code = % x\nwant % x", code, want) t.Errorf("code = % x\nwant % x", code, want)
} }
} }
// TestLOONG64IndirectBranch pins the indirect branch encodings: JMP (Rj) and
// JAL (Rj) lower to jirl, and the raw JIRL spelling encodes the written
// offset (the Go loong64 assembler deletes raw JIRL instructions entirely,
// so this form is a gasm-only superset with faithful semantics).
func TestLOONG64IndirectBranch(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (R4)
JIRL R0, R4, 8
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x4C000080, // jirl r0, r4, 0
0x4C002080, // jirl r0, r4, 8
0x4C000020, // jirl r0, r1, 0 (RET)
)
// JAL (R5) links, so the toolchain gives the function its autosize-8
// prologue and epilogue around the call and the closing RET.
fn = firstTextLOONG64(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JAL (R5)
RET
`)
code = assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x29FFE061, // st.d r1, -8(r3) (prologue saves RA below the new SP)
0x02FFE063, // addi.d r3, r3, -8 (prologue opens the frame)
0x29C00061, // st.d r1, 0(r3) (prologue saves RA at SP)
0x4C0000A1, // jirl r1, r5, 0
0x28C00061, // ld.d r1, 0(r3) (epilogue restores RA)
0x02C02063, // addi.d r3, r3, 8
0x4C000020, // jirl r0, r1, 0 (RET)
)
}
+223 -12
View File
@@ -20,14 +20,20 @@ import (
// (the toolchain aligns frames with `if autosize&4 != 0 { autosize += 4 }`). // (the toolchain aligns frames with `if autosize&4 != 0 { autosize += 4 }`).
// A leaf function (no calls) with a zero frame gets no prologue at all. // A leaf function (no calls) with a zero frame gets no prologue at all.
// //
// Prologue (autosize > 0), byte-identical to the toolchain: // Prologue (autosize > 0, small), byte-identical to the toolchain:
// //
// MOVV R1, -autosize(R3) // save LR below the new SP (traceback-safe) // MOVV R1, -autosize(R3) // save LR below the new SP (traceback-safe)
// ADDV $-autosize, R3 // open the frame // ADDV $-autosize, R3 // open the frame
// MOVV R1, 0(R3) // save LR again at SP (signal-safety) // MOVV R1, 0(R3) // save LR again at SP (signal-safety)
// //
// Large frames (autosize past the 12-bit offset or immediate ranges) expand
// the store and the adjust through REGTMP (R30) exactly as the toolchain's
// assembler does: the store via the rounding LU12IW split, the adjust via
// the floor LU12IW/ORI split.
//
// Epilogue: MOVV 0(R3), R1; ADDV $autosize, R3 (non-leaf only for the LR // Epilogue: MOVV 0(R3), R1; ADDV $autosize, R3 (non-leaf only for the LR
// restore); the RET's jirl r0, r1, 0 follows. // restore; the adjust materialised when the immediate does not fit); the
// RET's jirl r0, r1, 0 follows.
// loong64FrameInfo holds the frame layout derived from a TEXT directive. // loong64FrameInfo holds the frame layout derived from a TEXT directive.
type loong64FrameInfo struct { type loong64FrameInfo struct {
@@ -36,6 +42,11 @@ type loong64FrameInfo struct {
args int // the declared -argsize args int // the declared -argsize
noSplit bool // the NOSPLIT flag noSplit bool // the NOSPLIT flag
leaf bool // no call instructions in the body leaf bool // no call instructions in the body
// Stack-split guard state: like amd64 and arm64, a leaf function with a
// small autosize is auto-marked NOSPLIT by the toolchain.
needSplit bool
splitClass int // 0: <=StackSmall, 1: <=StackBig, 2: >StackBig
} }
// loong64ComputeFrame derives the frame layout for a TEXT function. // loong64ComputeFrame derives the frame layout for a TEXT function.
@@ -59,9 +70,157 @@ func loong64ComputeFrame(t *ast.Text) loong64FrameInfo {
// A zero-frame non-leaf function still opens an 8-byte frame for LR. // A zero-frame non-leaf function still opens an 8-byte frame for LR.
fi.autosize = 8 fi.autosize = 8
} }
switch {
case fi.noSplit:
case fi.autosize < stackSmall && fi.leaf:
// Auto-NOSPLIT, as the toolchain's leaf mark concludes.
default:
fi.needSplit = true
switch {
case fi.autosize <= stackSmall:
fi.splitClass = 0
case fi.autosize <= stackBig:
fi.splitClass = 1
default:
fi.splitClass = 2
}
}
return fi return fi
} }
// loong64GuardLen returns the byte length of the stack-split guard prefix
// (zero when the function needs no guard). The big class materialises two
// constants through R30; each materialisation shrinks by one word when the
// constant's low 12 bits are zero.
func loong64GuardLen(fi loong64FrameInfo) int {
if !fi.needSplit {
return 0
}
off := int64(fi.autosize - stackSmall)
switch fi.splitClass {
case 0:
return 12
case 1:
if off <= 2048 {
return 16 // ADDV $-off fits the signed 12-bit immediate
}
return 24 // MOVV + LU12IW + ORI + ADDV + SGTU + BEQ
default:
// MOVV + [mat] + SGTU + BNE + [mat] + ADDV + SGTU + BEQ
return (6 + loong64MatLen(off) + loong64MatLen(-off)) * 4
}
}
// loong64MatLen reports the word count of materialising v in R30: a value
// with a zero high part needs only the ORI (the toolchain's MOVW $v, R30),
// one with a zero low part only the LU12IW.
func loong64MatLen(v int64) int {
if v>>12 == 0 || v&0xFFF == 0 {
return 1
}
return 2
}
// loong64MatWords appends the words that materialise v in R30, splitting it
// as v>>12 plus the zero-extended low 12 bits.
func loong64MatWords(ws []uint32, v int64) []uint32 {
hi := v >> 12
lo := v & 0xFFF
if hi == 0 {
return append(ws, l64irr(l64OriOp, int(v), 0, 30))
}
ws = append(ws, l64ir(l64Lu12iwOp, int(hi), 30))
if lo != 0 {
ws = append(ws, l64irr(l64OriOp, int(lo), 30, 30))
}
return ws
}
// The LU12IW and ORI opcode bases (2RI20 and 2RI12 formats); the ORI reads
// and writes rd itself.
const (
l64Lu12iwOp = 0x0a << 25
l64OriOp = 0x0e << 22
)
// loong64Imm12 reports whether v fits a signed 12-bit immediate.
func loong64Imm12(v int64) bool { return v >= -2048 && v <= 2047 }
// loong64GuardBytes emits the stack-split guard prefix. blockStart is the
// function-relative address of the morestack call at the end of the function;
// branch displacements are in instructions and are computed from each
// branch's own position.
func loong64GuardBytes(fi loong64FrameInfo, blockStart int) []byte {
// MOVV 16(g), R20 (g.stackguard0), g = R22.
ws := []uint32{l64irr(l64loadStoreTable["MOVV"].ld, 16, 22, 20)}
off := int64(fi.autosize - stackSmall)
// beq appends BEQ R20, blockStart from the branch's own position.
beq := func() {
ws = append(ws, loong64Beqz(20, int32((blockStart-len(ws)*4)>>2)))
}
switch fi.splitClass {
case 0:
// SGTU SP, R20, R20; BEQ R20, more
ws = append(ws, l64rrr(l64DualTable["SGTU"].rrr, 3, 20, 20))
beq()
case 1:
ws = append(ws, loong64MediumWords(off)...)
ws = append(ws, l64rrr(l64DualTable["SGTU"].rrr, 24, 20, 20))
beq()
default:
// SGTU $off, SP, R24 catches the SP underflow a huge frame would
// cause; BNE jumps to morestack in that case.
ws = append(ws, loong64MatWords(nil, off)...)
ws = append(ws, l64rrr(l64DualTable["SGTU"].rrr, 30, 3, 24))
ws = append(ws, loong64Bnez(24, int32((blockStart-len(ws)*4)>>2)))
ws = append(ws, loong64MatWords(nil, -off)...)
ws = append(ws, l64rrr(l64DualTable["ADDV"].rrr, 30, 3, 24))
ws = append(ws, l64rrr(l64DualTable["SGTU"].rrr, 24, 20, 20))
beq()
}
return l64WordsLE(ws...)
}
// loong64MediumWords emits the medium-class stack check for offset off: the
// ADDV immediate when it fits, otherwise the same sequence with the constant
// materialised in R30.
func loong64MediumWords(off int64) []uint32 {
if off <= 2048 {
return []uint32{l64irr(l64DualTable["ADDV"].imm, int(-off), 3, 24)}
}
ws := loong64MatWords(nil, -off)
return append(ws, l64rrr(l64DualTable["ADDV"].rrr, 30, 3, 24))
}
// loong64Beqz/loong64Bnez build the 21-bit conditional branches against R0
// that the toolchain emits for its guard compares.
func loong64Beqz(rj int, dispInstr int32) uint32 {
return l64ir21(l64branch21Table["BEQZ"], int(dispInstr), rj)
}
func loong64Bnez(rj int, dispInstr int32) uint32 {
return l64ir21(l64branch21Table["BNEZ"], int(dispInstr), rj)
}
// loong64MoreStackBlock emits the trailing block: MOVV R1, R31 (save LR, the
// toolchain's OR R1, R0, R31 expansion), BL runtime.morestack_noctxt, B back
// to the function entry.
func loong64MoreStackBlock(blockStart int) ([]byte, Reloc) {
ws := []uint32{
l64rrr(l64DualTable["OR"].rrr, 0, 1, 31), // MOVV R1, R31 (OR R1, R0, R31)
l64bbl(l64jumpTable["BL"], 0), // BL, patched by the linker
}
disp := (-(blockStart + 8)) >> 2
ws = append(ws, l64bbl(l64jumpTable["B"], int(disp)))
reloc := Reloc{
Off: blockStart + 4,
After: blockStart + 8,
Name: "runtime\u00b7morestack_noctxt",
Kind: RelLoong64Branch,
}
return l64WordsLE(ws...), reloc
}
// loong64IsLeaf reports whether a function contains no call instructions // loong64IsLeaf reports whether a function contains no call instructions
// (JAL/BL/CALL), matching the toolchain's LEAF mark, which drives the frame // (JAL/BL/CALL), matching the toolchain's LEAF mark, which drives the frame
// and the epilogue shape. // and the epilogue shape.
@@ -79,17 +238,37 @@ func loong64IsLeaf(t *ast.Text) bool {
return true return true
} }
// loong64Prologue returns the prologue bytes for a loong64 function. // loong64Prologue returns the prologue bytes for a loong64 function. When
// the LR store offset leaves the toolchain's 12-bit store range ([-2046,
// 2045], BIG_12 = 2046) or the SP adjust immediate its 12-bit immediate
// range, each switches to the R30 materialisation the assembler expands it
// to: the store uses the rounding %hi/%lo split (LU12IW of (v+2048)>>12,
// REGTMP += SP, store at the raw offset), the adjust the floor split
// (LU12IW, ORI when the low part is non-zero, REGTMP += SP).
func loong64Prologue(fi loong64FrameInfo) []byte { func loong64Prologue(fi loong64FrameInfo) []byte {
if fi.autosize == 0 { if fi.autosize == 0 {
return nil return nil
} }
addiD := l64DualTable["ADDV"].imm addiD := l64DualTable["ADDV"].imm
return l64WordsLE( var ws []uint32
l64irr(l64loadStoreTable["MOVV"].st, -fi.autosize, 3, 1), // MOVV R1, -autosize(R3) storeBase := 3
l64irr(addiD, -fi.autosize, 3, 3), // ADDV $-autosize, R3 if fi.autosize > 2046 {
l64irr(l64loadStoreTable["MOVV"].st, 0, 3, 1), // MOVV R1, 0(R3) // The store goes through REGTMP: LU12IW of the rounding split,
) // REGTMP += SP, then the store at REGTMP with the truncated offset.
v := -int64(fi.autosize)
ws = append(ws, l64ir(l64Lu12iwOp, int((v+2048)>>12), 30))
ws = append(ws, l64rrr(l64DualTable["ADDV"].rrr, 3, 30, 30))
storeBase = 30
}
ws = append(ws, l64irr(l64loadStoreTable["MOVV"].st, -fi.autosize, storeBase, 1)) // MOVV R1, -autosize(base)
if loong64Imm12(-int64(fi.autosize)) {
ws = append(ws, l64irr(addiD, -fi.autosize, 3, 3)) // ADDV $-autosize, R3
} else {
ws = append(ws, loong64MatWords(nil, -int64(fi.autosize))...)
ws = append(ws, l64rrr(l64DualTable["ADDV"].rrr, 30, 3, 3))
}
ws = append(ws, l64irr(l64loadStoreTable["MOVV"].st, 0, 3, 1)) // MOVV R1, 0(R3)
return l64WordsLE(ws...)
} }
// loong64Return returns the bytes for a RET: the epilogue (restore LR and // loong64Return returns the bytes for a RET: the epilogue (restore LR and
@@ -98,17 +277,49 @@ func loong64Return(fi loong64FrameInfo) []byte {
var ws []uint32 var ws []uint32
if fi.autosize != 0 { if fi.autosize != 0 {
if !fi.leaf { if !fi.leaf {
// MOVV 0(R3), R1 — restore the link register. // MOVV 0(R3), R1, restore the link register.
ws = append(ws, l64irr(l64loadStoreTable["MOVV"].ld, 0, 3, 1)) ws = append(ws, l64irr(l64loadStoreTable["MOVV"].ld, 0, 3, 1))
} }
// ADDV $autosize, R3 — close the frame. // ADDV $autosize, R3, close the frame (materialised when the
ws = append(ws, l64irr(l64DualTable["ADDV"].imm, fi.autosize, 3, 3)) // immediate does not fit).
if loong64Imm12(int64(fi.autosize)) {
ws = append(ws, l64irr(l64DualTable["ADDV"].imm, fi.autosize, 3, 3))
} else {
ws = append(ws, loong64MatWords(nil, int64(fi.autosize))...)
ws = append(ws, l64rrr(l64DualTable["ADDV"].rrr, 30, 3, 3))
}
} }
// jirl r0, r1, 0 — return. // jirl r0, r1, 0, return.
ws = append(ws, l64irr16(l64branchTable["JIRL"], 0, 1, 0)) ws = append(ws, l64irr16(l64branchTable["JIRL"], 0, 1, 0))
return l64WordsLE(ws...) return l64WordsLE(ws...)
} }
// loong64StoreWords reports the prologue word count of the LR store, and
// loong64AdjustWords the word count of an SP adjust of v: the immediate
// forms when they fit, otherwise the R30 materialisation sequences.
func loong64StoreWords(autosize int) int {
if autosize > 2046 {
return 3
}
return 1
}
func loong64AdjustWords(v int64) int {
if loong64Imm12(v) {
return 1
}
return loong64MatLen(v) + 1
}
// loong64EpilogueWords reports the epilogue word count the RET expands to.
func loong64EpilogueWords(fi loong64FrameInfo) int {
n := loong64AdjustWords(int64(fi.autosize))
if !fi.leaf {
n++
}
return n
}
// loong64ResolvePseudo translates a pseudo-register memory reference into a // loong64ResolvePseudo translates a pseudo-register memory reference into a
// hardware base register and offset. x+N(FP) → (N + autosize + 8)(SP); // hardware base register and offset. x+N(FP) → (N + autosize + 8)(SP);
// x-N(SP) → (autosize - N)(SP). Returns base = -1 for an unresolvable // x-N(SP) → (autosize - N)(SP). Returns base = -1 for an unresolvable
+78 -5
View File
@@ -280,7 +280,7 @@ done:
// TestLOONG64_pcsp checks the stack-adjustment table of a framed function: // TestLOONG64_pcsp checks the stack-adjustment table of a framed function:
// the prologue raises the SP delta by autosize (in effect from the third // the prologue raises the SP delta by autosize (in effect from the third
// instruction) and the RET's epilogue restores it to zero, with the pc deltas // instruction) and the RET's epilogue restores it to zero, with the pc deltas
// in MinLC (4) units — byte-identical to `go tool asm`. // in MinLC (4) units; byte-identical to `go tool asm`.
func TestLOONG64_pcsp(t *testing.T) { func TestLOONG64_pcsp(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
@@ -347,21 +347,94 @@ TEXT ·sb(SB), NOSPLIT, $0
} }
} }
// TestLOONG64_movImmToFp checks the immediate-to-FP move forms. // TestLOONG64_movImmToFp checks the immediate-to-FP move: MOVW $c, Fd is the
// only spelling the toolchain accepts, expanding to ori (or addi.w for the
// negative span) into R30 plus movgr2fr.w. The pinned words are the
// toolchain's own bytes; the other widths and out-of-range constants are
// illegal combinations there and are diagnosed here.
func TestLOONG64_movImmToFp(t *testing.T) { func TestLOONG64_movImmToFp(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h" fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·fpmov(SB), NOSPLIT, $0 TEXT ·fpmov(SB), NOSPLIT, $0
MOVV $0x1, F0 MOVW $0x1, F0
MOVW $0x2, F4 MOVW $0x2, F4
MOVW $-1, F4
RET RET
`) `)
code := assembleLOONG64Helper(t, fn) code := assembleLOONG64Helper(t, fn)
want := []byte{ want := []byte{
0x00, 0x04, 0x80, 0x03, // ori f0, r0, 1 0x1e, 0x04, 0x80, 0x03, // ori r30, r0, 1
0x04, 0x08, 0x80, 0x03, // ori f4, r0, 2 0xc0, 0xa7, 0x14, 0x01, // movgr2fr.w f0, r30
0x1e, 0x08, 0x80, 0x03, // ori r30, r0, 2
0xc4, 0xa7, 0x14, 0x01, // movgr2fr.w f4, r30
0x1e, 0xfc, 0xbf, 0x02, // addi.w r30, r0, -1
0xc4, 0xa7, 0x14, 0x01, // movgr2fr.w f4, r30
0x20, 0x00, 0x00, 0x4c, // jirl r0, r1, 0 0x20, 0x00, 0x00, 0x4c, // jirl r0, r1, 0
} }
if !bytes.Equal(code, want) { if !bytes.Equal(code, want) {
t.Errorf("code = % x\nwant % x", code, want) t.Errorf("code = % x\nwant % x", code, want)
} }
} }
// TestLOONG64_movImmToFpErrors checks the immediate-to-FP diagnostics: the
// widths the toolchain rejects as illegal combinations, and constants beyond
// the 12-bit ori/addi.w span (the toolchain never materialises a wider
// constant on this path).
func TestLOONG64_movImmToFpErrors(t *testing.T) {
cases := []string{
"MOVV $1, F0",
"MOVF $2, F4",
"MOVD $2, F4",
"MOVW $100000, F1",
"MOVW $-2049, F1",
"MOVW $4096, F1",
}
for _, src := range cases {
fn := firstTextLOONG64(t, "#include \"textflag.h\"\nTEXT ·e(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleLOONG64(fn); err == nil {
t.Errorf("%s: expected an error, got none", src)
}
}
}
// TestLOONG64_branch16Unsigned pins the unsigned two-operand branches: with
// one register BLTU/BGEU keep the register-register form against R0 (never
// taken), the toolchain's encoding, where a beqz would test the wrong
// condition; the three-operand forms are unchanged.
func TestLOONG64_branch16Unsigned(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·u(SB), NOSPLIT, $0
BLTU R4, done
BGEU R5, done
BLTU R6, R7, done
BGEU R8, R9, done
done:
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x68001080, // bltu r4, r0, +4
0x6C000CA0, // bgeu r5, r0, +3
0x680008C7, // bltu r6, r7, +2
0x6C000509, // bgeu r8, r9, +1
0x4C000020, // jirl r0, r1, 0
)
}
// TestLOONG64_bitFieldRange checks the BSTRINS/BSTRPICK bit-number
// validation, mirroring the toolchain's "illegal bit number" rule: 0..31 for
// the .w forms, 0..63 for the .d forms, and lsb <= msb.
func TestLOONG64_bitFieldRange(t *testing.T) {
cases := []string{
"BSTRINSW $32, R4, $0, R5",
"BSTRPICKW $31, R4, $32, R5",
"BSTRINSV $64, R4, $0, R5",
"BSTRPICKV $3, R4, $4, R5",
"BSTRINSW $-1, R4, $0, R5",
}
for _, src := range cases {
fn := firstTextLOONG64(t, "#include \"textflag.h\"\nTEXT ·e(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleLOONG64(fn); err == nil {
t.Errorf("%s: expected an error, got none", src)
}
}
}
+42
View File
@@ -0,0 +1,42 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// TestLOONG64RelocOffsetsIncludePrologue pins the function-relative
// relocation offsets of a framed loong64 function: the offsets used to
// exclude the prologue, so every relocation landed on a prologue
// instruction in the GOOBJ/ELF output.
func TestLOONG64RelocOffsetsIncludePrologue(t *testing.T) {
f, errs := parser.Parse("k_loong64.s", "TEXT \u00b7f(SB), $16-0\n"+
"\tMOVV $gdata(SB), R4\n"+
"\tRET\n"+
"GLOBL gdata(SB), $8\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
fn := img.Funcs[0]
// Layout: 12-byte prologue (autosize 32), pcalau12i+addi.d (12, 16),
// epilogue with RET.
if len(fn.Relocs) != 2 {
t.Fatalf("relocs = %d, want 2", len(fn.Relocs))
}
hi, lo := fn.Relocs[0], fn.Relocs[1]
if hi.Kind != RelLoong64AddrHi || hi.Off != 12 || hi.After != 12 {
t.Errorf("hi reloc = {off %d after %d kind %d}, want {off 12 after 12 kind RelLoong64AddrHi}", hi.Off, hi.After, hi.Kind)
}
if lo.Kind != RelLoong64AddrLo || lo.Off != 16 || lo.After != 16 {
t.Errorf("lo reloc = {off %d after %d kind %d}, want {off 16 after 16 kind RelLoong64AddrLo}", lo.Off, lo.After, lo.Kind)
}
}
+9 -9
View File
@@ -12,33 +12,33 @@ import "maps"
import "strings" import "strings"
// Reg is an x86-64 register. In Plan 9 assembly the classic names (AX, BX, …) // Reg is an x86-64 register. In Plan 9 assembly the classic names (AX, BX, …)
// are size-agnostic — the instruction suffix (MOVQ vs MOVL) fixes the width — // are size-agnostic, the instruction suffix (MOVQ vs MOVL) fixes the width
// so the encoder keys off the register's index and lets the mnemonic supply the // so the encoder keys off the register's index and lets the mnemonic supply the
// size. The high flag marks the legacy high-byte registers AH/CH/DH/BH, which // size. The high flag marks the legacy high-byte registers AH/CH/DH/BH, which
// occupy indices 4–7 yet take no REX prefix, unlike SPL/BPL/SIL/DIL that share // occupy indices 4-7 yet take no REX prefix, unlike SPL/BPL/SIL/DIL that share
// those indices but require one. The mask flag marks the AVX-512 opmask // those indices but require one. The mask flag marks the AVX-512 opmask
// registers K0–K7. // registers K0-K7.
type Reg struct { type Reg struct {
idx int idx int
size int // informational width implied by the name; the mnemonic decides size int // informational width implied by the name; the mnemonic decides
high bool // AH/CH/DH/BH high bool // AH/CH/DH/BH
mask bool // K0–K7 opmask register mask bool // K0-K7 opmask register
} }
// Index returns the register number (0–15 for GPRs, 0–31 for vectors). // Index returns the register number (0-15 for GPRs, 0-31 for vectors).
func (r Reg) Index() int { return r.idx } func (r Reg) Index() int { return r.idx }
// Size returns the width in bytes implied by the register's name. // Size returns the width in bytes implied by the register's name.
func (r Reg) Size() int { return r.size } func (r Reg) Size() int { return r.size }
// IsMask reports whether r is an AVX-512 opmask register (K0–K7). // IsMask reports whether r is an AVX-512 opmask register (K0-K7).
func (r Reg) IsMask() bool { return r.mask } func (r Reg) IsMask() bool { return r.mask }
func (r Reg) isOperand() {} func (r Reg) isOperand() {}
// needsREX reports whether this register forces a REX prefix at the given // needsREX reports whether this register forces a REX prefix at the given
// operand size: the extended registers R8–R15 always do, and at byte size the // operand size: the extended registers R8-R15 always do, and at byte size the
// low registers SPL/BPL/SIL/DIL (indices 4–7, not high) do as well. // low registers SPL/BPL/SIL/DIL (indices 4-7, not high) do as well.
func (r Reg) needsREX(opSize int) bool { func (r Reg) needsREX(opSize int) bool {
if r.idx >= 8 { if r.idx >= 8 {
return true return true
@@ -133,7 +133,7 @@ func buildRegByName() map[string]Reg {
} }
// Vector: X0..X31 (128-bit, size 16), Y0..Y31 (256-bit, size 32), // Vector: X0..X31 (128-bit, size 16), Y0..Y31 (256-bit, size 32),
// Z0..Z31 (512-bit, size 64). Indices 16–31 are only encodable in EVEX // Z0..Z31 (512-bit, size 64). Indices 16-31 are only encodable in EVEX
// (AVX-512) instructions; the encoder validates that through its tables. // (AVX-512) instructions; the encoder validates that through its tables.
for i := 0; i <= 31; i++ { for i := 0; i <= 31; i++ {
m["X"+itoa(i)] = Reg{idx: i, size: 16} m["X"+itoa(i)] = Reg{idx: i, size: 16}
+393 -43
View File
@@ -15,14 +15,19 @@ import (
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) { func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
fi := riscvComputeFrame(t) fi := riscvComputeFrame(t)
prologue := riscvPrologue(fi) prologue := riscvPrologue(fi)
guardLen, err := riscvGuardLen(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
}
var relocs []Reloc var relocs []Reloc
var spadj []SpadjStep var spadj []SpadjStep
// The prologue raises the SP delta by autosize; the boundary is reported // The prologue raises the SP delta by autosize; the boundary is reported
// at the pc just past its ADDI, exactly as the toolchain's pctospadj does. // at the pc just past its ADDI, exactly as the toolchain's pctospadj does.
// The guard prefix shifts its PC.
if fi.autosize != 0 { if fi.autosize != 0 {
spadj = append(spadj, SpadjStep{PC: riscvPrologueSpadjPC(fi), Value: fi.autosize}) spadj = append(spadj, SpadjStep{PC: guardLen + riscvPrologueSpadjPC(fi), Value: fi.autosize})
} }
// Pass 1: collect instructions and compute label offsets assuming 4 bytes // Pass 1: collect instructions and compute label offsets assuming 4 bytes
@@ -34,7 +39,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
} }
var recs []instrRec var recs []instrRec
offsets := map[string]int{} offsets := map[string]int{}
pos := len(prologue) pos := guardLen + len(prologue)
for _, stmt := range t.Body { for _, stmt := range t.Body {
switch s := stmt.(type) { switch s := stmt.(type) {
case *ast.Label: case *ast.Label:
@@ -64,27 +69,38 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
} }
} }
// Pass 4: recompute offsets with actual sizes. // Pass 4: recompute offsets with actual sizes. recs holds the
// instructions in emission order, so an index into it walks t.Body in
// lockstep (the same single pass Pass 1 uses) instead of rescanning the
// whole slice per statement.
offsets = map[string]int{} offsets = map[string]int{}
pos = len(prologue) pos = guardLen + len(prologue)
ri := 0
for _, stmt := range t.Body { for _, stmt := range t.Body {
switch s := stmt.(type) { switch s := stmt.(type) {
case *ast.Label: case *ast.Label:
offsets[s.Name.Text] = pos offsets[s.Name.Text] = pos
case *ast.Instr: case *ast.Instr:
for _, r := range recs { pos += len(recs[ri].code)
if r.instr == s { ri++
pos += len(r.code)
break
}
}
} }
} }
// Pass 5: re-encode branches with corrected offsets. Record relocations // Pass 5: re-encode branches with corrected offsets. Record relocations
// during this final pass (relocation offsets are relative to instruction start). // during this final pass (relocation offsets are relative to instruction
out := append([]byte(nil), prologue...) // start). The guard prefix precedes the prologue; its branches target
pc = len(prologue) // the morestack block at the end of the function, which the previous
// passes have sized.
var out []byte
guardBytes, guardReloc, err := riscvGuard(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
}
if fi.needSplit {
out = append(out, guardBytes...)
}
out = append(out, prologue...)
pc = guardLen + len(prologue)
preCount := len(relocs) preCount := len(relocs)
var lines []LineEntry var lines []LineEntry
for _, r := range recs { for _, r := range recs {
@@ -119,9 +135,49 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
pc += len(code) pc += len(code)
} }
} }
if fi.needSplit {
relocs = append(relocs, guardReloc)
}
return out, offsets, relocs, lines, spadj, nil return out, offsets, relocs, lines, spadj, nil
} }
// riscvImmAlias maps the R-type ALU mnemonics onto their I-type immediate
// forms: the toolchain accepts ADD $imm, rj, rd and emits addi. Applied
// whenever the first operand is an immediate.
var riscvImmAlias = map[string]string{
"ADD": "ADDI",
"ADDW": "ADDIW",
"AND": "ANDI",
"OR": "ORI",
"XOR": "XORI",
"SLL": "SLLI",
"SRL": "SRLI",
"SRA": "SRAI",
"SLLW": "SLLIW",
"SRLW": "SRLIW",
"SRAW": "SRAIW",
}
// riscvNormaliseImmAlias rewrites the mnemonic to its immediate form when the
// first operand is an immediate: the toolchain accepts ADD $imm, rj, rd and
// emits addi, and SUB $imm becomes addi with the negated immediate. The
// second result reports that negation; the operand itself is left untouched
// because several passes normalise the same instruction.
func riscvNormaliseImmAlias(mnem string, ops []*ast.Operand) (string, bool) {
if len(ops) >= 2 && isImmOperand(ops[0]) {
switch strings.ToUpper(mnem) {
case "SUB":
return "ADDI", true
case "SUBW":
return "ADDIW", true
}
if alias, ok := riscvImmAlias[strings.ToUpper(mnem)]; ok {
return alias, false
}
}
return mnem, false
}
// riscvInstrSize returns the encoded size in bytes of a RISC-V instruction. // riscvInstrSize returns the encoded size in bytes of a RISC-V instruction.
// Most instructions are 4 bytes; MOV with a large immediate and I-type // Most instructions are 4 bytes; MOV with a large immediate and I-type
// arithmetic with a large immediate expand to several (possibly compressed) // arithmetic with a large immediate expand to several (possibly compressed)
@@ -129,10 +185,12 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int { func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
mnem := instr.Mnemonic.Text mnem := instr.Mnemonic.Text
ops := instr.Operands ops := instr.Operands
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
if mnem == "RET" { if mnem == "RET" {
return len(riscvReturn(fi)) return len(riscvReturn(fi))
} }
if mnem == "MOV" && len(ops) == 2 { if strings.HasPrefix(mnem, "MOV") && len(ops) == 2 {
// MOV $sym(SB), rd → 8 bytes (AUIPC + ADDI). // MOV $sym(SB), rd → 8 bytes (AUIPC + ADDI).
if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil && ops[0].Imm.Sym.Pseudo == "SB" { if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil && ops[0].Imm.Sym.Pseudo == "SB" {
return 8 return 8
@@ -149,10 +207,22 @@ func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
if isImmOperand(ops[0]) && ops[0].Imm.Sym == nil { if isImmOperand(ops[0]) && ops[0].Imm.Sym == nil {
return riscvMovImmSize(regFromOperand(ops[1]), immFromOperand(ops[0])) return riscvMovImmSize(regFromOperand(ops[1]), immFromOperand(ops[0]))
} }
// Frame-relative loads and stores: a frame offset beyond the signed
// 12-bit range materialises the address in X31 first.
if isMemOperand(ops[0]) && !isMemOperand(ops[1]) {
return riscvFrameMemSize(ops[0], fi)
}
if isMemOperand(ops[1]) && !isMemOperand(ops[0]) {
return riscvFrameMemSize(ops[1], fi)
}
} }
// I-type arithmetic with a large immediate expands to several instructions. // I-type arithmetic with a large immediate expands to several instructions.
if (mnem == "ADDI" || mnem == "ANDI" || mnem == "ORI" || mnem == "XORI") && len(ops) >= 1 && isImmOperand(ops[0]) { if (mnem == "ADDI" || mnem == "ANDI" || mnem == "ORI" || mnem == "XORI") && len(ops) >= 1 && isImmOperand(ops[0]) {
return riscvItypeImmediateSize(mnem, immFromOperand(ops[0])) imm := immFromOperand(ops[0])
if immNeg {
imm = -imm
}
return riscvItypeImmediateSize(mnem, imm)
} }
return 4 return 4
} }
@@ -167,10 +237,31 @@ func isBranchLike(mnem string) bool {
return false return false
} }
// riscvCheckBranchOffset rejects a B-type displacement outside its signed
// 13-bit span [-4096, 4094]; the encoder masks to 13 bits, so an
// out-of-range offset would otherwise wrap to a wrong target.
func riscvCheckBranchOffset(target string, off int32) error {
if off < -4096 || off > 4094 {
return fmt.Errorf("branch to %q too far (13-bit range)", target)
}
return nil
}
// riscvCheckJumpOffset rejects a J-type displacement outside its signed
// 21-bit span [-1048576, 1048574].
func riscvCheckJumpOffset(target string, off int32) error {
if off < -1048576 || off > 1048574 {
return fmt.Errorf("jump to %q too far (21-bit range)", target)
}
return nil
}
// encodeRISCVInstr encodes a single RISC-V instruction. // encodeRISCVInstr encodes a single RISC-V instruction.
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) { func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) {
mnem := instr.Mnemonic.Text mnem := instr.Mnemonic.Text
ops := instr.Operands ops := instr.Operands
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
var word uint32 var word uint32
// Handle pseudo-instructions and special cases first. // Handle pseudo-instructions and special cases first.
@@ -187,25 +278,60 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
} }
op := ops[0] op := ops[0]
if op.Addr.Sym == nil || op.Addr.Sym.Pseudo != "SB" { if op.Addr.Sym == nil || op.Addr.Sym.Pseudo != "SB" {
// CALL (X5): an indirect call, the toolchain's JALR X1, 0(X5).
if op.Addr.Sym == nil && op.Addr.Base != "" {
if op.Addr.Offset != 0 || op.Addr.Index != "" {
return nil, fmt.Errorf("CALL: invalid indirect operand %q", op.Raw)
}
rs1 := riscvRegNum(op.Addr.Base)
if rs1 < 0 {
return nil, fmt.Errorf("CALL: unknown branch register %q", op.Addr.Base)
}
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
return nil, fmt.Errorf("CALL: local branch target is not supported (use CALL sym(SB))") return nil, fmt.Errorf("CALL: local branch target is not supported (use CALL sym(SB))")
} }
if relocs != nil { if relocs != nil {
*relocs = append(*relocs, Reloc{Off: 0, After: 4, Name: op.Addr.Sym.Name, Kind: RelRISCVJal, Addend: op.Addr.Sym.Offset}) *relocs = append(*relocs, Reloc{Off: 0, After: 4, Name: op.Addr.Sym.Name, Kind: RelRISCVJal, Addend: op.Addr.Sym.Offset})
} }
word = riscvJType(1, 0) // JAL X1, 0 — the linker fills the offset word = riscvJType(1, 0) // JAL X1, 0, the linker fills the offset
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
case "JMP": case "JMP":
// JMP = JAL X0, target. The Go assembler never compresses this to // JMP = JAL X0, target. The Go assembler never compresses this to
// C.J, so always emit the 32-bit JAL. // C.J, so always emit the 32-bit JAL.
var target string var target string
if len(ops) >= 1 { if len(ops) >= 1 {
// JMP sym(SB): a tail call, JAL X0 against a symbol relocation.
if ops[0].Addr.Sym != nil && ops[0].Addr.Sym.Pseudo == "SB" {
if relocs != nil {
*relocs = append(*relocs, Reloc{Off: 0, After: 4, Name: ops[0].Addr.Sym.Name, Kind: RelRISCVJal, Addend: ops[0].Addr.Sym.Offset})
}
word = riscvJType(0, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
target = labelFromOperand(ops[0]) target = labelFromOperand(ops[0])
// JMP (X5): an indirect branch, the toolchain's JALR X0, 0(X5).
if ops[0].Addr.Sym == nil && ops[0].Addr.Base != "" {
if ops[0].Addr.Offset != 0 || ops[0].Addr.Index != "" {
return nil, fmt.Errorf("JMP: invalid indirect operand %q", ops[0].Raw)
}
rs1 := riscvRegNum(ops[0].Addr.Base)
if rs1 < 0 {
return nil, fmt.Errorf("JMP: unknown branch register %q", ops[0].Addr.Base)
}
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
} }
targetOff, ok := offsets[target] targetOff, ok := offsets[target]
if !ok { if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets)) return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
} }
offset := int32(targetOff - pc) offset := int32(targetOff - pc)
if err := riscvCheckJumpOffset(target, offset); err != nil {
return nil, err
}
word = riscvJType(0, offset) word = riscvJType(0, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
case "JAL": case "JAL":
@@ -222,25 +348,74 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets)) return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
} }
offset := int32(targetOff - pc) offset := int32(targetOff - pc)
if err := riscvCheckJumpOffset(target, offset); err != nil {
return nil, err
}
word = riscvJType(rd, offset) word = riscvJType(rd, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
// MOV is a pseudo-instruction that the Go assembler uses for loads, // MOV is a pseudo-instruction that the Go assembler uses for loads,
// stores, register moves and immediate loads. // stores, register moves and immediate loads. The width suffixes
case "MOV": // (MOVB/MOVH/MOVW and unsigned forms) select the access width, and
// MOVD/MOVF address the FP registers.
case "MOV", "MOVB", "MOVBU", "MOVH", "MOVHU", "MOVW", "MOVWU", "MOVF", "MOVD":
return encodeRISCVMov(instr, fi, relocs) return encodeRISCVMov(instr, fi, relocs)
// JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1). // JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1).
case "JALR": case "JALR":
return encodeRISCVJALR(instr, fi) return encodeRISCVJALR(instr, fi)
// Branch-zero pseudos: BEQZ/BNEZ compare against X0, and BLTZ/BGEZ/
// BLEZ/BGTZ reorder the register operands of BLT/BGE accordingly.
case "BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ":
if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
}
rs := regFromOperand(ops[0])
if rs < 0 {
return nil, fmt.Errorf("%s: invalid register", mnem)
}
target := labelFromOperand(ops[1])
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
}
var enc riscvEnc
rs1, rs2 := rs, 0
switch mnem {
case "BEQZ":
enc = riscvEnc{0x63, 0x0, 0x00} // beq rs, x0
case "BNEZ":
enc = riscvEnc{0x63, 0x1, 0x00} // bne rs, x0
case "BLTZ":
enc = riscvEnc{0x63, 0x4, 0x00} // blt rs, x0
case "BGEZ":
enc = riscvEnc{0x63, 0x5, 0x00} // bge rs, x0
case "BLEZ":
enc, rs1, rs2 = riscvEnc{0x63, 0x5, 0x00}, 0, rs // bge x0, rs
case "BGTZ":
enc, rs1, rs2 = riscvEnc{0x63, 0x4, 0x00}, 0, rs // blt x0, rs
}
if err := riscvCheckBranchOffset(target, int32(targetOff-pc)); err != nil {
return nil, err
}
word = riscvBType(enc, rs1, rs2, int32(targetOff-pc))
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
// System instructions with no operands. // System instructions with no operands.
case "FENCE", "ECALL", "EBREAK": case "FENCE", "ECALL", "EBREAK":
enc, ok := riscvInstrTable[mnem] enc, ok := riscvInstrTable[mnem]
if !ok { if !ok {
return nil, fmt.Errorf("unsupported system instruction %q", mnem) return nil, fmt.Errorf("unsupported system instruction %q", mnem)
} }
word = riscvIType(enc, 0, 0, 0) // The bare FENCE expands to fence iorw, iorw: the predecessor and
// successor fields both carry 0xF in the I-type immediate
// (the toolchain's encodeFenceOperand TYPE_NONE default).
imm := int32(0)
if mnem == "FENCE" {
imm = 0x0FF
}
word = riscvIType(enc, 0, 0, imm)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
@@ -282,7 +457,10 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
} }
csr := immFromOperand(ops[0]) // CSR address (12-bit) csr := immFromOperand(ops[0]) // CSR address (12-bit)
rd := regFromOperand(ops[2]) // destination register if csr < 0 || csr > 0xFFF {
return nil, fmt.Errorf("%s: CSR address %d out of range 0-0xFFF", mnem, csr)
}
rd := regFromOperand(ops[2]) // destination register
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("invalid destination register in %s", mnem) return nil, fmt.Errorf("invalid destination register in %s", mnem)
} }
@@ -395,7 +573,7 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
} }
word = riscvSType(enc, rs1, rs2, imm) word = riscvSType(enc, rs1, rs2, imm)
// LR (load-reserved): INSTR (addr), dst — 2 operands. // LR (load-reserved): INSTR (addr), dst, 2 operands.
case len(ops) == 2 && isLRInstr(mnem): case len(ops) == 2 && isLRInstr(mnem):
rs1, _ := memFromOperandWithFrame(ops[0], fi) rs1, _ := memFromOperandWithFrame(ops[0], fi)
rd := regFromOperand(ops[1]) rd := regFromOperand(ops[1])
@@ -404,7 +582,7 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
} }
word = riscvAMOType(enc, rd, rs1, 0) // rs2=0 for LR word = riscvAMOType(enc, rd, rs1, 0) // rs2=0 for LR
// SC (store-conditional): INSTR src, (addr), dst — 3 operands. // SC (store-conditional): INSTR src, (addr), dst, 3 operands.
case len(ops) == 3 && isSCInstr(mnem): case len(ops) == 3 && isSCInstr(mnem):
rs2 := regFromOperand(ops[0]) rs2 := regFromOperand(ops[0])
rs1, _ := memFromOperandWithFrame(ops[1], fi) rs1, _ := memFromOperandWithFrame(ops[1], fi)
@@ -427,7 +605,10 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// I-type with immediate: Plan 9 order is INSTR $imm, rs1, rd; the // I-type with immediate: Plan 9 order is INSTR $imm, rs1, rd; the
// two-operand form INSTR $imm, rd uses rd as the source. // two-operand form INSTR $imm, rd uses rd as the source.
case len(ops) == 3 && isITypeInstr(mnem): case len(ops) == 3 && isITypeInstr(mnem):
imm := immFromOperand(ops[0]) // immediate imm, err := riscvImm32FromOperand(ops[0], immNeg) // immediate
if err != nil {
return nil, err
}
rs1 := regFromOperand(ops[1]) // source register rs1 := regFromOperand(ops[1]) // source register
rd := regFromOperand(ops[2]) // destination rd := regFromOperand(ops[2]) // destination
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
@@ -436,14 +617,17 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return encodeRISCVItypeImmediate(mnem, enc, rd, rs1, imm) return encodeRISCVItypeImmediate(mnem, enc, rd, rs1, imm)
case len(ops) == 2 && isITypeInstr(mnem): case len(ops) == 2 && isITypeInstr(mnem):
imm := immFromOperand(ops[0]) imm, err := riscvImm32FromOperand(ops[0], immNeg)
if err != nil {
return nil, err
}
rd := regFromOperand(ops[1]) rd := regFromOperand(ops[1])
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("invalid register in %s", mnem) return nil, fmt.Errorf("invalid register in %s", mnem)
} }
return encodeRISCVItypeImmediate(mnem, enc, rd, rd, imm) return encodeRISCVItypeImmediate(mnem, enc, rd, rd, imm)
// Loads: rd, offset(rs1) — Plan 9 order is LD src, dst. // Loads: rd, offset(rs1), Plan 9 order is LD src, dst.
case len(ops) == 2 && isLoadInstr(mnem): case len(ops) == 2 && isLoadInstr(mnem):
rd := regFromOperand(ops[1]) // destination (last operand) rd := regFromOperand(ops[1]) // destination (last operand)
rs1, imm := memFromOperandWithFrame(ops[0], fi) // memory source (first operand) rs1, imm := memFromOperandWithFrame(ops[0], fi) // memory source (first operand)
@@ -474,6 +658,9 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
if rs1 < 0 || rs2 < 0 { if rs1 < 0 || rs2 < 0 {
return nil, fmt.Errorf("invalid register in %s", mnem) return nil, fmt.Errorf("invalid register in %s", mnem)
} }
if err := riscvCheckBranchOffset(target, offset); err != nil {
return nil, err
}
// The Go assembler never compresses branches to C.BEQZ/C.BNEZ. // The Go assembler never compresses branches to C.BEQZ/C.BNEZ.
word = riscvBType(enc, rs1, rs2, offset) word = riscvBType(enc, rs1, rs2, offset)
@@ -538,7 +725,7 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
// Immediate → register. // Immediate → register.
if isImmOperand(src) { if isImmOperand(src) {
// MOV $sym(SB), rd — load address of a static symbol or external. // MOV $sym(SB), rd, load address of a static symbol or external.
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" { if src.Imm.Sym != nil && src.Imm.Sym.Pseudo == "SB" {
rd := regFromOperand(dst) rd := regFromOperand(dst)
if rd < 0 { if rd < 0 {
@@ -546,7 +733,7 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
} }
return encodeRISCVSBAddr(src.Imm.Sym, rd, relocs), nil return encodeRISCVSBAddr(src.Imm.Sym, rd, relocs), nil
} }
// MOV $sym(FP/SP), rd — not supported: immediate symbol references // MOV $sym(FP/SP), rd, not supported: immediate symbol references
// other than SB cannot be encoded as a simple immediate. // other than SB cannot be encoded as a simple immediate.
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo != "" { if src.Imm.Sym != nil && src.Imm.Sym.Pseudo != "" {
return nil, fmt.Errorf("MOV $%s(%s): unsupported immediate symbol reference (only SB is supported)", src.Imm.Sym.Name, src.Imm.Sym.Pseudo) return nil, fmt.Errorf("MOV $%s(%s): unsupported immediate symbol reference (only SB is supported)", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
@@ -555,14 +742,17 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("MOV $imm: invalid destination register") return nil, fmt.Errorf("MOV $imm: invalid destination register")
} }
imm := immFromOperand(src) imm, err := riscvImm32FromOperand(src, false)
if err != nil {
return nil, err
}
return encodeRISCVLoadImm(rd, imm), nil return encodeRISCVLoadImm(rd, imm), nil
} }
// Memory → register (load). // Memory → register (load).
if isMemOperand(src) && !isMemOperand(dst) { if isMemOperand(src) && !isMemOperand(dst) {
rd := regFromOperand(dst) rd := regFromOperand(dst)
// MOV sym(SB), rd — load from static data. // MOV sym(SB), rd, load from static data.
if src.Addr.Sym != nil && src.Addr.Sym.Pseudo == "SB" { if src.Addr.Sym != nil && src.Addr.Sym.Pseudo == "SB" {
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("MOV sym(SB): invalid destination register") return nil, fmt.Errorf("MOV sym(SB): invalid destination register")
@@ -573,14 +763,13 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV load: invalid operand") return nil, fmt.Errorf("MOV load: invalid operand")
} }
word := riscvIType(riscvEnc{0x03, 0x3, 0x00}, rd, rs1, off) return riscvFrameMemOp(riscvMovEnc(strings.ToUpper(instr.Mnemonic.Text), false), false, rd, rs1, off), nil
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
// Register → memory (store). // Register → memory (store).
if !isMemOperand(src) && isMemOperand(dst) { if !isMemOperand(src) && isMemOperand(dst) {
rs2 := regFromOperand(src) rs2 := regFromOperand(src)
// MOV rd, sym(SB) — store to static data. // MOV rd, sym(SB), store to static data.
if dst.Addr.Sym != nil && dst.Addr.Sym.Pseudo == "SB" { if dst.Addr.Sym != nil && dst.Addr.Sym.Pseudo == "SB" {
if rs2 < 0 { if rs2 < 0 {
return nil, fmt.Errorf("MOV rd, sym(SB): invalid source register") return nil, fmt.Errorf("MOV rd, sym(SB): invalid source register")
@@ -591,22 +780,108 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rs2 < 0 || rs1 < 0 { if rs2 < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV store: invalid operand") return nil, fmt.Errorf("MOV store: invalid operand")
} }
word := riscvSType(riscvEnc{0x23, 0x3, 0x00}, rs1, rs2, off) return riscvFrameMemOp(riscvMovEnc(strings.ToUpper(instr.Mnemonic.Text), true), true, rs2, rs1, off), nil
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
// Register → register (ADDI $0, src, dst). // Register → register: MOVD/MOVF are FP moves (fsgnj with rs2 = rs1),
// everything else is ADDI $0, src, dst.
{ {
rs1 := regFromOperand(src) rs1 := regFromOperand(src)
rd := regFromOperand(dst) rd := regFromOperand(dst)
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV: invalid register operand") return nil, fmt.Errorf("MOV: invalid register operand")
} }
mnem := strings.ToUpper(instr.Mnemonic.Text)
if mnem == "MOVD" || mnem == "MOVF" {
op := uint32(0x20000053) // FSGNJ.S
if mnem == "MOVD" {
op = 0x22000053 // FSGNJ.D
}
return wordLE(op | uint32(rs1)<<15 | uint32(rs1)<<20 | uint32(rd)<<7), nil
}
word := riscvIType(riscvEnc{0x13, 0x0, 0x00}, rd, rs1, 0) word := riscvIType(riscvEnc{0x13, 0x0, 0x00}, rd, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
} }
// riscvMovEnc returns the load (store=false) or store (store=true) opcode for
// a MOV-family mnemonic: the suffix selects the access width, MOVD and MOVF
// select the FP load/store opcodes, and bare MOV is the 64-bit integer form.
func riscvMovEnc(mnem string, store bool) riscvEnc {
if store {
switch mnem {
case "MOVB":
return riscvEnc{0x23, 0x0, 0x00} // SB
case "MOVH":
return riscvEnc{0x23, 0x1, 0x00} // SH
case "MOVW":
return riscvEnc{0x23, 0x2, 0x00} // SW
case "MOVF":
return riscvEnc{0x27, 0x2, 0x00} // FSW
case "MOVD":
return riscvEnc{0x27, 0x3, 0x00} // FSD
}
return riscvEnc{0x23, 0x3, 0x00} // SD
}
switch mnem {
case "MOVB":
return riscvEnc{0x03, 0x0, 0x00} // LB
case "MOVBU":
return riscvEnc{0x03, 0x4, 0x00} // LBU
case "MOVH":
return riscvEnc{0x03, 0x1, 0x00} // LH
case "MOVHU":
return riscvEnc{0x03, 0x5, 0x00} // LHU
case "MOVW":
return riscvEnc{0x03, 0x2, 0x00} // LW
case "MOVWU":
return riscvEnc{0x03, 0x6, 0x00} // LWU
case "MOVF":
return riscvEnc{0x07, 0x2, 0x00} // FLW
case "MOVD":
return riscvEnc{0x07, 0x3, 0x00} // FLD
}
return riscvEnc{0x03, 0x3, 0x00} // LD
}
// riscvFrameMemOp encodes a register-relative load (store=false, I-type
// width 0x03) or store (store=true, S-type width 0x23) of the 64-bit width
// at off(rs1). Offsets beyond the signed 12-bit range materialise the
// address in X31 first: LUI hi (the rounding split), then ADD X31, rs1,
// matching the toolchain's large-frame addressing; the access uses the
// sign-extended low part, which always fits.
func riscvFrameMemOp(enc riscvEnc, store bool, reg, rs1 int, off int32) []byte {
if fits12(off) {
var word uint32
if store {
word = riscvSType(enc, rs1, reg, off)
} else {
word = riscvIType(enc, reg, rs1, off)
}
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
}
lo := off - (splitHi(off) << 12)
out := riscvAddressInX31WithBase(off, rs1)
var word uint32
if store {
word = riscvSType(enc, 31, reg, lo)
} else {
word = riscvIType(enc, reg, 31, lo)
}
return append(out, wordLE(word)...)
}
// riscvFrameMemSize returns the encoded size of a frame-relative MOV for the
// layout pass: 4 bytes when the offset fits, otherwise the X31
// materialisation plus the access.
func riscvFrameMemSize(op *ast.Operand, fi riscvFrameInfo) int {
rs1, off := memFromOperandWithFrame(op, fi)
if fits12(off) {
return 4
}
return len(riscvAddressInX31WithBase(off, rs1)) + 4
}
// encodeRISCVLoadImm encodes loading an immediate into a register (MOV $imm, // encodeRISCVLoadImm encodes loading an immediate into a register (MOV $imm,
// rd), matching the toolchain's instructionsForMOVConst. For 12-bit // rd), matching the toolchain's instructionsForMOVConst. For 12-bit
// immediates it emits ADDI $imm, ZERO, rd (compressed to C.LI when it fits // immediates it emits ADDI $imm, ZERO, rd (compressed to C.LI when it fits
@@ -821,25 +1096,34 @@ func word16(w uint16) []byte {
} }
// encodeRISCVJALR encodes the JALR indirect jump/call instruction. // encodeRISCVJALR encodes the JALR indirect jump/call instruction.
// Plan 9: JALR rs1, rd (2 regs) or JALR offset(rs1) (memory → rd=X1). // Plan 9: JALR rs1, rd (2 regs), JALR rd, offset(rs1) (the trampoline
// form), or JALR offset(rs1) (memory → rd=X1).
func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) { func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
ops := instr.Operands ops := instr.Operands
// JALR rd, offset(rs1): the memory operand's base is the jump-target
// register, not the destination.
if len(ops) == 2 && isMemOperand(ops[1]) {
rd := regFromOperand(ops[0])
rs1, imm := memFromOperandWithFrame(ops[1], fi)
if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid register operand")
}
return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, imm)), nil
}
if len(ops) == 2 { if len(ops) == 2 {
rs1 := regFromOperand(ops[0]) rs1 := regFromOperand(ops[0])
rd := regFromOperand(ops[1]) rd := regFromOperand(ops[1])
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid register operand") return nil, fmt.Errorf("JALR: invalid register operand")
} }
word := riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, 0) return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, 0)), nil
return wordLE(word), nil
} }
if len(ops) == 1 { if len(ops) == 1 {
rs1, imm := memFromOperandWithFrame(ops[0], fi) rs1, imm := memFromOperandWithFrame(ops[0], fi)
if rs1 < 0 { if rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid memory operand") return nil, fmt.Errorf("JALR: invalid memory operand")
} }
word := riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, imm) return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, imm)), nil
return wordLE(word), nil
} }
return nil, fmt.Errorf("JALR expects 1 or 2 operands, got %d", len(ops)) return nil, fmt.Errorf("JALR expects 1 or 2 operands, got %d", len(ops))
} }
@@ -847,7 +1131,7 @@ func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
// tryCompressRVC attempts to compress a RISC-V instruction to its 16-bit // tryCompressRVC attempts to compress a RISC-V instruction to its 16-bit
// RVC form. It returns the compressed instruction word and true on success. // RVC form. It returns the compressed instruction word and true on success.
func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) { func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
mnem := instr.Mnemonic.Text mnem := riscvCompressMnem(instr)
ops := instr.Operands ops := instr.Operands
switch mnem { switch mnem {
@@ -1006,7 +1290,7 @@ func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
} }
case "ADDW", "SUBW": case "ADDW", "SUBW":
// C.ADDW (0x27,1) / C.SUBW (0x27,0) — CA-type, prime regs. // C.ADDW (0x27,1) / C.SUBW (0x27,0), CA-type, prime regs.
if len(ops) == 3 { if len(ops) == 3 {
funct2 := uint32(0x0) funct2 := uint32(0x0)
if mnem == "ADDW" { if mnem == "ADDW" {
@@ -1097,6 +1381,49 @@ func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
return 0, false return 0, false
} }
// riscvCompressMnem maps a MOV-family load or store onto the base mnemonic
// the toolchain lowers it to (MOVW 4(SP), X9 is LW under another name), so
// the width spellings compress exactly like their base forms. Register and
// immediate forms keep their own mnemonic: the C.MV path matches "MOV" and
// nothing else in the switch has a width case.
func riscvCompressMnem(instr *ast.Instr) string {
mnem := instr.Mnemonic.Text
ops := instr.Operands
if !strings.HasPrefix(mnem, "MOV") || len(ops) != 2 {
return mnem
}
load := isMemOperand(ops[0]) && !isMemOperand(ops[1])
store := !isMemOperand(ops[0]) && isMemOperand(ops[1])
if !load && !store {
return mnem
}
switch mnem {
case "MOVW":
if load {
return "LW"
}
return "SW"
case "MOVF":
if load {
return "FLW"
}
return "FSW"
case "MOVD":
if load {
return "FLD"
}
return "FSD"
case "MOV":
if load {
return "LD"
}
return "SD"
}
// MOVB/MOVBU/MOVH/MOVHU/MOVWU have no compressed form; their base
// mnemonics (LB/LBU/LH/LHU/LWU, SB/SH) match no case either.
return mnem
}
// extractLDParams extracts rd, rs1, and immediate offset for a load instruction. // extractLDParams extracts rd, rs1, and immediate offset for a load instruction.
func extractLDParams(instr *ast.Instr, fi riscvFrameInfo) (rd, rs1 int, imm int32) { func extractLDParams(instr *ast.Instr, fi riscvFrameInfo) (rd, rs1 int, imm int32) {
ops := instr.Operands ops := instr.Operands
@@ -1273,6 +1600,29 @@ func immFromOperand(op *ast.Operand) int32 {
return 0 return 0
} }
// riscvImm32FromOperand reads an immediate for the MOV/I-type paths as a
// signed 32-bit value. The toolchain materialises wider constants through
// its SLLI expansion, which this assembler does not implement, so values
// outside the int32 span are diagnosed instead of silently truncated (MOV
// $0x123456789 must not assemble as $0x3456789). The neg flag carries the
// SUB $imm alias, whose negated value may fit when the written one does not.
func riscvImm32FromOperand(op *ast.Operand, neg bool) (int32, error) {
var v int64
if op.Imm.HasVal {
v = op.Imm.Val
if op.Imm.Neg {
v = -v
}
}
if neg {
v = -v
}
if int64(int32(v)) != v {
return 0, fmt.Errorf("immediate %d out of range; 64-bit materialisation not supported", v)
}
return int32(v), nil
}
func memFromOperand(op *ast.Operand) (rs1 int, imm int32) { func memFromOperand(op *ast.Operand) (rs1 int, imm int32) {
rs1 = riscvRegNum(op.Addr.Base) rs1 = riscvRegNum(op.Addr.Base)
imm = int32(op.Addr.Offset) imm = int32(op.Addr.Offset)
@@ -1314,7 +1664,7 @@ func suggestLabel(target string, offsets map[string]int) string {
} }
// Only suggest if the distance is small enough. // Only suggest if the distance is small enough.
if bestDist <= 3 && bestDist < len(target)/2+1 { if bestDist <= 3 && bestDist < len(target)/2+1 {
return fmt.Sprintf(" — did you mean %q?", best) return fmt.Sprintf("; did you mean %q?", best)
} }
return "" return ""
} }
+23 -19
View File
@@ -65,7 +65,7 @@ func riscvRegNum(name string) int {
return 25 return 25
case "X26", "S10": case "X26", "S10":
return 26 return 26
case "X27", "S11": case "X27", "S11", "g":
return 27 return 27
case "X28", "T3": case "X28", "T3":
return 28 return 28
@@ -154,7 +154,7 @@ type riscvEnc struct {
// riscvInstrTable maps RISC-V mnemonics to their encoding. // riscvInstrTable maps RISC-V mnemonics to their encoding.
var riscvInstrTable = map[string]riscvEnc{ var riscvInstrTable = map[string]riscvEnc{
// RV64I — R-type arithmetic/logic. // RV64I, R-type arithmetic/logic.
"ADD": {0x33, 0x0, 0x00}, "ADD": {0x33, 0x0, 0x00},
"SUB": {0x33, 0x0, 0x20}, "SUB": {0x33, 0x0, 0x20},
"SLL": {0x33, 0x1, 0x00}, "SLL": {0x33, 0x1, 0x00},
@@ -165,20 +165,20 @@ var riscvInstrTable = map[string]riscvEnc{
"SRA": {0x33, 0x5, 0x20}, "SRA": {0x33, 0x5, 0x20},
"OR": {0x33, 0x6, 0x00}, "OR": {0x33, 0x6, 0x00},
"AND": {0x33, 0x7, 0x00}, "AND": {0x33, 0x7, 0x00},
// RV64I — 32-bit variants (W suffix). // RV64I, 32-bit variants (W suffix).
"ADDW": {0x3B, 0x0, 0x00}, "ADDW": {0x3B, 0x0, 0x00},
"SUBW": {0x3B, 0x0, 0x20}, "SUBW": {0x3B, 0x0, 0x20},
"SLLW": {0x3B, 0x1, 0x00}, "SLLW": {0x3B, 0x1, 0x00},
"SRLW": {0x3B, 0x5, 0x00}, "SRLW": {0x3B, 0x5, 0x00},
"SRAW": {0x3B, 0x5, 0x20}, "SRAW": {0x3B, 0x5, 0x20},
// RV64I — I-type shift-immediate (shamt in rs2 field). // RV64I, I-type shift-immediate (shamt in rs2 field).
"SLLI": {0x13, 0x1, 0x00}, "SLLI": {0x13, 0x1, 0x00},
"SRLI": {0x13, 0x5, 0x00}, "SRLI": {0x13, 0x5, 0x00},
"SRAI": {0x13, 0x5, 0x20}, "SRAI": {0x13, 0x5, 0x20},
"SLLIW": {0x1B, 0x1, 0x00}, "SLLIW": {0x1B, 0x1, 0x00},
"SRLIW": {0x1B, 0x5, 0x00}, "SRLIW": {0x1B, 0x5, 0x00},
"SRAIW": {0x1B, 0x5, 0x20}, "SRAIW": {0x1B, 0x5, 0x20},
// RV64M — multiply/divide. // RV64M, multiply/divide.
"MUL": {0x33, 0x0, 0x01}, "MUL": {0x33, 0x0, 0x01},
"MULH": {0x33, 0x1, 0x01}, "MULH": {0x33, 0x1, 0x01},
"MULHSU": {0x33, 0x2, 0x01}, "MULHSU": {0x33, 0x2, 0x01},
@@ -187,13 +187,13 @@ var riscvInstrTable = map[string]riscvEnc{
"DIVU": {0x33, 0x5, 0x01}, "DIVU": {0x33, 0x5, 0x01},
"REM": {0x33, 0x6, 0x01}, "REM": {0x33, 0x6, 0x01},
"REMU": {0x33, 0x7, 0x01}, "REMU": {0x33, 0x7, 0x01},
// RV64M — 32-bit variants. // RV64M, 32-bit variants.
"MULW": {0x3B, 0x0, 0x01}, "MULW": {0x3B, 0x0, 0x01},
"DIVW": {0x3B, 0x4, 0x01}, "DIVW": {0x3B, 0x4, 0x01},
"DIVUW": {0x3B, 0x5, 0x01}, "DIVUW": {0x3B, 0x5, 0x01},
"REMW": {0x3B, 0x6, 0x01}, "REMW": {0x3B, 0x6, 0x01},
"REMUW": {0x3B, 0x7, 0x01}, "REMUW": {0x3B, 0x7, 0x01},
// RV64I — I-type arithmetic. // RV64I, I-type arithmetic.
"ADDI": {0x13, 0x0, 0x00}, "ADDI": {0x13, 0x0, 0x00},
"ADDIW": {0x1B, 0x0, 0x00}, "ADDIW": {0x1B, 0x0, 0x00},
"SLTI": {0x13, 0x2, 0x00}, "SLTI": {0x13, 0x2, 0x00},
@@ -228,10 +228,10 @@ var riscvInstrTable = map[string]riscvEnc{
"ECALL": {0x73, 0x0, 0x00}, "ECALL": {0x73, 0x0, 0x00},
"EBREAK": {0x73, 0x0, 0x00}, "EBREAK": {0x73, 0x0, 0x00},
"FENCE": {0x0F, 0x0, 0x00}, "FENCE": {0x0F, 0x0, 0x00},
// JALR — indirect jump/call (I-type). // JALR, indirect jump/call (I-type).
"JALR": {0x67, 0x0, 0x00}, "JALR": {0x67, 0x0, 0x00},
// RV64A — atomics (AMO opcode 0x2F). // RV64A, atomics (AMO opcode 0x2F).
// funct3: 0x2 = word, 0x3 = doubleword. funct5 in bits [31:27]. // funct3: 0x2 = word, 0x3 = doubleword. funct5 in bits [31:27].
"AMOSWAPW": {0x2F, 0x2, 0x01 << 2}, "AMOSWAPW": {0x2F, 0x2, 0x01 << 2},
"AMOSWAPD": {0x2F, 0x3, 0x01 << 2}, "AMOSWAPD": {0x2F, 0x3, 0x01 << 2},
@@ -252,7 +252,7 @@ var riscvInstrTable = map[string]riscvEnc{
"AMOMINUW": {0x2F, 0x2, 0x18 << 2}, "AMOMINUW": {0x2F, 0x2, 0x18 << 2},
"AMOMINUD": {0x2F, 0x3, 0x18 << 2}, "AMOMINUD": {0x2F, 0x3, 0x18 << 2},
// RV64F/D — floating-point arithmetic. // RV64F/D, floating-point arithmetic.
"FADDS": {0x53, 0x0, 0x00}, "FADDS": {0x53, 0x0, 0x00},
"FSUBS": {0x53, 0x0, 0x04}, "FSUBS": {0x53, 0x0, 0x04},
"FMULS": {0x53, 0x0, 0x08}, "FMULS": {0x53, 0x0, 0x08},
@@ -274,13 +274,13 @@ var riscvInstrTable = map[string]riscvEnc{
"FMIND": {0x53, 0x0, 0x15}, "FMIND": {0x53, 0x0, 0x15},
"FMAXD": {0x53, 0x1, 0x15}, "FMAXD": {0x53, 0x1, 0x15},
// RV64A — load-reserved / store-conditional (funct5 0x02 / 0x03). // RV64A, load-reserved / store-conditional (funct5 0x02 / 0x03).
"LRW": {0x2F, 0x2, 0x02 << 2}, "LRW": {0x2F, 0x2, 0x02 << 2},
"LRD": {0x2F, 0x3, 0x02 << 2}, "LRD": {0x2F, 0x3, 0x02 << 2},
"SCW": {0x2F, 0x2, 0x03 << 2}, "SCW": {0x2F, 0x2, 0x03 << 2},
"SCD": {0x2F, 0x3, 0x03 << 2}, "SCD": {0x2F, 0x3, 0x03 << 2},
// FP compare — result in integer register (funct7 0x50/0x51). // FP compare, result in integer register (funct7 0x50/0x51).
"FEQS": {0x53, 0x2, 0x50}, "FEQS": {0x53, 0x2, 0x50},
"FLTS": {0x53, 0x1, 0x50}, "FLTS": {0x53, 0x1, 0x50},
"FLES": {0x53, 0x0, 0x50}, "FLES": {0x53, 0x0, 0x50},
@@ -328,6 +328,8 @@ var riscvCvtTable = map[string]riscvCvtEnc{
"FCVTSWU": {0x68, 0x1, 0x53}, // uint32 → float32 "FCVTSWU": {0x68, 0x1, 0x53}, // uint32 → float32
"FCVTSL": {0x68, 0x2, 0x53}, // int64 → float32 "FCVTSL": {0x68, 0x2, 0x53}, // int64 → float32
"FCVTSLU": {0x68, 0x3, 0x53}, // uint64 → float32 "FCVTSLU": {0x68, 0x3, 0x53}, // uint64 → float32
"FCLASSS": {0x70, 0x0, 0x53}, // classify float32 → GPR mask
"FCLASSD": {0x70, 0x0, 0x53}, // classify float64 → GPR mask
"FCVTDW": {0x69, 0x0, 0x53}, // int32 → float64 "FCVTDW": {0x69, 0x0, 0x53}, // int32 → float64
"FCVTDWU": {0x69, 0x1, 0x53}, // uint32 → float64 "FCVTDWU": {0x69, 0x1, 0x53}, // uint32 → float64
"FCVTDL": {0x69, 0x2, 0x53}, // int64 → float64 "FCVTDL": {0x69, 0x2, 0x53}, // int64 → float64
@@ -371,7 +373,7 @@ var riscvFmaTable = map[string]riscvFmaEnc{
// riscvFmaType encodes an R4-type fused multiply-add instruction. // riscvFmaType encodes an R4-type fused multiply-add instruction.
func riscvFmaType(enc riscvFmaEnc, rd, rs1, rs2, rs3 int) uint32 { func riscvFmaType(enc riscvFmaEnc, rd, rs1, rs2, rs3 int) uint32 {
return (uint32(rs3) << 27) | (enc.fmt << 25) | (uint32(rs2) << 20) | return (uint32(rs3) << 27) | (enc.fmt << 25) | (uint32(rs2) << 20) |
(uint32(rs1) << 15) | (0x0 << 12) /* rm=dynamic */ | (uint32(rd) << 7) | enc.opcode (uint32(rs1) << 15) | (0x0 << 12) /* rm=RNE */ | (uint32(rd) << 7) | enc.opcode
} }
// CSR (Control and Status Register) instructions. // CSR (Control and Status Register) instructions.
@@ -442,10 +444,10 @@ func riscvJType(rd int, offset int32) uint32 {
// ---- RVC (compressed) encoding helpers ---- // ---- RVC (compressed) encoding helpers ----
// isRVCIntReg reports whether a register number can be encoded in the 3-bit // isRVCIntReg reports whether a register number can be encoded in the 3-bit
// prime register field used by compressed instructions (x8–x15). // prime register field used by compressed instructions (x8-x15).
func isRVCIntReg(r int) bool { return r >= 8 && r <= 15 } func isRVCIntReg(r int) bool { return r >= 8 && r <= 15 }
// rvcReg3 returns the 3-bit encoding for registers x8–x15 (0–7). // rvcReg3 returns the 3-bit encoding for registers x8-x15 (0-7).
func rvcReg3(r int) uint32 { return uint32(r - 8) } func rvcReg3(r int) uint32 { return uint32(r - 8) }
// rvcCR encodes a CR-type (register) compressed instruction. // rvcCR encodes a CR-type (register) compressed instruction.
@@ -455,7 +457,7 @@ func rvcCR(funct4, rd, rs2 uint32) uint16 {
} }
// rvcCI encodes a CI-type (immediate) compressed instruction. // rvcCI encodes a CI-type (immediate) compressed instruction.
// Used for C.ADDI, C.LI, C.LUI, C.ADDIW — linear 6-bit immediate. // Used for C.ADDI, C.LI, C.LUI, C.ADDIW, linear 6-bit immediate.
func rvcCI(funct3, rd uint32, imm uint32) uint16 { func rvcCI(funct3, rd uint32, imm uint32) uint16 {
return uint16((funct3 << 13) | ((imm>>5)&1)<<12 | (rd << 7) | (imm&0x1F)<<2 | 0x1) return uint16((funct3 << 13) | ((imm>>5)&1)<<12 | (rd << 7) | (imm&0x1F)<<2 | 0x1)
} }
@@ -520,11 +522,13 @@ func rvcCL(funct3, rd, rs1 uint32, imm uint32) uint16 {
// rvcCS encodes a register-relative compressed store (op=00 quadrant): C.SW // rvcCS encodes a register-relative compressed store (op=00 quadrant): C.SW
// (funct3=6), C.SD (funct3=7) or C.FSD (funct3=5). imm is the full byte // (funct3=6), C.SD (funct3=7) or C.FSD (funct3=5). imm is the full byte
// offset; the immediate bits are extracted per the RISC-V CS format. // offset; the immediate bits are extracted per the RISC-V CS format, with the
// same five-bit patterns as the load side ({5,4,3,7,6} and {5,4,3,2,6},
// matching the toolchain's encodeCS).
func rvcCS(funct3, rs2, rs1 uint32, imm uint32) uint16 { func rvcCS(funct3, rs2, rs1 uint32, imm uint32) uint16 {
pattern := []int{5, 3, 7, 6} pattern := []int{5, 4, 3, 7, 6}
if funct3 == 0x6 { if funct3 == 0x6 {
pattern = []int{5, 3, 2, 6} pattern = []int{5, 4, 3, 2, 6}
} }
packed := encodeRVCPattern(imm, pattern) packed := encodeRVCPattern(imm, pattern)
return uint16((funct3 << 13) | ((packed>>2)&0x7)<<10 | (rs1 << 7) | ((packed & 0x3) << 5) | (rs2 << 2)) return uint16((funct3 << 13) | ((packed>>2)&0x7)<<10 | (rs1 << 7) | ((packed & 0x3) << 5) | (rs2 << 2))
+211 -1
View File
@@ -5,6 +5,7 @@ package asm
import ( import (
"bytes" "bytes"
"strings"
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast" "sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -289,7 +290,7 @@ TEXT ·cmp(SB), NOSPLIT, $0
} }
func TestRISCV_forwardBranch(t *testing.T) { func TestRISCV_forwardBranch(t *testing.T) {
// Forward label reference — must not fail. // Forward label reference; must not fail.
fn := firstTextRISCV(t, `#include "textflag.h" fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·fwd(SB), NOSPLIT, $0 TEXT ·fwd(SB), NOSPLIT, $0
ADDI $1, X10, X10 ADDI $1, X10, X10
@@ -691,6 +692,81 @@ DATA answer<>+0(SB)/8, $42
} }
} }
// TestRISCV_RVC_StorePatterns pins the register-relative compressed store
// encodings for offsets with immediate bits 4 and 5 set, byte-identical to
// the toolchain's encodeCS (patterns {5,4,3,7,6} and {5,4,3,2,6}).
// Regression: the store-side patterns dropped imm[4], so every such store
// silently encoded the wrong address while the loads stayed correct.
func TestRISCV_RVC_StorePatterns(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·csstores(SB), NOSPLIT, $0
SD X9, 24(X8)
SW X10, 16(X11)
FSD F8, 40(X12)
LD 24(X8), X9
LW 16(X11), X10
FLD 40(X12), F8
RET
`)
code := assembleRISCVHelper(t, fn)
want := []byte{
0x04, 0xec, // c.sd x9, 24(x8)
0x88, 0xc9, // c.sw x10, 16(x11)
0x00, 0xb6, // c.fsd f8, 40(x12)
0x04, 0x6c, // c.ld x9, 24(x8)
0x88, 0x49, // c.lw x10, 16(x11)
0x00, 0x36, // c.fld f8, 40(x12)
0x67, 0x80, 0x00, 0x00, // jalr x0, 0(x1)
}
if !bytes.Equal(code, want) {
t.Errorf("code = % x\nwant % x", code, want)
}
}
// TestRISCV_FENCE pins the FENCE encoding: the toolchain expands the bare
// mnemonic to fence iorw, iorw (0x0FF0000F), not fence 0,0.
func TestRISCV_FENCE(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·fence(SB), NOSPLIT, $0
FENCE
RET
`)
code := assembleRISCVHelper(t, fn)
want := []byte{
0x0f, 0x00, 0xf0, 0x0f, // fence iorw, iorw
0x67, 0x80, 0x00, 0x00, // jalr x0, 0(x1)
}
if !bytes.Equal(code, want) {
t.Errorf("code = % x\nwant % x", code, want)
}
}
// TestRISCV_RVC_WidthSpellings pins the compression of the GOROOT width
// spellings: MOVW and MOVD lower to their base load/store and compress
// exactly like LW/SW/FLD/FSD would (the toolchain compresses these shapes;
// before the normalisation they stayed 4 bytes).
func TestRISCV_RVC_WidthSpellings(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·widths(SB), NOSPLIT, $0-16
MOVW w+0(FP), X9
MOVW X9, v+4(FP)
MOVD d+0(FP), F8
MOVD F8, r+8(FP)
RET
`)
code := assembleRISCVHelper(t, fn)
want := []byte{
0xa2, 0x44, // c.lwsp x9, 8
0x26, 0xc6, // c.swsp x9, 12
0x22, 0x24, // c.fldsp f8, 8
0x22, 0xa8, // c.fsdsp f8, 16
0x67, 0x80, 0x00, 0x00, // jalr x0, 0(x1)
}
if !bytes.Equal(code, want) {
t.Errorf("code = % x\nwant % x", code, want)
}
}
func TestRISCV_system_instrs(t *testing.T) { func TestRISCV_system_instrs(t *testing.T) {
// Test FENCE, ECALL, EBREAK encoding. // Test FENCE, ECALL, EBREAK encoding.
fn := firstTextRISCV(t, `#include "textflag.h" fn := firstTextRISCV(t, `#include "textflag.h"
@@ -761,3 +837,137 @@ sub:
t.Error("expected error for CALL to local label, got nil") t.Error("expected error for CALL to local label, got nil")
} }
} }
// TestRISCVIndirectBranch pins the indirect branch encodings: JMP (X5) is the
// toolchain's JALR X0, 0(X5), and the trampoline form JALR rd, offset(rs1)
// takes its destination from the first operand (regression: the base
// register was once read as the destination, silently jumping to X0).
func TestRISCVIndirectBranch(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (X5)
JALR X0, 0(X6)
JALR X28, 0(X9)
RET
`)
code := assembleRISCVHelper(t, fn)
wantWords(t, code,
0x00028067, // jalr x0, 5(x0), 0
0x00030067, // jalr x0, 6(x0), 0
0x00048e67, // jalr x28, 9(x0), 0
0x00008067, // jalr x0, 1(x0), 0 (RET)
)
}
// encodeOneInstrRISCV encodes a single parsed instruction against a synthetic
// offsets map, the smallest honest harness for the branch-range diagnostics:
// the spans are far larger than any source a test would want to spell out.
func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]int) ([]byte, error) {
t.Helper()
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
instr := fn.Body[0].(*ast.Instr)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil)
}
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
// [-4096, 4094] and the J-type span [-1048576, 1048574] are diagnosed instead
// of wrapping silently to a wrong target.
func TestRISCVBranchJumpRange(t *testing.T) {
cases := []struct {
name string
src string
off int // the target's function-relative offset (pc 0)
ok bool
}{
{"branch max", "BEQ X10, X11, tgt\nRET\n", 4094, true},
{"branch past max", "BEQ X10, X11, tgt\nRET\n", 4096, false},
{"branch back max", "BEQ X10, X11, tgt\nRET\n", -4096, true},
{"branch back past max", "BEQ X10, X11, tgt\nRET\n", -4098, false},
{"branchz past max", "BEQZ X10, tgt\nRET\n", 4096, false},
{"jump max", "JMP tgt\nRET\n", 1048574, true},
{"jump past max", "JMP tgt\nRET\n", 1048576, false},
{"jump back max", "JMP tgt\nRET\n", -1048576, true},
{"jump back past max", "JMP tgt\nRET\n", -1048578, false},
{"jal past max", "JAL tgt\nRET\n", 1048576, false},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
_, err := encodeOneInstrRISCV(t, "TEXT ·f(SB), NOSPLIT, $0\n\t"+c.src, 0, map[string]int{"tgt": c.off})
if c.ok && err != nil {
t.Fatalf("unexpected error: %v", err)
}
if !c.ok && err == nil {
t.Fatal("expected an out-of-range diagnostic, got none")
}
})
}
}
// TestRISCVBranchFarBody drives the range check through the full two-pass
// assembler: a forward branch over a body larger than the B-type span must
// error rather than wrap.
func TestRISCVBranchFarBody(t *testing.T) {
var sb strings.Builder
sb.WriteString("#include \"textflag.h\"\nTEXT ·far(SB), NOSPLIT, $0\n\tBEQ X10, X11, done\n")
for range 1100 {
sb.WriteString("\tADD X10, X11, X12\n")
}
sb.WriteString("done:\n\tRET\n")
fn := firstTextRISCV(t, sb.String())
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected a branch-out-of-range error, got none")
}
}
// TestRISCV_CSRRange checks the CSR address range: the 12-bit field is
// diagnosed rather than masked, so CSRRW $4096 does not silently address
// CSR 0.
func TestRISCV_CSRRange(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·csrhi(SB), NOSPLIT, $0
CSRRW $4096, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected an out-of-range error for CSR $4096, got none")
}
fn = firstTextRISCV(t, `#include "textflag.h"
TEXT ·csrmax(SB), NOSPLIT, $0
CSRRW $4095, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("CSR $4095 must assemble: %v", err)
}
}
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
// span are diagnosed instead of silently truncated to their low 32 bits (the
// toolchain materialises such constants via SLLI expansion, which this
// assembler does not implement).
func TestRISCV_Imm64Rejected(t *testing.T) {
cases := []string{
"MOV $0x123456789, X10",
"ADDI $0x100000000, X10, X11",
"ANDI $-0x800000001, X10, X11",
"SUB $0x100000000, X10, X11",
}
for _, src := range cases {
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Errorf("%s: expected an out-of-range error, got none", src)
}
}
// The full signed 32-bit span still assembles, including the SUB form
// whose negated immediate only just fits.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·edge(SB), NOSPLIT, $0
MOV $2147483647, X10
MOV $-2147483648, X11
SUB $0x80000000, X12, X13
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("int32-span immediates must assemble: %v", err)
}
}
+211 -29
View File
@@ -4,6 +4,7 @@
package asm package asm
import ( import (
"fmt"
"strings" "strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast" "sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -32,6 +33,13 @@ import (
// riscvFrameInfo holds the frame layout derived from a TEXT directive. // riscvFrameInfo holds the frame layout derived from a TEXT directive.
type riscvFrameInfo struct { type riscvFrameInfo struct {
autosize int // the real SP adjustment (locals + saved LR) autosize int // the real SP adjustment (locals + saved LR)
// Stack-split guard state: the toolchain emits the check for every
// non-NOSPLIT function whose autosize is nonzero (a zero autosize is
// "effectively NOSPLIT"); unlike amd64 and arm64 there is no leaf
// auto-NOSPLIT.
needSplit bool
splitClass int // 0: <=StackSmall, 1: <=StackBig, 2: >StackBig
} }
// riscvComputeFrame derives the frame layout for a TEXT function. // riscvComputeFrame derives the frame layout for a TEXT function.
@@ -40,11 +48,34 @@ func riscvComputeFrame(t *ast.Text) riscvFrameInfo {
if frame != 0 || !riscvIsLeaf(t) { if frame != 0 || !riscvIsLeaf(t) {
// FixedFrameSize = 8: space for the saved link register. A // FixedFrameSize = 8: space for the saved link register. A
// zero-frame non-leaf function still opens an 8-byte frame for LR. // zero-frame non-leaf function still opens an 8-byte frame for LR.
return riscvFrameInfo{autosize: frame + 8} autosize := frame + 8
fi := riscvFrameInfo{autosize: autosize}
if !hasNoSplitFlag(t) {
fi.needSplit = true
switch {
case autosize <= stackSmall:
fi.splitClass = 0
case autosize <= stackBig:
fi.splitClass = 1
default:
fi.splitClass = 2
}
}
return fi
} }
return riscvFrameInfo{} return riscvFrameInfo{}
} }
// hasNoSplitFlag reports whether the TEXT directive carries NOSPLIT.
func hasNoSplitFlag(t *ast.Text) bool {
for _, f := range t.Flags {
if strings.EqualFold(f, "NOSPLIT") {
return true
}
}
return false
}
// riscvIsLeaf reports whether a function contains no call instructions. // riscvIsLeaf reports whether a function contains no call instructions.
// CALL always links; JAL/JALR link only when their destination register is // CALL always links; JAL/JALR link only when their destination register is
// the link register (X1), matching cmd/internal/obj/riscv's containsCall. // the link register (X1), matching cmd/internal/obj/riscv's containsCall.
@@ -58,17 +89,24 @@ func riscvIsLeaf(t *ast.Text) bool {
case "CALL": case "CALL":
return false return false
case "JAL": case "JAL":
// JAL rd, target — a call only when rd is the link register. // JAL rd, target, a call only when rd is the link register.
if len(in.Operands) >= 2 && regFromOperand(in.Operands[0]) == 1 { if len(in.Operands) >= 2 && regFromOperand(in.Operands[0]) == 1 {
return false return false
} }
case "JALR": case "JALR":
// JALR rs1, rd — a call when rd is X1; JALR offset(rs1) always // JALR rd, offset(rs1) links when the destination register (the
// links to X1. // first operand) is X1; JALR rs1, rd links when the second
// register is X1; JALR offset(rs1) always links to X1.
if len(in.Operands) == 1 { if len(in.Operands) == 1 {
return false return false
} }
if len(in.Operands) >= 2 && regFromOperand(in.Operands[1]) == 1 { if isMemOperand(in.Operands[1]) {
if regFromOperand(in.Operands[0]) == 1 {
return false
}
continue
}
if regFromOperand(in.Operands[1]) == 1 {
return false return false
} }
} }
@@ -84,28 +122,99 @@ func riscvPrologue(fi riscvFrameInfo) []byte {
return nil return nil
} }
var out []byte var out []byte
// MOV LR, -autosize(SP) — SD X1, -autosize(X2). The negative offset is // MOV LR, -autosize(SP), SD X1, -autosize(X2). The negative offset is
// not compressible to C.SDSP (unsigned), so it stays 4 bytes. // not compressible to C.SDSP (unsigned), so it stays 4 bytes. Beyond
out = append(out, wordLE(riscvSType(riscvEnc{0x23, 0x3, 0x00}, 2, 1, int32(-fi.autosize)))...) // the imm12 range the toolchain materialises the address in X31.
// ADDI $-autosize, SP, SP — open the frame (C.ADDI when it fits). if fits12(int32(-fi.autosize)) {
out = append(out, riscvSPAdjust(int32(-fi.autosize))...) out = append(out, wordLE(riscvSType(riscvEnc{0x23, 0x3, 0x00}, 2, 1, int32(-fi.autosize)))...)
// MOV LR, 0(SP) — SD X1, 0(X2) → C.SDSP X1, 0. } else {
out = append(out, riscvAddressInX31(int32(-fi.autosize))...)
lo := int32(-fi.autosize) - (splitHi(int32(-fi.autosize)) << 12)
out = append(out, wordLE(riscvSType(riscvEnc{0x23, 0x3, 0x00}, 31, 1, lo))...)
}
// ADDI $-autosize, SP, SP, open the frame (C.ADDI when it fits; X31
// materialisation beyond imm12).
if fits12(int32(-fi.autosize)) {
out = append(out, riscvSPAdjust(int32(-fi.autosize))...)
} else {
out = append(out, riscvAddToSP(int32(-fi.autosize))...)
}
// MOV LR, 0(SP), SD X1, 0(X2) → C.SDSP X1, 0.
c := rvcSSP(0x7, 1, 0) c := rvcSSP(0x7, 1, 0)
out = append(out, byte(c), byte(c>>8)) out = append(out, byte(c), byte(c>>8))
return out return out
} }
func fits12(v int32) bool { return v >= -2048 && v <= 2047 }
// splitHi returns the LUI half of the hi/lo split of v (what remains is the
// sign-extended 12-bit low part).
func splitHi(v int32) int32 {
_, high := splitRISCV32Imm(v)
return high
}
// riscvAddressInX31 materialises hi(v) into X31 against the stack pointer,
// matching the toolchain's large-frame addressing: C.LUI (or LUI) X31, hi;
// C.ADD (or ADD) X31, SP.
func riscvAddressInX31(v int32) []byte {
return riscvAddressInX31WithBase(v, 2)
}
// riscvAddressInX31WithBase materialises hi(v) into X31 against an arbitrary
// base register: LUI (or C.LUI) X31, hi; C.ADD X31, rs1. The CR rs2 field
// carries the full 5-bit register, so the compressed form is always
// available.
func riscvAddressInX31WithBase(v int32, rs1 int) []byte {
hi := splitHi(v)
var out []byte
if hi >= -32 && hi <= 31 {
c := rvcCI(0x3, 31, uint32(hi)&0x3F)
out = append(out, byte(c), byte(c>>8))
} else {
out = append(out, wordLE(riscvUType(riscvEnc{0x37, 0x0, 0x00}, 31, hi<<12))...)
}
c := rvcCR(0x9, 31, uint32(rs1))
return append(out, byte(c), byte(c>>8))
}
// riscvAddToSP adds v to SP through X31 for the values imm12 cannot carry:
// C.LUI X31, hi; C.ADDIW X31, lo; C.ADD SP, X31 (the toolchain's form).
func riscvAddToSP(v int32) []byte {
hi := splitHi(v)
lo := v - (hi << 12)
var out []byte
if hi >= -32 && hi <= 31 {
c := rvcCI(0x3, 31, uint32(hi)&0x3F)
out = append(out, byte(c), byte(c>>8))
} else {
out = append(out, wordLE(riscvUType(riscvEnc{0x37, 0x0, 0x00}, 31, hi<<12))...)
}
if lo >= -32 && lo <= 31 {
c := rvcCI(0x1, 31, uint32(lo)&0x3F)
out = append(out, byte(c), byte(c>>8))
} else {
out = append(out, wordLE(riscvIType(riscvEnc{0x1b, 0x0, 0x00}, 31, 31, lo))...)
}
c := rvcCR(0x9, 2, 31)
return append(out, byte(c), byte(c>>8))
}
// riscvReturn returns the bytes for a RET: the epilogue (restore LR and // riscvReturn returns the bytes for a RET: the epilogue (restore LR and
// deallocate the frame when present) followed by the uncompressed JALR X0, // deallocate the frame when present) followed by the uncompressed JALR X0,
// 0(X1) the toolchain emits for RET (it never compresses RET to C.JR). // 0(X1) the toolchain emits for RET (it never compresses RET to C.JR).
func riscvReturn(fi riscvFrameInfo) []byte { func riscvReturn(fi riscvFrameInfo) []byte {
var out []byte var out []byte
if fi.autosize != 0 { if fi.autosize != 0 {
// MOV 0(SP), LR — LD X1, 0(X2) → C.LDSP X1, 0. // MOV 0(SP), LR, LD X1, 0(X2) → C.LDSP X1, 0.
c := rvcLSP(0x3, 1, 0) c := rvcLSP(0x3, 1, 0)
out = append(out, byte(c), byte(c>>8)) out = append(out, byte(c), byte(c>>8))
// ADDI $autosize, SP, SP — close the frame (C.ADDI when it fits). // ADDI $autosize, SP, SP, close the frame (C.ADDI when it fits).
out = append(out, riscvSPAdjust(int32(fi.autosize))...) if fits12(int32(fi.autosize)) {
out = append(out, riscvSPAdjust(int32(fi.autosize))...)
} else {
out = append(out, riscvAddToSP(int32(fi.autosize))...)
}
} }
// JALR X0, 0(X1). // JALR X0, 0(X1).
return append(out, wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, 1, 0))...) return append(out, wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, 1, 0))...)
@@ -133,33 +242,36 @@ func riscvFitsCAddi(imm int32) bool {
} }
// riscvPrologueSpadjPC returns the function-relative byte offset where the // riscvPrologueSpadjPC returns the function-relative byte offset where the
// prologue has finished decrementing SP (the delta becomes autosize). // prologue has finished decrementing SP (the delta becomes autosize). It is
// computed from the same expansion functions the prologue emits, so the
// large-frame X31 materialisations are counted: C.LUI + C.ADD before the SD,
// C.LUI + ADDIW + C.ADD for the SP adjust.
func riscvPrologueSpadjPC(fi riscvFrameInfo) int { func riscvPrologueSpadjPC(fi riscvFrameInfo) int {
if fi.autosize == 0 { if fi.autosize == 0 {
return 0 return 0
} }
// SD (4 bytes) + ADDI/C.ADDI (2 or 4 bytes). adj := int32(-fi.autosize)
return 4 + riscvSPAdjustLen(int32(-fi.autosize)) if fits12(adj) {
// SD (4 bytes) + ADDI/C.ADDI (2 or 4 bytes).
return 4 + len(riscvSPAdjust(adj))
}
return len(riscvAddressInX31(adj)) + 4 + len(riscvAddToSP(adj))
} }
// riscvReturnEpilogueLen returns the byte length of the RET's epilogue up to // riscvReturnEpilogueLen returns the byte length of the RET's epilogue up to
// (but not including) the final JALR — the point where SP is restored. // (but not including) the final JALR, the point where SP is restored. The
// small frame closes with C.LDSP + ADDI/C.ADDI; the large frame materialises
// the adjustment through X31 (C.LUI + ADDIW + C.ADD).
func riscvReturnEpilogueLen(fi riscvFrameInfo) int { func riscvReturnEpilogueLen(fi riscvFrameInfo) int {
if fi.autosize == 0 { if fi.autosize == 0 {
return 0 return 0
} }
// C.LDSP (2 bytes) + ADDI/C.ADDI (2 or 4 bytes). adj := int32(fi.autosize)
return 2 + riscvSPAdjustLen(int32(fi.autosize)) if fits12(adj) {
} // C.LDSP (2 bytes) + ADDI/C.ADDI (2 or 4 bytes).
return 2 + len(riscvSPAdjust(adj))
func riscvSPAdjustLen(imm int32) int {
if imm != 0 && imm%16 == 0 && imm >= -512 && imm <= 511 {
return 2
} }
if riscvFitsCAddi(imm) { return 2 + len(riscvAddToSP(adj))
return 2
}
return 4
} }
// riscvResolvePseudo translates a pseudo-register memory reference into a // riscvResolvePseudo translates a pseudo-register memory reference into a
@@ -180,3 +292,73 @@ func riscvResolvePseudo(sym *ast.Symbol, fi riscvFrameInfo) (base int, off int32
} }
return -1, 0 return -1, 0
} }
// riscvGuardLen returns the byte length of the stack-split guard prefix
// including the inline morestack call (zero when the function needs no
// guard). Unlike amd64 and arm64, the toolchain places the morestack call
// between the guard and the body: the guard branches forward over it.
func riscvGuardLen(fi riscvFrameInfo) (int, error) {
g, _, err := riscvGuard(fi)
if err != nil {
return 0, err
}
return len(g), nil
}
// riscvGuard emits the stack-split guard prefix with the inline morestack
// call: the branch skips forward over JAL X5 and JAL X0 straight into the
// body; the JAL X5 carries the R_RISCV_JAL relocation. All offsets are
// relative to the guard itself, which sits at function offset 0.
func riscvGuard(fi riscvFrameInfo) ([]byte, Reloc, error) {
if !fi.needSplit {
return nil, Reloc{}, nil
}
// MOV 16(g), X6 (g.stackguard0), g = X27.
out := wordLE(riscvIType(riscvEnc{0x03, 0x3, 0x00}, 6, 27, 16))
jalBack := func() []byte {
// JAL X0 back to the function start: it sits right after the JAL X5,
// so its displacement is minus the current offset.
return wordLE(riscvJType(0, int32(-len(out))))
}
var reloc Reloc
switch fi.splitClass {
case 0:
// BLTU X6, SP, done (+12: over the CALL and the JMP back)
out = append(out, wordLE(riscvBType(riscvEnc{0x63, 0x06, 0x00}, 6, 2, 12))...)
call := len(out)
reloc = Reloc{Off: call, After: call + 4, Name: "runtime\u00b7morestack_noctxt", Kind: RelRISCVJal}
out = append(out, wordLE(riscvJType(5, 0))...)
out = append(out, jalBack()...)
case 1:
// ADDI $-(framesize-StackSmall), SP, X7; BLTU X6, X7, done (+12)
off := int32(fi.autosize - stackSmall)
out = append(out, wordLE(riscvIType(riscvEnc{0x13, 0x0, 0x00}, 7, 2, -off))...)
out = append(out, wordLE(riscvBType(riscvEnc{0x63, 0x06, 0x00}, 6, 7, 12))...)
call := len(out)
reloc = Reloc{Off: call, After: call + 4, Name: "runtime\u00b7morestack_noctxt", Kind: RelRISCVJal}
out = append(out, wordLE(riscvJType(5, 0))...)
out = append(out, jalBack()...)
default:
// MOV $(framesize-StackSmall), X7; BLTU SP, X7, call;
// ADD $-(framesize-StackSmall), SP, X7; BLTU X6, X7, call
off := int32(fi.autosize - stackSmall)
mov := encodeRISCVLoadImm(7, off)
out = append(out, mov...)
addiLen := riscvItypeImmediateSize("ADDI", -off)
out = append(out, wordLE(riscvBType(riscvEnc{0x63, 0x06, 0x00}, 2, 7, int32(addiLen+8)))...)
addi, err := encodeRISCVItypeImmediate("ADDI", riscvEnc{0x13, 0x0, 0x00}, 7, 2, -off)
if err != nil {
// The ADDI expansion failed: the SP adjustment this class
// depends on is not emittable, and silently dropping it would
// corrupt every stack reference in the body.
return nil, Reloc{}, fmt.Errorf("stack-split guard: %w", err)
}
out = append(out, addi...)
out = append(out, wordLE(riscvBType(riscvEnc{0x63, 0x06, 0x00}, 6, 7, 12))...)
call := len(out)
reloc = Reloc{Off: call, After: call + 4, Name: "runtime\u00b7morestack_noctxt", Kind: RelRISCVJal}
out = append(out, wordLE(riscvJType(5, 0))...)
out = append(out, jalBack()...)
}
return out, reloc, nil
}
+41
View File
@@ -65,6 +65,47 @@ TEXT ·framed(SB), NOSPLIT, $16-16
} }
} }
// TestRISCVFrameSpadjLargeFrame checks the stack-adjustment boundaries of a
// frame past the imm12 range: the prologue materialises the LR-store address
// and the SP adjustment through X31 (C.LUI + C.ADD + SD, then C.LUI + ADDIW +
// C.ADD), so the SP boundary lands at PC 16, and the RET closes with
// C.LDSP plus the same X31 adjustment, 10 bytes. Regression: both helpers
// assumed the small-frame prologue and reported 8 and 6.
func TestRISCVFrameSpadjLargeFrame(t *testing.T) {
f, errs := parser.Parse("bigframe_riscv64.s", `#include "textflag.h"
TEXT ·big(SB), NOSPLIT, $9000-8
MOV a+0(FP), X10
RET
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
fn := img.Funcs[0]
// autosize = 9008. Prologue: C.LUI X31 + C.ADD X31,SP (4) + SD (4) +
// C.LUI X31 + ADDIW X31 + C.ADD SP,X31 (8) = 16 bytes to the SP boundary;
// C.SDSP X1 (2) follows, so the body starts at 18.
wantSpadj := []SpadjStep{{PC: 16, Value: 9008}, {PC: 36, Value: 0}}
if len(fn.Spadj) != len(wantSpadj) {
t.Fatalf("spadj = %v, want %v", fn.Spadj, wantSpadj)
}
for i := range wantSpadj {
if fn.Spadj[i] != wantSpadj[i] {
t.Errorf("spadj[%d] = %v, want %v", i, fn.Spadj[i], wantSpadj[i])
}
}
// The FP load materialises its 9016-byte offset through X31 as well
// (8 bytes), then RET's epilogue (C.LDSP + X31 adjust = 10) plus JALR.
if fn.Size != 18+8+14 {
t.Errorf("size = %d, want %d", fn.Size, 18+8+14)
}
}
// TestRISCVRegAliases checks the Go ABI register aliases that the toolchain // TestRISCVRegAliases checks the Go ABI register aliases that the toolchain
// defines: LR is the link register (X1) and TMP is the assembler scratch // defines: LR is the link register (X1) and TMP is the assembler scratch
// register (X31/T6). // register (X31/T6).
+83
View File
@@ -74,6 +74,89 @@ DATA callee<>+0(SB)/8, $42
t.Error("ELF object missing R_RISCV_JAL relocation") t.Error("ELF object missing R_RISCV_JAL relocation")
} }
} }
// TestELFRISCVPCRELLO12Anchor checks the psABI's LO12 pairing rule: the
// R_RISCV_PCREL_LO12_I/S relocation must reference a symbol whose value is
// the AUIPC site of its HI20 partner (psABI §8.4.9; cmd/link generates one
// local text symbol per AUIPC for exactly this). The emitter pairs each
// HI20 (against the target symbol) with a LO12 against the .text section
// symbol whose addend is the AUIPC's section-relative offset, so S + A is
// the AUIPC address.
func TestELFRISCVPCRELLO12Anchor(t *testing.T) {
f, errs := parser.Parse("k_riscv64.s", `
#include "textflag.h"
TEXT ·sb(SB), NOSPLIT, $0-0
MOV $answer<>(SB), X10
MOV answer<>(SB), X11
MOV X12, answer<>(SB)
RET
GLOBL answer<>(SB), RODATA, $8
DATA answer<>+0(SB)/8, $42
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
obj, err := img.ELFRISCVObject()
if err != nil {
t.Fatalf("ELFRISCVObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse ELF: %v", err)
}
defer ef.Close()
if flags := binary.LittleEndian.Uint32(obj[48:]); flags != efRISCVFloatAbiDouble {
t.Errorf("e_flags = %#x, want %#x (EF_RISCV_FLOAT_ABI_DOUBLE)", flags, efRISCVFloatAbiDouble)
}
rela := ef.Section(".rela.text")
if rela == nil {
t.Fatal("missing .rela.text")
}
b, err := rela.Data()
if err != nil {
t.Fatal(err)
}
if len(b) != 6*24 {
t.Fatalf(".rela.text holds %d entries, want six (three HI20/LO12 pairs)", len(b)/24)
}
le := binary.LittleEndian
wantLo := []uint32{rRISCVPCRELLO12I, rRISCVPCRELLO12I, rRISCVPCRELLO12S}
for p := range 3 {
auipc := 8 * p
hi := b[p*2*24:]
lo := b[(p*2+1)*24:]
if off := le.Uint64(hi[0:]); off != uint64(auipc) {
t.Errorf("pair %d: HI20 r_offset = %d, want %d (the AUIPC)", p, off, auipc)
}
if typ := uint32(le.Uint64(hi[8:])); typ != rRISCVPCRELHI20 {
t.Errorf("pair %d: HI20 type = %d, want %d", p, typ, rRISCVPCRELHI20)
}
if sym := int(le.Uint64(hi[8:]) >> 32); sym == 0 || sym == 1 {
t.Errorf("pair %d: HI20 against symbol %d, want the target", p, sym)
}
if off := le.Uint64(lo[0:]); off != uint64(auipc+4) {
t.Errorf("pair %d: LO12 r_offset = %d, want %d", p, off, auipc+4)
}
if typ := uint32(le.Uint64(lo[8:])); typ != wantLo[p] {
t.Errorf("pair %d: LO12 type = %d, want %d", p, typ, wantLo[p])
}
// The LO12 must denote the AUIPC site: the .text section symbol
// (index 1) plus the AUIPC's section-relative offset as addend.
if sym := int(le.Uint64(lo[8:]) >> 32); sym != 1 {
t.Errorf("pair %d: LO12 against symbol %d, want 1 (the .text section symbol)", p, sym)
}
if add := int64(le.Uint64(lo[16:])); add != int64(auipc) {
t.Errorf("pair %d: LO12 addend = %d, want %d (S + A = the AUIPC address)", p, add, auipc)
}
}
}
func TestGOObjectRISCVStructure(t *testing.T) { func TestGOObjectRISCVStructure(t *testing.T) {
f, errs := parser.Parse("k_riscv64.s", ` f, errs := parser.Parse("k_riscv64.s", `
#include "textflag.h" #include "textflag.h"
+43 -43
View File
@@ -36,11 +36,11 @@ const (
vexNDS3Imm vexNDS3Imm
// vexExtract is the lane-extract form `OP $imm, ysrc, xdst`: ModRM.reg = // vexExtract is the lane-extract form `OP $imm, ysrc, xdst`: ModRM.reg =
// ysrc (op1), ModRM.rm = xdst or memory (op2), imm8 = op0. The YMM // ysrc (op1), ModRM.rm = xdst or memory (op2), imm8 = op0. The YMM
// source lives in the reg field, the destination in r/m — the PEXTR-style // source lives in the reg field, the destination in r/m, the PEXTR-style
// layout. VEXTRACTI128 and VEXTRACTF128 use this shape. // layout. VEXTRACTI128 and VEXTRACTF128 use this shape.
vexExtract vexExtract
// vexRMRev is the reversed two-operand form `OP src, dst` with the source // vexRMRev is the reversed two-operand form `OP src, dst` with the source
// in ModRM.reg and the destination in r/m — the layout of the EVEX // in ModRM.reg and the destination in r/m, the layout of the EVEX
// narrowing stores (VPMOVDW, VPMOVQD). // narrowing stores (VPMOVDW, VPMOVQD).
vexRMRev vexRMRev
// vexRMSrcLen is the two-operand conversion form `OP src, dst` whose // vexRMSrcLen is the two-operand conversion form `OP src, dst` whose
@@ -68,7 +68,7 @@ type vexSpec struct {
// incrementally; every entry is covered by a byte-for-byte ground-truth test // incrementally; every entry is covered by a byte-for-byte ground-truth test
// against the Go assembler. // against the Go assembler.
var vexTable = map[string]vexSpec{ var vexTable = map[string]vexSpec{
// VEX.128/256.66.0F.WIG — integer arithmetic / logic / compare. // VEX.128/256.66.0F.WIG, integer arithmetic / logic / compare.
"VPADDD": {1, 0xFE, 0, 1, -1, vexNDS3}, "VPADDD": {1, 0xFE, 0, 1, -1, vexNDS3},
"VPADDQ": {1, 0xD4, 0, 1, -1, vexNDS3}, "VPADDQ": {1, 0xD4, 0, 1, -1, vexNDS3},
"VPSUBD": {1, 0xFA, 0, 1, -1, vexNDS3}, "VPSUBD": {1, 0xFA, 0, 1, -1, vexNDS3},
@@ -82,7 +82,7 @@ var vexTable = map[string]vexSpec{
"VPUNPCKHDQ": {1, 0x6A, 0, 1, -1, vexNDS3}, "VPUNPCKHDQ": {1, 0x6A, 0, 1, -1, vexNDS3},
"VPUNPCKLQDQ": {1, 0x6C, 0, 1, -1, vexNDS3}, "VPUNPCKLQDQ": {1, 0x6C, 0, 1, -1, vexNDS3},
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3}, "VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3},
// VEX.256.66.0F38.W0 — dword permute (three-operand NDS form). // VEX.256.66.0F38.W0, dword permute (three-operand NDS form).
"VPERMD": {2, 0x36, 0, 1, -1, vexNDS3}, "VPERMD": {2, 0x36, 0, 1, -1, vexNDS3},
// VEX.128/256.66.0F38.WIG. // VEX.128/256.66.0F38.WIG.
"VPMULLD": {2, 0x40, 0, 1, -1, vexNDS3}, "VPMULLD": {2, 0x40, 0, 1, -1, vexNDS3},
@@ -90,14 +90,14 @@ var vexTable = map[string]vexSpec{
"VPSHUFB": {2, 0x00, 0, 1, -1, vexNDS3}, "VPSHUFB": {2, 0x00, 0, 1, -1, vexNDS3},
"VPCMPGTQ": {2, 0x37, 0, 1, -1, vexNDS3}, "VPCMPGTQ": {2, 0x37, 0, 1, -1, vexNDS3},
// VEX.128/256.66.0F.WIG — packed double-precision arithmetic / logic. // VEX.128/256.66.0F.WIG, packed double-precision arithmetic / logic.
"VADDPD": {1, 0x58, 0, 1, -1, vexNDS3}, "VADDPD": {1, 0x58, 0, 1, -1, vexNDS3},
"VMULPD": {1, 0x59, 0, 1, -1, vexNDS3}, "VMULPD": {1, 0x59, 0, 1, -1, vexNDS3},
"VSUBPD": {1, 0x5C, 0, 1, -1, vexNDS3}, "VSUBPD": {1, 0x5C, 0, 1, -1, vexNDS3},
"VDIVPD": {1, 0x5E, 0, 1, -1, vexNDS3}, "VDIVPD": {1, 0x5E, 0, 1, -1, vexNDS3},
"VMINPD": {1, 0x5D, 0, 1, -1, vexNDS3}, "VMINPD": {1, 0x5D, 0, 1, -1, vexNDS3},
"VMAXPD": {1, 0x5F, 0, 1, -1, vexNDS3}, "VMAXPD": {1, 0x5F, 0, 1, -1, vexNDS3},
// VEX.128/256.0F.WIG — packed single-precision arithmetic. // VEX.128/256.0F.WIG, packed single-precision arithmetic.
"VADDPS": {1, 0x58, 0, 0, -1, vexNDS3}, "VADDPS": {1, 0x58, 0, 0, -1, vexNDS3},
"VMULPS": {1, 0x59, 0, 0, -1, vexNDS3}, "VMULPS": {1, 0x59, 0, 0, -1, vexNDS3},
"VSUBPS": {1, 0x5C, 0, 0, -1, vexNDS3}, "VSUBPS": {1, 0x5C, 0, 0, -1, vexNDS3},
@@ -107,7 +107,7 @@ var vexTable = map[string]vexSpec{
"VXORPD": {1, 0x57, 0, 1, -1, vexNDS3}, "VXORPD": {1, 0x57, 0, 1, -1, vexNDS3},
"VUNPCKHPD": {1, 0x15, 0, 1, -1, vexNDS3}, "VUNPCKHPD": {1, 0x15, 0, 1, -1, vexNDS3},
"VUNPCKLPD": {1, 0x14, 0, 1, -1, vexNDS3}, "VUNPCKLPD": {1, 0x14, 0, 1, -1, vexNDS3},
// VEX.128.F2.0F.WIG — scalar double-precision arithmetic (the packed // VEX.128.F2.0F.WIG, scalar double-precision arithmetic (the packed
// opcodes with an F2 pp). // opcodes with an F2 pp).
"VADDSD": {1, 0x58, 0, 3, -1, vexNDS3}, "VADDSD": {1, 0x58, 0, 3, -1, vexNDS3},
"VSUBSD": {1, 0x5C, 0, 3, -1, vexNDS3}, "VSUBSD": {1, 0x5C, 0, 3, -1, vexNDS3},
@@ -115,7 +115,7 @@ var vexTable = map[string]vexSpec{
"VDIVSD": {1, 0x5E, 0, 3, -1, vexNDS3}, "VDIVSD": {1, 0x5E, 0, 3, -1, vexNDS3},
"VMINSD": {1, 0x5D, 0, 3, -1, vexNDS3}, "VMINSD": {1, 0x5D, 0, 3, -1, vexNDS3},
"VMAXSD": {1, 0x5F, 0, 3, -1, vexNDS3}, "VMAXSD": {1, 0x5F, 0, 3, -1, vexNDS3},
// VEX.128.F3.0F.WIG — scalar single-precision arithmetic (the packed // VEX.128.F3.0F.WIG, scalar single-precision arithmetic (the packed
// opcodes with an F3 pp). // opcodes with an F3 pp).
"VADDSS": {1, 0x58, 0, 2, -1, vexNDS3}, "VADDSS": {1, 0x58, 0, 2, -1, vexNDS3},
"VSUBSS": {1, 0x5C, 0, 2, -1, vexNDS3}, "VSUBSS": {1, 0x5C, 0, 2, -1, vexNDS3},
@@ -123,10 +123,10 @@ var vexTable = map[string]vexSpec{
"VDIVSS": {1, 0x5E, 0, 2, -1, vexNDS3}, "VDIVSS": {1, 0x5E, 0, 2, -1, vexNDS3},
"VMINSS": {1, 0x5D, 0, 2, -1, vexNDS3}, "VMINSS": {1, 0x5D, 0, 2, -1, vexNDS3},
"VMAXSS": {1, 0x5F, 0, 2, -1, vexNDS3}, "VMAXSS": {1, 0x5F, 0, 2, -1, vexNDS3},
// VEX.128/256.66.0F38.W1 — fused multiply-add (NDS form). // VEX.128/256.66.0F38.W1, fused multiply-add (NDS form).
"VFMADD231PD": {2, 0xB8, 1, 1, -1, vexNDS3}, "VFMADD231PD": {2, 0xB8, 1, 1, -1, vexNDS3},
// VEX.128/256.66.0F38.WIG — sign/zero extend and broadcast (reg=dst, rm=src, // VEX.128/256.66.0F38.WIG, sign/zero extend and broadcast (reg=dst, rm=src,
// no vvvv). // no vvvv).
"VPMOVSXWD": {2, 0x23, 0, 1, -1, vexRM}, "VPMOVSXWD": {2, 0x23, 0, 1, -1, vexRM},
"VPMOVSXDQ": {2, 0x25, 0, 1, -1, vexRM}, "VPMOVSXDQ": {2, 0x25, 0, 1, -1, vexRM},
@@ -143,70 +143,70 @@ var vexTable = map[string]vexSpec{
"VPBROADCASTQ": {2, 0x59, 0, 1, -1, vexRM}, "VPBROADCASTQ": {2, 0x59, 0, 1, -1, vexRM},
"VPBROADCASTB": {2, 0x78, 0, 1, -1, vexRM}, "VPBROADCASTB": {2, 0x78, 0, 1, -1, vexRM},
"VPBROADCASTW": {2, 0x79, 0, 1, -1, vexRM}, "VPBROADCASTW": {2, 0x79, 0, 1, -1, vexRM},
// VEX.128/256.F3.0F.WIG — signed dword to packed double conversion // VEX.128/256.F3.0F.WIG, signed dword to packed double conversion
// (reg=dst, rm=src, no vvvv; the length follows the destination). // (reg=dst, rm=src, no vvvv; the length follows the destination).
"VCVTDQ2PD": {1, 0xE6, 0, 2, -1, vexRM}, "VCVTDQ2PD": {1, 0xE6, 0, 2, -1, vexRM},
// VEX.128/256.0F.WIG — signed dword to packed single conversion // VEX.128/256.0F.WIG, signed dword to packed single conversion
// (reg=dst, rm=src, no vvvv, no mandatory prefix). // (reg=dst, rm=src, no vvvv, no mandatory prefix).
"VCVTDQ2PS": {1, 0x5B, 0, 0, -1, vexRM}, "VCVTDQ2PS": {1, 0x5B, 0, 0, -1, vexRM},
// VEX.128/256.0F.WIG — packed single to packed double conversion // VEX.128/256.0F.WIG, packed single to packed double conversion
// (reg=dst, rm=src; the destination is the wide operand and sets the // (reg=dst, rm=src; the destination is the wide operand and sets the
// length). Intel's maps prescribe the F3 prefix here (VEX.pp = 10), but // length). Intel's maps prescribe the F3 prefix here (VEX.pp = 10), but
// the Go assembler emits the instruction with pp = 00, and gasm follows // the Go assembler emits the instruction with pp = 00, and gasm follows
// the Go assembler's bytes — its machine code is the oracle, not the // the Go assembler's bytes, its machine code is the oracle, not the
// manual. // manual.
"VCVTPS2PD": {1, 0x5A, 0, 0, -1, vexRM}, "VCVTPS2PD": {1, 0x5A, 0, 0, -1, vexRM},
// VEX.128.F2.0F.WIG — duplicate the low double of each 128-bit lane // VEX.128.F2.0F.WIG, duplicate the low double of each 128-bit lane
// (reg=dst, rm=src, no vvvv; the length follows the destination). // (reg=dst, rm=src, no vvvv; the length follows the destination).
"VMOVDDUP": {1, 0x12, 0, 3, -1, vexRM}, "VMOVDDUP": {1, 0x12, 0, 3, -1, vexRM},
// VEX.128/256.66.0F.WIG — move mask to a GPR (reg=gpr dst, rm=vec src). // VEX.128/256.66.0F.WIG, move mask to a GPR (reg=gpr dst, rm=vec src).
"VPMOVMSKB": {1, 0xD7, 0, 1, -1, vexRM}, "VPMOVMSKB": {1, 0xD7, 0, 1, -1, vexRM},
"VMOVMSKPS": {1, 0x50, 0, 0, -1, vexRM}, // no 66 prefix (that would be VMOVMSKPD) "VMOVMSKPS": {1, 0x50, 0, 0, -1, vexRM}, // no 66 prefix (that would be VMOVMSKPD)
// VEX.128/256.66.0F.WIG — immediate shifts (opdigit selects the shift). // VEX.128/256.66.0F.WIG, immediate shifts (opdigit selects the shift).
"VPSLLD": {1, 0x72, 0, 1, 6, vexShiftImm}, "VPSLLD": {1, 0x72, 0, 1, 6, vexShiftImm},
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm}, "VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm},
"VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm}, "VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm},
"VPSRLQ": {1, 0x73, 0, 1, 2, vexShiftImm}, "VPSRLQ": {1, 0x73, 0, 1, 2, vexShiftImm},
"VPSLLQ": {1, 0x73, 0, 1, 6, vexShiftImm}, "VPSLLQ": {1, 0x73, 0, 1, 6, vexShiftImm},
// VEX.128/256.66.0F.WIG — immediate shuffle (reg=dst, rm=src, imm8). // VEX.128/256.66.0F.WIG, immediate shuffle (reg=dst, rm=src, imm8).
"VPSHUFD": {1, 0x70, 0, 1, -1, vexImmRM}, "VPSHUFD": {1, 0x70, 0, 1, -1, vexImmRM},
// VEX.256.66.0F3A.W1 — qword permute (reg=dst, rm=src, imm8). // VEX.256.66.0F3A.W1, qword permute (reg=dst, rm=src, imm8).
"VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM}, "VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM},
// VEX.128/256.66.0F.WIG — two-source shuffle (reg=dst, vvvv=src1, rm=src2, // VEX.128/256.66.0F.WIG, two-source shuffle (reg=dst, vvvv=src1, rm=src2,
// imm8). // imm8).
"VSHUFPD": {1, 0xC6, 0, 1, -1, vexNDS3Imm}, "VSHUFPD": {1, 0xC6, 0, 1, -1, vexNDS3Imm},
// VEX.256.66.0F3A.W0 — permute / insert (same shape; VINSERTI128's rm is // VEX.256.66.0F3A.W0, permute / insert (same shape; VINSERTI128's rm is
// the XMM or memory source). // the XMM or memory source).
"VPERM2I128": {3, 0x46, 0, 1, -1, vexNDS3Imm}, "VPERM2I128": {3, 0x46, 0, 1, -1, vexNDS3Imm},
"VINSERTI128": {3, 0x38, 0, 1, -1, vexNDS3Imm}, "VINSERTI128": {3, 0x38, 0, 1, -1, vexNDS3Imm},
// VEX.256.66.0F3A.W0 — lane extract (reg=YMM src, rm=XMM/memory dst, imm8). // VEX.256.66.0F3A.W0, lane extract (reg=YMM src, rm=XMM/memory dst, imm8).
"VEXTRACTI128": {3, 0x39, 0, 1, -1, vexExtract}, "VEXTRACTI128": {3, 0x39, 0, 1, -1, vexExtract},
"VEXTRACTF128": {3, 0x19, 0, 1, -1, vexExtract}, "VEXTRACTF128": {3, 0x19, 0, 1, -1, vexExtract},
// VEX.128/256.66.0F3A.W0 — half-precision convert back ($imm, src, dst: // VEX.128/256.66.0F3A.W0, half-precision convert back ($imm, src, dst:
// reg=src, rm=XMM/memory dst, imm8 — the extract layout). // reg=src, rm=XMM/memory dst, imm8, the extract layout).
"VCVTPS2PH": {3, 0x1D, 0, 1, -1, vexExtract}, "VCVTPS2PH": {3, 0x1D, 0, 1, -1, vexExtract},
// VEX.128.0F.W0 — no operands. // VEX.128.0F.W0, no operands.
"VZEROUPPER": {1, 0x77, 0, 0, -1, vexZero}, "VZEROUPPER": {1, 0x77, 0, 0, -1, vexZero},
// VEX.128.0F.W0 — mask-register test (KTESTW k1, k2: reg = dst, rm = src). // VEX.128.0F.W0, mask-register test (KTESTW k1, k2: reg = dst, rm = src).
"KTESTW": {1, 0x99, 0, 0, -1, vexRM}, "KTESTW": {1, 0x99, 0, 0, -1, vexRM},
// VEX.66.0F38.W0 — broadcast a single/double to all lanes (reg=dst, // VEX.66.0F38.W0, broadcast a single/double to all lanes (reg=dst,
// rm=scalar memory; SD is 256-bit only). // rm=scalar memory; SD is 256-bit only).
"VBROADCASTSS": {2, 0x18, 0, 1, -1, vexRM}, "VBROADCASTSS": {2, 0x18, 0, 1, -1, vexRM},
"VBROADCASTSD": {2, 0x19, 0, 1, -1, vexRM}, "VBROADCASTSD": {2, 0x19, 0, 1, -1, vexRM},
// VEX.66.0F38.W0 — half-precision convert (reg=dst, rm=half-width // VEX.66.0F38.W0, half-precision convert (reg=dst, rm=half-width
// source). // source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM}, "VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM},
// VEX.F3.0F.WIG — replicate even/odd singles (reg=dst, rm=src). // VEX.F3.0F.WIG, replicate even/odd singles (reg=dst, rm=src).
"VMOVSLDUP": {1, 0x12, 0, 2, -1, vexRM}, "VMOVSLDUP": {1, 0x12, 0, 2, -1, vexRM},
"VMOVSHDUP": {1, 0x16, 0, 2, -1, vexRM}, "VMOVSHDUP": {1, 0x16, 0, 2, -1, vexRM},
// VEX.66.0F.WIG — packed double to packed single conversion, the X/Y // VEX.66.0F.WIG, packed double to packed single conversion, the X/Y
// spellings: the destination is always XMM and the spelling fixes the // spellings: the destination is always XMM and the spelling fixes the
// source length (X = 128, Y = 256). // source length (X = 128, Y = 256).
"VCVTPD2PSX": {1, 0x5A, 0, 1, -1, vexRMSrcLen}, "VCVTPD2PSX": {1, 0x5A, 0, 1, -1, vexRMSrcLen},
@@ -230,14 +230,14 @@ var vexTable = map[string]vexSpec{
"VCVTSI2SSL": {1, 0x2A, 0, 2, -1, vexNDS3}, "VCVTSI2SSL": {1, 0x2A, 0, 2, -1, vexNDS3},
"VCVTSI2SSQ": {1, 0x2A, 1, 2, -1, vexNDS3}, "VCVTSI2SSQ": {1, 0x2A, 1, 2, -1, vexNDS3},
// VEX.128/256.66.0F.WIG — word shifts (opdigit selects the shift). // VEX.128/256.66.0F.WIG, word shifts (opdigit selects the shift).
"VPSRLW": {1, 0x71, 0, 1, 2, vexShiftImm}, "VPSRLW": {1, 0x71, 0, 1, 2, vexShiftImm},
"VPSRAW": {1, 0x71, 0, 1, 4, vexShiftImm}, "VPSRAW": {1, 0x71, 0, 1, 4, vexShiftImm},
"VPSLLW": {1, 0x71, 0, 1, 6, vexShiftImm}, "VPSLLW": {1, 0x71, 0, 1, 6, vexShiftImm},
// VEX.F2.0F — packed double to packed dword conversions, truncating and // VEX.F2.0F, packed double to packed dword conversions, truncating and
// non-truncating. The destination is always XMM; the X/Y spellings fix // non-truncating. The destination is always XMM; the X/Y spellings fix
// the source length (XMM/YMM), and VEX.L follows it — see vexSrcLen. // the source length (XMM/YMM), and VEX.L follows it, see vexSrcLen.
"VCVTPD2DQX": {1, 0xE6, 0, 3, -1, vexRMSrcLen}, "VCVTPD2DQX": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
"VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen}, "VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
"VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen}, "VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
@@ -257,7 +257,7 @@ var vexSrcLen = map[string]int{
"VCVTPD2PSY": 1, "VCVTPD2PSY": 1,
} }
// vexVarShift maps the shift mnemonics to their variable-count opcode — the // vexVarShift maps the shift mnemonics to their variable-count opcode, the
// form whose count comes from an XMM register or memory (VPSRLQ X0, Y8, Y8), // form whose count comes from an XMM register or memory (VPSRLQ X0, Y8, Y8),
// an ordinary NDS encoding rather than the /digit immediate form above. // an ordinary NDS encoding rather than the /digit immediate form above.
var vexVarShift = map[string]byte{ var vexVarShift = map[string]byte{
@@ -288,20 +288,20 @@ type vexMoveSpec struct {
// vexMoveTable maps an upper-case move mnemonic to its encoding. // vexMoveTable maps an upper-case move mnemonic to its encoding.
var vexMoveTable = map[string]vexMoveSpec{ var vexMoveTable = map[string]vexMoveSpec{
// VEX.128/256.F3.0F.WIG — unaligned integer move. // VEX.128/256.F3.0F.WIG, unaligned integer move.
"VMOVDQU": {1, 2, 0x6F, 0x7F, 0, 0, 0, 0, true, false, false}, "VMOVDQU": {1, 2, 0x6F, 0x7F, 0, 0, 0, 0, true, false, false},
// VEX.128/256.66.0F.WIG — unaligned packed double move. // VEX.128/256.66.0F.WIG, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 0, 0, 0, 0, true, false, false}, "VMOVUPD": {1, 1, 0x10, 0x11, 0, 0, 0, 0, true, false, false},
// VEX.128.66.0F.W0 — 32-bit GPR/memory ↔ XMM. // VEX.128.66.0F.W0, 32-bit GPR/memory ↔ XMM.
"VMOVD": {1, 1, 0x6E, 0x7E, 0, 0, 0, 0, false, true, true}, "VMOVD": {1, 1, 0x6E, 0x7E, 0, 0, 0, 0, false, true, true},
// VMOVQ — 66 6E W1 (r/m→xmm), 66 7E W1 (xmm→r/m), 66 D6 W0 (xmm→xmm). // VMOVQ, 66 6E W1 (r/m→xmm), 66 7E W1 (xmm→r/m), 66 D6 W0 (xmm→xmm).
"VMOVQ": {1, 1, 0x6E, 0x7E, 1, 1, 0xD6, 0, true, true, true}, "VMOVQ": {1, 1, 0x6E, 0x7E, 1, 1, 0xD6, 0, true, true, true},
// VEX.128.F2.0F.WIG — scalar double move, memory operands only (the // VEX.128.F2.0F.WIG, scalar double move, memory operands only (the
// register form takes three operands and is not supported yet). // register form takes three operands and is not supported yet).
"VMOVSD": {1, 3, 0x10, 0x11, 0, 0, 0, 0, false, false, true}, "VMOVSD": {1, 3, 0x10, 0x11, 0, 0, 0, 0, false, false, true},
// VEX.128.F3.0F.WIG — scalar single move, memory operands only. // VEX.128.F3.0F.WIG, scalar single move, memory operands only.
"VMOVSS": {1, 2, 0x10, 0x11, 0, 0, 0, 0, false, false, true}, "VMOVSS": {1, 2, 0x10, 0x11, 0, 0, 0, 0, false, false, true},
// VEX.128/256 — aligned packed moves. // VEX.128/256, aligned packed moves.
"VMOVAPS": {1, 0, 0x28, 0x29, 0, 0, 0, 0, true, false, false}, "VMOVAPS": {1, 0, 0x28, 0x29, 0, 0, 0, 0, true, false, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 0, 0, 0, 0, true, false, false}, "VMOVAPD": {1, 1, 0x28, 0x29, 0, 0, 0, 0, true, false, false},
} }
@@ -317,7 +317,7 @@ func isVex(mnemUpper string) bool {
// encodeVex encodes a VEX instruction with operands in Plan 9 order. // encodeVex encodes a VEX instruction with operands in Plan 9 order.
func (e *enc) encodeVex(mnemUpper string, ops []Operand) error { func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
// Vector register indices 16–31 exist only in EVEX encodings; fail // Vector register indices 16-31 exist only in EVEX encodings; fail
// loudly rather than silently truncating the index. // loudly rather than silently truncating the index.
for _, op := range ops { for _, op := range ops {
if r, ok := op.(Reg); ok && r.isVec() && r.idx >= 16 { if r, ok := op.(Reg); ok && r.isVec() && r.idx >= 16 {
@@ -420,7 +420,7 @@ func (e *enc) encodeVexRM(spec vexSpec, ops []Operand) error {
} }
// encodeVexRMSrcLen encodes a length-narrowing conversion: OP src, dst with // encodeVexRMSrcLen encodes a length-narrowing conversion: OP src, dst with
// the destination always XMM and the VEX.L bit following the source — fixed // the destination always XMM and the VEX.L bit following the source, fixed
// by the mnemonic's spelling (VCVTPD2DQX = 128, VCVTPD2DQY = 256) even when // by the mnemonic's spelling (VCVTPD2DQX = 128, VCVTPD2DQY = 256) even when
// the source is memory. // the source is memory.
func (e *enc) encodeVexRMSrcLen(mnem string, spec vexSpec, ops []Operand) error { func (e *enc) encodeVexRMSrcLen(mnem string, spec vexSpec, ops []Operand) error {
+5 -5
View File
@@ -173,7 +173,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VPMULLD Y1,Y2,Y3", "VPMULLD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d40d9", ""}, {"VPMULLD Y1,Y2,Y3", "VPMULLD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d40d9", ""},
{"VPUNPCKLDQ Y4,Y3,Y5", "VPUNPCKLDQ", []Operand{vreg(t, "Y4"), vreg(t, "Y3"), vreg(t, "Y5")}, "c5e562ec", ""}, {"VPUNPCKLDQ Y4,Y3,Y5", "VPUNPCKLDQ", []Operand{vreg(t, "Y4"), vreg(t, "Y3"), vreg(t, "Y5")}, "c5e562ec", ""},
{"VPERMD Y1,Y2,Y3", "VPERMD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d36d9", ""}, {"VPERMD Y1,Y2,Y3", "VPERMD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d36d9", ""},
// Floating point (packed and scalar) and FMA — same NDS form, the pp // Floating point (packed and scalar) and FMA; same NDS form, the pp
// bits and map select the operation. // bits and map select the operation.
{"VADDPD Y9,Y8,Y8", "VADDPD", []Operand{vreg(t, "Y9"), vreg(t, "Y8"), vreg(t, "Y8")}, "c4413d58c1", ""}, {"VADDPD Y9,Y8,Y8", "VADDPD", []Operand{vreg(t, "Y9"), vreg(t, "Y8"), vreg(t, "Y8")}, "c4413d58c1", ""},
{"VADDPD X1,X2,X3", "VADDPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e958d9", ""}, {"VADDPD X1,X2,X3", "VADDPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e958d9", ""},
@@ -217,7 +217,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VEXTRACTI128 $1,Y8,X9", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d39c101", ""}, {"VEXTRACTI128 $1,Y8,X9", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d39c101", ""},
{"VEXTRACTI128 $1,Y8,(DI)", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), Ptr(DI, 0, 16)}, "c4637d390701", ""}, {"VEXTRACTI128 $1,Y8,(DI)", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), Ptr(DI, 0, 16)}, "c4637d390701", ""},
{"VEXTRACTF128 $1,Y8,X9", "VEXTRACTF128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d19c101", ""}, {"VEXTRACTF128 $1,Y8,X9", "VEXTRACTF128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d19c101", ""},
// Moves — each direction picks its own opcode and VEX.W. // Moves; each direction picks its own opcode and VEX.W.
{"VMOVDQU (SI),Y1", "VMOVDQU", []Operand{Ptr(SI, 0, 32), vreg(t, "Y1")}, "c5fe6f0e", ""}, {"VMOVDQU (SI),Y1", "VMOVDQU", []Operand{Ptr(SI, 0, 32), vreg(t, "Y1")}, "c5fe6f0e", ""},
{"VMOVDQU Y3,(DI)", "VMOVDQU", []Operand{vreg(t, "Y3"), Ptr(DI, 0, 32)}, "c5fe7f1f", ""}, {"VMOVDQU Y3,(DI)", "VMOVDQU", []Operand{vreg(t, "Y3"), Ptr(DI, 0, 32)}, "c5fe7f1f", ""},
{"VMOVDQU X1,X2", "VMOVDQU", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fa7fca", ""}, {"VMOVDQU X1,X2", "VMOVDQU", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fa7fca", ""},
@@ -234,7 +234,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VMOVD AX,X0", "VMOVD", []Operand{AX, vreg(t, "X0")}, "c5f96ec0", ""}, {"VMOVD AX,X0", "VMOVD", []Operand{AX, vreg(t, "X0")}, "c5f96ec0", ""},
{"VMOVSD (SI),X8", "VMOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X8")}, "c57b1006", ""}, {"VMOVSD (SI),X8", "VMOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X8")}, "c57b1006", ""},
{"VMOVSD X8,(SI)", "VMOVSD", []Operand{vreg(t, "X8"), Ptr(SI, 0, 8)}, "c57b1106", ""}, {"VMOVSD X8,(SI)", "VMOVSD", []Operand{vreg(t, "X8"), Ptr(SI, 0, 8)}, "c57b1106", ""},
// Packed double arithmetic and unpack — the NDS form, the opcode // Packed double arithmetic and unpack; the NDS form, the opcode
// selects the operation. // selects the operation.
{"VSUBPD Y1,Y2,Y3", "VSUBPD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ed5cd9", ""}, {"VSUBPD Y1,Y2,Y3", "VSUBPD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ed5cd9", ""},
{"VDIVPD X1,X2,X3", "VDIVPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e95ed9", ""}, {"VDIVPD X1,X2,X3", "VDIVPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e95ed9", ""},
@@ -255,12 +255,12 @@ func TestVexGroundTruth(t *testing.T) {
{"VMINSS X6,X7,X8", "VMINSS", []Operand{vreg(t, "X6"), vreg(t, "X7"), vreg(t, "X8")}, "c5425dc6", ""}, {"VMINSS X6,X7,X8", "VMINSS", []Operand{vreg(t, "X6"), vreg(t, "X7"), vreg(t, "X8")}, "c5425dc6", ""},
{"VMAXSS X1,X2,X3", "VMAXSS", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5ea5fd9", ""}, {"VMAXSS X1,X2,X3", "VMAXSS", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5ea5fd9", ""},
{"VADDSD 8(AX),X1,X2", "VADDSD", []Operand{Ptr(AX, 8, 8), vreg(t, "X1"), vreg(t, "X2")}, "c5f3585008", ""}, {"VADDSD 8(AX),X1,X2", "VADDSD", []Operand{Ptr(AX, 8, 8), vreg(t, "X1"), vreg(t, "X2")}, "c5f3585008", ""},
// VMOVDDUP — duplicate the low double (reg=dst, rm=src, F2 pp). // VMOVDDUP; duplicate the low double (reg=dst, rm=src, F2 pp).
{"VMOVDDUP X1,X2", "VMOVDDUP", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fb12d1", ""}, {"VMOVDDUP X1,X2", "VMOVDDUP", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fb12d1", ""},
{"VMOVDDUP Y1,Y2", "VMOVDDUP", []Operand{vreg(t, "Y1"), vreg(t, "Y2")}, "c5ff12d1", ""}, {"VMOVDDUP Y1,Y2", "VMOVDDUP", []Operand{vreg(t, "Y1"), vreg(t, "Y2")}, "c5ff12d1", ""},
{"VMOVDDUP 8(AX),X1", "VMOVDDUP", []Operand{Ptr(AX, 8, 8), vreg(t, "X1")}, "c5fb124808", ""}, {"VMOVDDUP 8(AX),X1", "VMOVDDUP", []Operand{Ptr(AX, 8, 8), vreg(t, "X1")}, "c5fb124808", ""},
// Conversions: DQ→PS (no prefix), PS→PD (Go emits it without the F3 // Conversions: DQ→PS (no prefix), PS→PD (Go emits it without the F3
// prefix — see the table comment), DQ→PD. // prefix; see the table comment), DQ→PD.
{"VCVTDQ2PS X1,X2", "VCVTDQ2PS", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85bd1", ""}, {"VCVTDQ2PS X1,X2", "VCVTDQ2PS", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85bd1", ""},
{"VCVTDQ2PS Y3,Y4", "VCVTDQ2PS", []Operand{vreg(t, "Y3"), vreg(t, "Y4")}, "c5fc5be3", ""}, {"VCVTDQ2PS Y3,Y4", "VCVTDQ2PS", []Operand{vreg(t, "Y3"), vreg(t, "Y4")}, "c5fc5be3", ""},
{"VCVTPS2PD X1,X2", "VCVTPS2PD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85ad1", ""}, {"VCVTPS2PD X1,X2", "VCVTPS2PD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85ad1", ""},
+3 -2
View File
@@ -17,7 +17,7 @@ type File struct {
Orphans []Stmt // labels/instructions seen before any TEXT directive Orphans []Stmt // labels/instructions seen before any TEXT directive
// Macros holds the names introduced by #define directives in this file. // Macros holds the names introduced by #define directives in this file.
// The linter uses it to avoid flagging macro invocations as unknown // The linter uses it to avoid flagging macro invocations as unknown
// instructions (macro expansion itself is out of scope — see the docs). // instructions (macro expansion itself is out of scope, see the docs).
Macros map[string]bool Macros map[string]bool
} }
@@ -114,6 +114,7 @@ type Symbol struct {
Pkg string // package prefix before the middle dot ("" = current package) Pkg string // package prefix before the middle dot ("" = current package)
Name string // identifier without the middle dot or <> Name string // identifier without the middle dot or <>
Static bool // the <> marker is present Static bool // the <> marker is present
ABI string // the <NAME> ABI marker, e.g. ABIInternal ("" when absent)
Pseudo string // FP, SP, SB or PC ("" for a bare name) Pseudo string // FP, SP, SB or PC ("" for a bare name)
Offset int64 Offset int64
HasOff bool HasOff bool
@@ -156,5 +157,5 @@ type Address struct {
Scale int // index scale; 0 when absent Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base) Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<<2" Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
} }
+215 -3
View File
@@ -37,21 +37,33 @@ import (
// construction and are excluded from the diff; the other architectures list // construction and are excluded from the diff; the other architectures list
// their conditional branches outright. // their conditional branches outright.
func cmdAuditInstructions(args []string) error { func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [amd64|arm64|riscv64|loong64]", ` fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj), known-but-unencodable names (the backlog) and go- gasm asm --format goobj), known-but-unencodable names (the backlog) and go-
only names (feature gaps). The Go side is probed black-box with a battery only names (feature gaps). The Go side is probed black-box with a battery
of bare mnemonics, so the audit tracks whatever toolchain `+"`go env GOROOT`"+` of bare mnemonics, so the audit tracks whatever toolchain `+"`go env GOROOT`"+`
provides. provides.
With --corpus the audit changes shape: it assembles every .s file under the
given directory (default GOROOT/src) with the gasm encoder only, no
toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
`) `)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return err return err
} }
if *corpus {
return cmdAuditCorpus(fs.Args())
}
archName := "amd64" archName := "amd64"
switch n := len(fs.Args()); { switch n := len(fs.Args()); {
case n > 1: case n > 1:
return fmt.Errorf("audit-instructions takes at most one architecture argument") return &usageError{fmt.Errorf("audit-instructions takes at most one architecture argument")}
case n == 1: case n == 1:
archName = strings.ToLower(fs.Arg(0)) archName = strings.ToLower(fs.Arg(0))
} }
@@ -124,7 +136,7 @@ func auditArch(name string) (arch.Arch, error) {
case "loong64", "loong": case "loong64", "loong":
return arch.LOONG64, nil return arch.LOONG64, nil
} }
return arch.Unknown, fmt.Errorf("unknown architecture %q: want amd64, arm64, riscv64 or loong64", name) return arch.Unknown, &usageError{fmt.Errorf("unknown architecture %q: want amd64, arm64, riscv64 or loong64", name)}
} }
// goarchName maps an arch identifier onto its GOARCH spelling. // goarchName maps an arch identifier onto its GOARCH spelling.
@@ -205,6 +217,13 @@ func probeGoAsm(goarch string, names []string) (map[string]bool, error) {
cmd := exec.Command(asmBin, "-p", "probe", "-o", filepath.Join(dir, "probe.o"), probePath) cmd := exec.Command(asmBin, "-p", "probe", "-o", filepath.Join(dir, "probe.o"), probePath)
cmd.Env = append(os.Environ(), "GOARCH="+goarch, "GOOS="+runtime.GOOS) cmd.Env = append(os.Environ(), "GOARCH="+goarch, "GOOS="+runtime.GOOS)
out, _ := cmd.CombinedOutput() out, _ := cmd.CombinedOutput()
// The expected failure mode is a non-zero exit with compiler diagnostics
// on stdout; empty output means the probe broke at the exec level (a
// killed child, a tool that would not start), and seeding every name as
// recognized on that silence would fake a clean audit.
if len(out) == 0 {
return nil, fmt.Errorf("go tool asm probe for GOARCH=%s produced no output", goarch)
}
result := map[string]bool{} result := map[string]bool{}
for _, name := range names { for _, name := range names {
@@ -305,3 +324,196 @@ func gasmAssembles(a arch.Arch, name, shape string) bool {
func sanitize(name string) string { func sanitize(name string) string {
return strings.NewReplacer(".", "_", "$", "_").Replace(name) return strings.NewReplacer(".", "_", "$", "_").Replace(name)
} }
// --- corpus audit -----------------------------------------------------------
// corpusTarget is one architecture row of the corpus report.
type corpusTarget struct {
a arch.Arch
name string
}
// corpusTally accumulates one architecture's attempts over the corpus.
type corpusTally struct {
attempted int
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
}
func (t *corpusTally) fail(path, reason string) {
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
}
// cmdAuditCorpus implements audit-instructions --corpus.
func cmdAuditCorpus(args []string) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
root := ""
if len(args) == 1 {
root = args[0]
} else {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
return fmt.Errorf("locate GOROOT: %w", err)
}
root = filepath.Join(strings.TrimSpace(string(out)), "src")
}
stats, err := runCorpusAudit(root)
if err != nil {
return err
}
printCorpusStats(stats)
return nil
}
// corpusStats is the outcome of one corpus audit run.
type corpusStats struct {
root string
files int
generic int // files attempted for all four architectures
full int // files that assembled for every target architecture
targets []corpusTarget
tallies []*corpusTally
}
// runCorpusAudit assembles every .s file under root and returns the stats.
func runCorpusAudit(root string) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
return nil, err
}
targets := []corpusTarget{
{arch.AMD64, "amd64"},
{arch.ARM64, "arm64"},
{arch.RISCV, "riscv64"},
{arch.LOONG64, "loong64"},
}
tallies := make([]*corpusTally, len(targets))
for i := range tallies {
tallies[i] = &corpusTally{reasons: map[string]int{}, example: map[string]string{}}
}
// full is the north-star number: a file counts when every architecture
// its name allows assembles it.
full, generic := 0, 0
for _, path := range files {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
for i, tg := range targets {
if tg.a == a {
wanted = append(wanted, i)
}
}
} else {
generic++
for i := range targets {
wanted = append(wanted, i)
}
}
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
continue
}
t.assembled++
}
if ok && len(wanted) > 0 {
full++
}
}
return &corpusStats{
root: root,
files: len(files),
generic: generic,
full: full,
targets: targets,
tallies: tallies,
}, nil
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures)\n", s.root, s.files, s.generic)
fmt.Printf(" assemble for every target architecture: %d (%.1f%%)\n", s.full, 100*float64(s.full)/float64(max(s.files, 1)))
for i, tg := range s.targets {
t := s.tallies[i]
fmt.Printf(" %s: %d/%d attempted\n", tg.name, t.assembled, t.attempted)
for _, r := range topReasons(t) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
}
}
// corpusReason buckets an assembly or parse failure for the histogram.
func corpusReason(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
return "instruction not encodable"
case strings.Contains(msg, "undefined label"):
return "undefined label"
case strings.Contains(msg, "undefined symbol"), strings.Contains(msg, "external symbol"), strings.Contains(msg, "file-level assembly"):
return "undefined symbol or external"
case strings.Contains(msg, "operand"), strings.Contains(msg, "operand form"):
return "unsupported operand form"
default:
return "other: " + firstLine(msg)
}
}
// topReasons returns at most five reasons, most frequent first.
func topReasons(t *corpusTally) []string {
type kv struct {
k string
n int
}
var kvs []kv
for k, n := range t.reasons {
kvs = append(kvs, kv{k, n})
}
slices.SortFunc(kvs, func(a, b kv) int { return b.n - a.n })
if len(kvs) > 5 {
kvs = kvs[:5]
}
out := make([]string, len(kvs))
for i, kv := range kvs {
out[i] = kv.k
}
return out
}
// firstLine returns the first line of an error message, truncated.
func firstLine(msg string) string {
if i := strings.IndexByte(msg, '\n'); i >= 0 {
msg = msg[:i]
}
if len(msg) > 80 {
msg = msg[:80]
}
return msg
}
+9 -5
View File
@@ -53,7 +53,7 @@ REPL commands:
bufSpec := fs.String("buf", "", "buffer specification: name:size:pattern[,name:size:pattern...] where pattern is zero, ones, seq, or hex") bufSpec := fs.String("buf", "", "buffer specification: name:size:pattern[,name:size:pattern...] where pattern is zero, ones, seq, or hex")
script := fs.String("script", "", "run REPL commands from a file (one per line) and exit; '-' reads stdin") script := fs.String("script", "", "run REPL commands from a file (one per line) and exit; '-' reads stdin")
cover := fs.Bool("cover", false, "run to completion with a breakpoint on every instruction and report which executed and how often") cover := fs.Bool("cover", false, "run to completion with a breakpoint on every instruction and report which executed and how often")
timeout := fs.Duration("timeout", 0, "kill the debuggee after this duration (e.g. 30s); for headless --script runs") timeout := fs.Duration("timeout", 0, "kill the debuggee after this duration (e.g. 30s); for headless --script runs; a timeout exits 3")
fs.Parse(args) fs.Parse(args)
// --- Debuggee mode (internal, spawned by the debugger) --- // --- Debuggee mode (internal, spawned by the debugger) ---
@@ -84,7 +84,7 @@ REPL commands:
if *timeout > 0 { if *timeout > 0 {
go func() { go func() {
time.Sleep(*timeout) time.Sleep(*timeout)
fmt.Fprintf(os.Stderr, "gasm debug: timeout (%s) — killing the debuggee\n", *timeout) fmt.Fprintf(os.Stderr, "gasm debug: timeout (%s), killing the debuggee\n", *timeout)
os.Exit(3) os.Exit(3)
}() }()
} }
@@ -140,7 +140,6 @@ REPL commands:
} }
// Construct the argument block with buffer pointers at the correct positions. // Construct the argument block with buffer pointers at the correct positions.
bufIdx := 0
for _, arg := range layout { for _, arg := range layout {
if !arg.IsPtr { if !arg.IsPtr {
continue continue
@@ -170,12 +169,10 @@ REPL commands:
argBlock[off+16+j] = byte(size >> (j * 8)) argBlock[off+16+j] = byte(size >> (j * 8))
} }
} }
bufIdx++
break break
} }
} }
} }
_ = bufIdx
} else { } else {
argBlock = make([]byte, fl.Args) argBlock = make([]byte, fl.Args)
sess, err = debug.Launch("", path, *funcName, argBlock) sess, err = debug.Launch("", path, *funcName, argBlock)
@@ -234,6 +231,13 @@ REPL commands:
if sess.Exited() { if sess.Exited() {
break break
} }
// A genuine signal-delivery-stop (a fault in the kernel): the
// run cannot make progress, because resuming would restart the
// faulting instruction and fault forever. Report and stop.
if sig := sess.LastSignal(); sig != 0 {
fmt.Printf("gasm debug: cover: stopped on signal %v\n", sig)
break
}
regs, rerr := sess.GetRegs() regs, rerr := sess.GetRegs()
if rerr != nil { if rerr != nil {
break break
+148
View File
@@ -0,0 +1,148 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"fmt"
"os"
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// cmdDis disassembles machine code: either a raw binary (standard input with
// "-") whose architecture is given with -a, or a .s file, which is assembled
// first so the listing shows the real function and label layout.
func cmdDis(args []string) int {
fs := newCommand("dis", "gasm dis [-a arch] <file>", `
Disassemble machine code to instruction text (via golang.org/x/arch).
With a .s file, the file is assembled first and the listing follows the
real layout: one block per TEXT function, local labels printed at their
offsets. The architecture comes from the file name suffix, or from -a.
With any other file, or "-" for standard input, the bytes are disassembled
linearly and -a selects the architecture (amd64, arm64, riscv64 or
loong64).
`)
archName := fs.String("a", "", "architecture for raw input: amd64, arm64, riscv64 or loong64")
fs.Parse(args)
if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm dis [-a arch] <file>")
return 2
}
path := fs.Arg(0)
var target arch.Arch
if *archName != "" {
var err error
target, err = auditArch(*archName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 2
}
}
if strings.HasSuffix(path, ".s") {
if target == arch.Unknown {
target = arch.FromFilename(path)
}
if target == arch.Unknown {
fmt.Fprintln(os.Stderr, "gasm dis: cannot infer the architecture from the file name; use -a")
return 2
}
return disSource(path, target)
}
if target == arch.Unknown {
fmt.Fprintln(os.Stderr, "gasm dis: raw input needs -a (amd64, arm64, riscv64 or loong64)")
return 2
}
src, err := readSource(path)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm dis:", err)
return 1
}
printListing(target, []byte(src), 0, nil)
return 0
}
// disSource assembles a .s file and prints one listing block per function.
func disSource(path string, target arch.Arch) int {
src, err := readSource(path)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm dis:", err)
return 1
}
f, errs := parser.Parse(path, src)
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return 1
}
img, err := assembleFile(target, f)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 1
}
if len(img.Funcs) == 0 {
fmt.Fprintln(os.Stderr, "gasm dis: no assemblable TEXT functions found")
return 1
}
for _, fn := range img.Funcs {
code := img.Code[fn.Offset : fn.Offset+fn.Size]
fmt.Printf("%s: %d bytes\n", fn.Name, fn.Size)
labels := make(map[int][]string, len(fn.Labels))
for name, off := range fn.Labels {
labels[off] = append(labels[off], name)
}
for off := range labels {
sort.Strings(labels[off])
}
printListing(target, code, uint64(fn.Offset), labels)
}
if len(img.Data) > 0 {
fmt.Printf("data: %d bytes at 0x%x\n", len(img.Data), len(img.Code))
}
return 0
}
// printListing decodes code linearly from offset base, printing label lines
// (label name to offset within the block) as they are reached.
func printListing(a arch.Arch, code []byte, base uint64, labels map[int][]string) {
pc := 0
for pc < len(code) {
for _, name := range labels[pc] {
fmt.Printf("%s:\n", name)
}
ins, err := disasm.Decode(a, code[pc:], base+uint64(pc))
if err != nil {
break
}
end := min(pc+ins.Len, len(code))
fmt.Printf(" %04x: %-16s %s\n", base+uint64(pc), hexBytes(code[pc:end]), ins.Text)
if ins.Len <= 0 {
break
}
pc += ins.Len
}
}
// hexBytes renders up to 8 bytes as contiguous hex.
func hexBytes(b []byte) string {
var sb strings.Builder
for i, c := range b {
if i == 8 {
break
}
if i > 0 {
sb.WriteByte(' ')
}
fmt.Fprintf(&sb, "%02x", c)
}
return sb.String()
}
+286 -369
View File
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org) // Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause // SPDX-License-Identifier: BSD-3-Clause
// Command gasm is the developer frontend for GAsm — Go's Plan 9 assembler. // Command gasm is the developer frontend for GAsm, Go's Plan 9 assembler.
// It bundles a token dumper, a parser, a formatter, a linter and a language // It bundles a token dumper, a parser, a formatter, a linter and a language
// server into one binary. Every subcommand works headlessly so it can be // server into one binary. Every subcommand works headlessly so it can be
// driven from scripts and CI as well as from an editor. // driven from scripts and CI as well as from an editor.
@@ -10,6 +10,7 @@ package main
import ( import (
"bytes" "bytes"
"encoding/json" "encoding/json"
"errors"
"flag" "flag"
"fmt" "fmt"
"io" "io"
@@ -18,6 +19,7 @@ import (
"os/exec" "os/exec"
"path/filepath" "path/filepath"
"runtime" "runtime"
"runtime/debug"
"slices" "slices"
"sort" "sort"
"strconv" "strconv"
@@ -36,9 +38,32 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/verify" "sourcedock.dev/petrbalvin/gasm-devkit/verify"
) )
// version is the release version, stamped at build time via // version reports the release the toolchain recorded for this build: the
// -ldflags "-X main.version=…" (defaulting to the current release). // tag on a tag, a pseudo-version below one, and (devel) outside version
var version = "0.32.0" // control. Nothing is injected; the recorded value cannot go stale.
func version() string {
bi, ok := debug.ReadBuildInfo()
if !ok || bi.Main.Version == "" {
return "(devel)"
}
return bi.Main.Version
}
// usageError marks an error the caller's arguments caused, which exits 2
// instead of the 1 a runtime failure gets.
type usageError struct{ err error }
func (e *usageError) Error() string { return e.err.Error() }
func (e *usageError) Unwrap() error { return e.err }
// exitCodeFor maps an error onto the process exit status: 2 for a usage
// error, 1 for anything else.
func exitCodeFor(err error) int {
if _, ok := errors.AsType[*usageError](err); ok {
return 2
}
return 1
}
func main() { func main() {
if len(os.Args) < 2 { if len(os.Args) < 2 {
@@ -56,6 +81,8 @@ func main() {
os.Exit(cmdLint(os.Args[2:])) os.Exit(cmdLint(os.Args[2:]))
case "asm": case "asm":
os.Exit(cmdAsm(os.Args[2:])) os.Exit(cmdAsm(os.Args[2:]))
case "dis":
os.Exit(cmdDis(os.Args[2:]))
case "verify": case "verify":
os.Exit(cmdVerify(os.Args[2:])) os.Exit(cmdVerify(os.Args[2:]))
case "debug": case "debug":
@@ -67,12 +94,12 @@ func main() {
case "audit-instructions": case "audit-instructions":
if err := cmdAuditInstructions(os.Args[2:]); err != nil { if err := cmdAuditInstructions(os.Args[2:]); err != nil {
fmt.Fprintln(os.Stderr, err) fmt.Fprintln(os.Stderr, err)
os.Exit(1) os.Exit(exitCodeFor(err))
} }
case "scaffold": case "scaffold":
if err := cmdScaffold(os.Args[2:]); err != nil { if err := cmdScaffold(os.Args[2:]); err != nil {
fmt.Fprintln(os.Stderr, err) fmt.Fprintln(os.Stderr, err)
os.Exit(1) os.Exit(exitCodeFor(err))
} }
case "lsp": case "lsp":
os.Exit(cmdLSP(os.Args[2:])) os.Exit(cmdLSP(os.Args[2:]))
@@ -81,27 +108,27 @@ func main() {
case "help", "--help", "-h": case "help", "--help", "-h":
usage(os.Stdout) usage(os.Stdout)
default: default:
fmt.Fprintf(os.Stderr, "gasm: unknown command %q — run \"gasm --help\" for usage\n", os.Args[1]) fmt.Fprintf(os.Stderr, "gasm: unknown command %q; run \"gasm --help\" for usage\n", os.Args[1])
os.Exit(2) os.Exit(2)
} }
} }
// cmdVersion prints the release version. // cmdVersion prints the recorded version.
func cmdVersion() int { func cmdVersion() int {
fmt.Printf("gasm %s\n", version) fmt.Printf("gasm %s\n", version())
return 0 return 0
} }
// ANSI color helpers for terminal output. // ANSI colour helpers for terminal output.
const ( const (
colorReset = "\033[0m" colourReset = "\033[0m"
colorBold = "\033[1m" colourBold = "\033[1m"
colorCyan = "\033[36m" colourCyan = "\033[36m"
colorYellow = "\033[33m" colourYellow = "\033[33m"
colorGray = "\033[90m" colourGrey = "\033[90m"
) )
// isTTY reports whether the writer is a terminal (for color output). // isTTY reports whether the writer is a terminal (for colour output).
func isTTY(w io.Writer) bool { func isTTY(w io.Writer) bool {
if f, ok := w.(*os.File); ok { if f, ok := w.(*os.File); ok {
stat, _ := f.Stat() stat, _ := f.Stat()
@@ -112,12 +139,12 @@ func isTTY(w io.Writer) bool {
func usage(w io.Writer) { func usage(w io.Writer) {
useColor := isTTY(w) useColor := isTTY(w)
bold, cyan, yellow, gray, reset := "", "", "", "", "" bold, cyan, yellow, grey, reset := "", "", "", "", ""
if useColor { if useColor {
bold, cyan, yellow, gray, reset = colorBold, colorCyan, colorYellow, colorGray, colorReset bold, cyan, yellow, grey, reset = colourBold, colourCyan, colourYellow, colourGrey, colourReset
} }
fmt.Fprintf(w, "%sgasm %s%s — developer tooling for Go's Plan 9 assembler (GAsm)%s\n\n", bold, version, reset, reset) fmt.Fprintf(w, "%sgasm %s%s: developer tooling for Go's Plan 9 assembler (GAsm)%s\n\n", bold, version(), reset, reset)
fmt.Fprintf(w, "gasm bundles a lexer, parser, formatter, linter, standalone assembler and\n") fmt.Fprintf(w, "gasm bundles a lexer, parser, formatter, linter, standalone assembler and\n")
fmt.Fprintf(w, "language server for Plan 9 assembly into one self-contained binary.\n\n") fmt.Fprintf(w, "language server for Plan 9 assembly into one self-contained binary.\n\n")
@@ -132,6 +159,7 @@ func usage(w io.Writer) {
{"fmt", "canonicalise formatting (gofmt for assembly)"}, {"fmt", "canonicalise formatting (gofmt for assembly)"},
{"lint", "run static checks"}, {"lint", "run static checks"},
{"asm", "assemble .s files to machine code (amd64, arm64, riscv64, loong64)"}, {"asm", "assemble .s files to machine code (amd64, arm64, riscv64, loong64)"},
{"dis", "disassemble machine code (raw bytes or an assembled .s file)"},
{"verify", "JIT-assemble and run dynamic checks (amd64, arm64, riscv64, loong64)"}, {"verify", "JIT-assemble and run dynamic checks (amd64, arm64, riscv64, loong64)"},
{"debug", "interactive source-level debugger (amd64, arm64, riscv64, loong64)"}, {"debug", "interactive source-level debugger (amd64, arm64, riscv64, loong64)"},
{"diff", "compare machine code of two .s files"}, {"diff", "compare machine code of two .s files"},
@@ -142,12 +170,12 @@ func usage(w io.Writer) {
{"version", "print the version (same as --version)"}, {"version", "print the version (same as --version)"},
} }
for _, c := range commands { for _, c := range commands {
fmt.Fprintf(w, " %s%-10s%s %s%s%s\n", cyan, c.name, reset, gray, c.desc, reset) fmt.Fprintf(w, " %s%-10s%s %s%s%s\n", cyan, c.name, reset, grey, c.desc, reset)
} }
fmt.Fprintf(w, "\n%sFlags:%s\n", yellow, reset) fmt.Fprintf(w, "\n%sFlags:%s\n", yellow, reset)
fmt.Fprintf(w, " %s-h, --help%s %sshow this help%s\n", cyan, reset, gray, reset) fmt.Fprintf(w, " %s-h, --help%s %sshow this help%s\n", cyan, reset, grey, reset)
fmt.Fprintf(w, " %s-V, --version%s %sprint the version%s\n", cyan, reset, gray, reset) fmt.Fprintf(w, " %s-V, --version%s %sprint the version%s\n", cyan, reset, grey, reset)
fmt.Fprintf(w, "\nRun \"gasm <command> -h\" for a command's usage and flags.\n\n") fmt.Fprintf(w, "\nRun \"gasm <command> -h\" for a command's usage and flags.\n\n")
@@ -161,7 +189,7 @@ func usage(w io.Writer) {
} }
for _, e := range examples { for _, e := range examples {
if e.desc != "" { if e.desc != "" {
fmt.Fprintf(w, " %s%s%s %s%s%s\n", cyan, e.cmd, reset, gray, e.desc, reset) fmt.Fprintf(w, " %s%s%s %s%s%s\n", cyan, e.cmd, reset, grey, e.desc, reset)
} else { } else {
fmt.Fprintf(w, " %s%s%s\n", cyan, e.cmd, reset) fmt.Fprintf(w, " %s%s%s\n", cyan, e.cmd, reset)
} }
@@ -263,24 +291,34 @@ standard input.
funcs++ funcs++
} }
} }
fmt.Printf("%s: OK — %d declarations, %d functions\n", path, len(file.Decls), funcs) fmt.Printf("%s: OK, %d declarations, %d functions\n", path, len(file.Decls), funcs)
return 0 return 0
} }
func cmdFmt(args []string) int { func cmdFmt(args []string) int {
fs := newCommand("fmt", "gasm fmt [-w] [path...]", ` fs := newCommand("fmt", "gasm fmt [-w|-l|-d] [path...]", `
Canonicalise the formatting of Plan 9 assembly sources: indentation, operand Canonicalise the formatting of Plan 9 assembly sources: indentation, operand
spacing, per-function mnemonic alignment and blank-line layout (exactly one spacing, per-function mnemonic alignment and blank-line layout (exactly one
blank line before each label, TEXT and GLOBL block). Formatting is blank line before each label, TEXT and GLOBL block). Formatting is
idempotent and preserves every line, comments included. idempotent and preserves every line, comments included.
With no paths — or a directory path — every .s file below it is reformatted With no paths, or a directory path, every .s file below it is reformatted
in place and the changed files are listed, the way go fmt does; "." and "_" in place and the changed files are listed, the way go fmt does; "." and "_"
directories are skipped. Explicit file paths print to stdout unless -w is directories are skipped. Explicit file paths print to stdout unless -w is
given. given.
-l and -d rewrite nothing: -l prints the paths whose formatting differs
from gasm's (empty output means everything is formatted, which is what a CI
check wants), -d prints the diffs. They are mutually exclusive.
`) `)
write := fs.Bool("w", false, "write result to the source file") write := fs.Bool("w", false, "write result to the source file")
list := fs.Bool("l", false, "list files whose formatting differs from gasm's")
diffMode := fs.Bool("d", false, "print diffs instead of rewriting files")
fs.Parse(args) fs.Parse(args)
if *list && *diffMode {
fmt.Fprintln(os.Stderr, "gasm fmt: -l and -d are mutually exclusive")
return 2
}
// Like go fmt: with no arguments, or with a directory argument, every .s // Like go fmt: with no arguments, or with a directory argument, every .s
// file below the directory is formatted in place and the names of the // file below the directory is formatted in place and the names of the
// changed files are listed; explicit file arguments keep the -w / stdout // changed files are listed; explicit file arguments keep the -w / stdout
@@ -318,6 +356,16 @@ given.
continue continue
} }
out := format.Source(src) out := format.Source(src)
if *list || *diffMode {
if out != src {
if *list {
fmt.Println(path)
} else {
fmt.Print(unifiedDiff(path, strings.Split(src, "\n"), strings.Split(out, "\n")))
}
}
continue
}
if dirMode || *write { if dirMode || *write {
if out != src { if out != src {
if err := os.WriteFile(path, []byte(out), 0o644); err != nil { if err := os.WriteFile(path, []byte(out), 0o644); err != nil {
@@ -337,7 +385,7 @@ given.
} }
// asmFiles collects the .s files below dir, skipping directories whose name // asmFiles collects the .s files below dir, skipping directories whose name
// starts with "." or "_" — as the go tooling does, which keeps .git and // starts with "." or "_", as the go tooling does, which keeps .git and
// scratch or reference trees (e.g. _refs) untouched. // scratch or reference trees (e.g. _refs) untouched.
func asmFiles(dir string) ([]string, error) { func asmFiles(dir string) ([]string, error) {
var out []string var out []string
@@ -419,6 +467,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
`) `)
fs.Parse(args) fs.Parse(args)
srv := lsp.New(os.Stdin, os.Stdout) srv := lsp.New(os.Stdin, os.Stdout)
srv.SetVersion(version())
if err := srv.Run(); err != nil { if err := srv.Run(); err != nil {
fmt.Fprintln(os.Stderr, "gasm lsp:", err) fmt.Fprintln(os.Stderr, "gasm lsp:", err)
return 1 return 1
@@ -427,7 +476,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
} }
func cmdAsm(args []string) int { func cmdAsm(args []string) int {
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-o out] <file>", ` fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>", `
Assemble FILE without the Go toolchain: every TEXT function is encoded to Assemble FILE without the Go toolchain: every TEXT function is encoded to
machine code and printed as a hex dump. Supported architectures: amd64 machine code and printed as a hex dump. Supported architectures: amd64
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP, (including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
@@ -445,13 +494,30 @@ requires -p, the package path, and the installed Go toolchain).
out := fs.String("o", "", "write the output to this file") out := fs.String("o", "", "write the output to this file")
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)") format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)") pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
fs.Parse(args) fs.Parse(args)
if fs.NArg() != 1 { if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-o out] <file>") fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>")
return 2
}
// The format is validated before anything else, so a bogus value exits 2
// with or without -o instead of silently dumping the hex of a raw image.
switch *format {
case "raw", "elf", "goobj":
default:
fmt.Fprintf(os.Stderr, "gasm asm: unknown format %q (want raw, elf or goobj)\n", *format)
return 2 return 2
} }
path := fs.Arg(0) path := fs.Arg(0)
targetArch := arch.FromFilename(path) targetArch := arch.FromFilename(path)
if *archName != "" {
a, err := auditArch(*archName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
return 2
}
targetArch = a
}
src, err := readSource(path) src, err := readSource(path)
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "gasm:", err) fmt.Fprintln(os.Stderr, "gasm:", err)
@@ -470,42 +536,48 @@ requires -p, the package path, and the installed Go toolchain).
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err) fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1 return 1
} }
if len(img.Funcs) == 0 { if len(img.Funcs) == 0 && len(img.Data) == 0 {
fmt.Fprintln(os.Stderr, "gasm asm: no assemblable TEXT functions found") // A file with neither code nor data assembles to nothing, which is
// almost always a wrong architecture rather than an intent.
fmt.Fprintln(os.Stderr, "gasm asm: no assemblable TEXT functions or GLOBL data found")
return 1 return 1
} }
for _, fn := range img.Funcs { // Without -o the hex dump on stdout is the output; with -o the file is,
code := img.Code[fn.Offset : fn.Offset+fn.Size] // and the dump is skipped, as the -o help text promises.
fmt.Printf("%s: %d bytes\n", fn.Name, fn.Size) if *out == "" {
for i := 0; i < len(code); i += 16 { for _, fn := range img.Funcs {
end := min(i+16, len(code)) code := img.Code[fn.Offset : fn.Offset+fn.Size]
fmt.Printf(" %04x:", i) fmt.Printf("%s: %d bytes\n", fn.Name, fn.Size)
for _, b := range code[i:end] { for i := 0; i < len(code); i += 16 {
fmt.Printf(" %02x", b) end := min(i+16, len(code))
fmt.Printf(" %04x:", i)
for _, b := range code[i:end] {
fmt.Printf(" %02x", b)
}
fmt.Println()
} }
fmt.Println()
} }
} if len(img.Data) > 0 {
if len(img.Data) > 0 { fmt.Printf("data: %d bytes at 0x%x\n", len(img.Data), len(img.Code))
fmt.Printf("data: %d bytes at 0x%x\n", len(img.Data), len(img.Code)) for _, d := range f.Decls {
for _, d := range f.Decls { g, ok := d.(*ast.Globl)
g, ok := d.(*ast.Globl) if !ok || g.Name == nil || g.Name.Pseudo != "SB" {
if !ok || g.Name == nil || g.Name.Pseudo != "SB" { continue
continue }
size := 0
if g.Size != nil && g.Size.Imm.HasVal {
size = int(g.Size.Imm.Val)
}
fmt.Printf(" %s: %d bytes at 0x%x\n", g.Name.Name, size, img.Symbols[g.Name.Name])
} }
size := 0 for i := 0; i < len(img.Data); i += 16 {
if g.Size != nil && g.Size.Imm.HasVal { end := min(i+16, len(img.Data))
size = int(g.Size.Imm.Val) fmt.Printf(" %04x:", len(img.Code)+i)
for _, b := range img.Data[i:end] {
fmt.Printf(" %02x", b)
}
fmt.Println()
} }
fmt.Printf(" %s: %d bytes at 0x%x\n", g.Name.Name, size, img.Symbols[g.Name.Name])
}
for i := 0; i < len(img.Data); i += 16 {
end := min(i+16, len(img.Data))
fmt.Printf(" %04x:", len(img.Code)+i)
for _, b := range img.Data[i:end] {
fmt.Printf(" %02x", b)
}
fmt.Println()
} }
} }
if *out != "" { if *out != "" {
@@ -543,9 +615,6 @@ requires -p, the package path, and the installed Go toolchain).
obj, err = img.GOObject(*pkg, path) obj, err = img.GOObject(*pkg, path)
} }
kind = "Go object" kind = "Go object"
default:
fmt.Fprintf(os.Stderr, "gasm asm: unknown format %q (want raw, elf or goobj)\n", *format)
return 2
} }
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "gasm asm:", err) fmt.Fprintln(os.Stderr, "gasm asm:", err)
@@ -562,7 +631,7 @@ requires -p, the package path, and the installed Go toolchain).
// cmdDiff compares the machine code of two assembly files. // cmdDiff compares the machine code of two assembly files.
func cmdDiff(args []string) int { func cmdDiff(args []string) int {
set := newCommand("diff", "gasm diff <file1.s> <file2.s>", ` set := newCommand("diff", "gasm diff [-GOARCH arch] <file1.s> <file2.s>", `
Compare the machine code produced by assembling two files. Compare the machine code produced by assembling two files.
Shows which functions differ and the byte-level differences. Shows which functions differ and the byte-level differences.
Useful for verifying that two implementations produce identical code, Useful for verifying that two implementations produce identical code,
@@ -572,12 +641,22 @@ Use --map to compare functions whose names differ between the files,
e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix. e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
`) `)
mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names") mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names")
archName := set.String("GOARCH", "", "target architecture for both files: amd64, arm64, riscv64 or loong64")
set.Parse(args) set.Parse(args)
if set.NArg() != 2 { if set.NArg() != 2 {
fmt.Fprintln(os.Stderr, "usage: gasm diff <file1.s> <file2.s>") fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>")
return 2 return 2
} }
path1, path2 := set.Arg(0), set.Arg(1) path1, path2 := set.Arg(0), set.Arg(1)
forced := arch.Unknown
if *archName != "" {
a, err := auditArch(*archName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %v\n", err)
return 2
}
forced = a
}
// Parse the name mapping (file1 name → file2 name). // Parse the name mapping (file1 name → file2 name).
nameMap := make(map[string]string) nameMap := make(map[string]string)
@@ -593,12 +672,12 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
} }
// Assemble both files. // Assemble both files.
img1, err := assemblePath(path1) img1, err := assemblePath(path1, forced)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err) fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err)
return 1 return 1
} }
img2, err := assemblePath(path2) img2, err := assemblePath(path2, forced)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err) fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err)
return 1 return 1
@@ -673,8 +752,9 @@ func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
} }
} }
// assemblePath reads, parses and assembles a file (used by cmdDiff). // assemblePath reads, parses and assembles a file (used by cmdDiff). A
func assemblePath(path string) (*asm.Image, error) { // non-Unknown forced architecture overrides the file-name suffix.
func assemblePath(path string, forced arch.Arch) (*asm.Image, error) {
src, err := readSource(path) src, err := readSource(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -686,7 +766,11 @@ func assemblePath(path string) (*asm.Image, error) {
if len(errs) > 0 { if len(errs) > 0 {
return nil, fmt.Errorf("parse errors") return nil, fmt.Errorf("parse errors")
} }
return assembleFile(arch.FromFilename(path), f) target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
} }
// printByteDiff shows the first few byte differences between two code blocks. // printByteDiff shows the first few byte differences between two code blocks.
@@ -765,9 +849,11 @@ gasm verify --fuzz which exercises the code paths.
return 0 return 0
} }
// cmdVerifyRISCV handles the verify subcommand for RISC-V files. // cmdVerifyNonJIT handles the verify subcommand for files whose architecture
// JIT requires RISC-V hardware; only ground-truth and profile are available. // the host cannot execute: only the ground-truth comparison and the static
func cmdVerifyRISCV(path string, groundTruth, profile bool) int { // profile are available there. Relocation sites are masked before the byte
// comparison, as the toolchain leaves them zero for the linker.
func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile bool) int {
src, err := readSource(path) src, err := readSource(path)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err) fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
@@ -780,63 +866,34 @@ func cmdVerifyRISCV(path string, groundTruth, profile bool) int {
if len(errs) > 0 { if len(errs) > 0 {
return 1 return 1
} }
img, err := asm.AssembleFileRISCV(f) img, err := assembleFile(targetArch, f)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err) fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1 return 1
} }
if groundTruth { if groundTruth {
gt, err := verify.GroundTruthRISCV(path) var gt map[string][]byte
switch targetArch {
case arch.RISCV:
gt, err = verify.GroundTruthRISCV(path)
case arch.LOONG64:
gt, err = verify.GroundTruthLOONG64(path)
case arch.ARM64:
gt, err = verify.GroundTruthARM64(path)
default:
gt, err = verify.GroundTruth(path)
}
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err) fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err)
return 1 return 1
} }
matched, total := 0, 0 matched, total, diffs := compareGroundTruth(img, gt)
for _, fn := range img.Funcs { if diffs > 0 {
gasmCode := img.Code[fn.Offset : fn.Offset+fn.Size] printCodeDiff(img, gt)
goCode, ok := gt[fn.Name]
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", fn.Name)
continue
}
total++
gasmCmp := make([]byte, len(gasmCode))
goCmp := make([]byte, len(goCode))
copy(gasmCmp, gasmCode)
copy(goCmp, goCode)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(gasmCmp); j++ {
gasmCmp[j] = 0
}
for j := r.Off; j < r.Off+4 && j < len(goCmp); j++ {
goCmp[j] = 0
}
}
if bytes.Equal(gasmCmp, goCmp) {
matched++
if len(fn.Relocs) > 0 {
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs))
} else {
fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size)
}
} else {
fmt.Printf(" %s: MISMATCH (%d vs %d bytes)\n", fn.Name, fn.Size, len(goCode))
for i := 0; i < len(gasmCode) || i < len(goCode); i += 16 {
var gb, gs string
for j := i; j < i+16 && j < len(gasmCode); j++ {
gb += fmt.Sprintf(" %02x", gasmCode[j])
}
for j := i; j < i+16 && j < len(goCode); j++ {
gs += fmt.Sprintf(" %02x", goCode[j])
}
fmt.Printf(" %04x: gasm:%s\n", i, gb)
fmt.Printf(" %04x: gt: %s\n", i, gs)
}
}
} }
fmt.Printf("%s: %d/%d matched\n", path, matched, total) fmt.Printf("%s: %d/%d matched\n", path, matched, total)
if matched < total { if matched < total || diffs > 0 {
return 1 return 1
} }
return 0 return 0
@@ -856,193 +913,103 @@ func cmdVerifyRISCV(path string, groundTruth, profile bool) int {
return 0 return 0
} }
// cmdVerifyLOONG64 verifies a loong64 source file against `go tool asm` // compareGroundTruth compares the image's functions against the go tool asm
// (GOARCH=loong64) — the ground-truth oracle — since gasm cannot JIT-load // output byte-for-byte, masking relocation sites (disp32 fields the Go linker
// LoongArch code on an amd64 host. Relocation sites are masked before the // fills at link time). It prints one line per function and returns the
// byte comparison, as the toolchain leaves them zero for the linker. // matched and compared counts plus the number of functions with byte diffs.
func cmdVerifyLOONG64(path string, groundTruth, profile bool) int { func compareGroundTruth(img *asm.Image, gt map[string][]byte) (matched, total, diffs int) {
src, err := readSource(path)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
}
f, errs := parser.Parse(path, src)
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return 1
}
img, err := asm.AssembleFileLOONG64(f)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
}
if groundTruth {
gt, err := verify.GroundTruthLOONG64(path)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err)
return 1
}
matched, total := 0, 0
for _, fn := range img.Funcs {
gasmCode := img.Code[fn.Offset : fn.Offset+fn.Size]
goCode, ok := gt[fn.Name]
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", fn.Name)
continue
}
total++
gasmCmp := make([]byte, len(gasmCode))
goCmp := make([]byte, len(goCode))
copy(gasmCmp, gasmCode)
copy(goCmp, goCode)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(gasmCmp); j++ {
gasmCmp[j] = 0
}
for j := r.Off; j < r.Off+4 && j < len(goCmp); j++ {
goCmp[j] = 0
}
}
if bytes.Equal(gasmCmp, goCmp) {
matched++
if len(fn.Relocs) > 0 {
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs))
} else {
fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size)
}
} else {
fmt.Printf(" %s: MISMATCH (%d vs %d bytes)\n", fn.Name, fn.Size, len(goCode))
for i := 0; i < len(gasmCode) || i < len(goCode); i += 16 {
var gb, gs string
for j := i; j < i+16 && j < len(gasmCode); j++ {
gb += fmt.Sprintf(" %02x", gasmCode[j])
}
for j := i; j < i+16 && j < len(goCode); j++ {
gs += fmt.Sprintf(" %02x", goCode[j])
}
fmt.Printf(" %04x: gasm:%s\n", i, gb)
fmt.Printf(" %04x: gt: %s\n", i, gs)
}
}
}
fmt.Printf("%s: %d/%d matched\n", path, matched, total)
if matched < total {
return 1
}
return 0
}
if profile {
for _, fn := range img.Funcs {
fmt.Printf("%s: %d bytes, labels: %v\n", fn.Name, fn.Size, fn.Labels)
}
return 0
}
fmt.Printf("%s: %d functions assembled\n", path, len(img.Funcs))
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
fmt.Printf(" %s: %d bytes\n", fn.Name, fn.Size) gasmCode := img.Code[fn.Offset : fn.Offset+fn.Size]
goCode, ok := gt[fn.Name]
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", fn.Name)
continue
}
total++
gasmCmp := make([]byte, len(gasmCode))
goCmp := make([]byte, len(goCode))
copy(gasmCmp, gasmCode)
copy(goCmp, goCode)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(gasmCmp); j++ {
gasmCmp[j] = 0
}
for j := r.Off; j < r.Off+4 && j < len(goCmp); j++ {
goCmp[j] = 0
}
}
// The toolchain pads text symbols to 16-byte boundaries with
// zeros, so a function whose size is not a multiple of 16
// carries trailing zeros in the ground truth that are not part
// of the encoding. Compare up to the shorter side and require
// the remainder of whichever is longer to be zero, so padding
// never masks a real difference.
cmpLen := min(len(gasmCmp), len(goCmp))
equal := bytes.Equal(gasmCmp[:cmpLen], goCmp[:cmpLen])
if equal {
for _, b := range gasmCmp[cmpLen:] {
if b != 0 {
equal = false
break
}
}
}
if equal {
for _, b := range goCmp[cmpLen:] {
if b != 0 {
equal = false
break
}
}
}
if equal {
matched++
switch {
case len(fn.Relocs) > 0 && len(goCmp) > cmpLen:
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked, %d padding)\n", fn.Name, fn.Size, len(fn.Relocs), len(goCmp)-cmpLen)
case len(fn.Relocs) > 0:
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs))
case len(goCmp) > cmpLen:
fmt.Printf(" %s: MATCH (%d bytes, %d padding)\n", fn.Name, fn.Size, len(goCmp)-cmpLen)
default:
fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size)
}
} else {
diffs++
fmt.Printf(" %s: MISMATCH (%d vs %d bytes)\n", fn.Name, fn.Size, len(goCode))
}
} }
return 0 return matched, total, diffs
} }
func cmdVerifyARM64(path string, groundTruth, profile bool) int { // printCodeDiff shows a 16-byte hex dump per function whose gasm bytes differ
src, err := readSource(path) // from the go tool asm output.
if err != nil { func printCodeDiff(img *asm.Image, gt map[string][]byte) {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
}
f, errs := parser.Parse(path, src)
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return 1
}
img, err := asm.AssembleFileARM64(f)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
}
if groundTruth {
gt, err := verify.GroundTruthARM64(path)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err)
return 1
}
matched, total := 0, 0
for _, fn := range img.Funcs {
gasmCode := img.Code[fn.Offset : fn.Offset+fn.Size]
goCode, ok := gt[fn.Name]
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", fn.Name)
continue
}
total++
gasmCmp := make([]byte, len(gasmCode))
goCmp := make([]byte, len(goCode))
copy(gasmCmp, gasmCode)
copy(goCmp, goCode)
for _, r := range fn.Relocs {
for j := r.Off; j < r.Off+4 && j < len(gasmCmp); j++ {
gasmCmp[j] = 0
}
for j := r.Off; j < r.Off+4 && j < len(goCmp); j++ {
goCmp[j] = 0
}
}
if bytes.Equal(gasmCmp, goCmp) {
matched++
if len(fn.Relocs) > 0 {
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs))
} else {
fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size)
}
} else {
fmt.Printf(" %s: MISMATCH (%d vs %d bytes)\n", fn.Name, fn.Size, len(goCode))
for i := 0; i < len(gasmCode) || i < len(goCode); i += 16 {
var gb, gs string
for j := i; j < i+16 && j < len(gasmCode); j++ {
gb += fmt.Sprintf(" %02x", gasmCode[j])
}
for j := i; j < i+16 && j < len(goCode); j++ {
gs += fmt.Sprintf(" %02x", goCode[j])
}
fmt.Printf(" %04x: gasm:%s\n", i, gb)
fmt.Printf(" %04x: gt: %s\n", i, gs)
}
}
}
fmt.Printf("%s: %d/%d matched\n", path, matched, total)
if matched < total {
return 1
}
return 0
}
if profile {
for _, fn := range img.Funcs {
fmt.Printf("%s: %d bytes, labels: %v\n", fn.Name, fn.Size, fn.Labels)
}
return 0
}
fmt.Printf("%s: %d functions assembled\n", path, len(img.Funcs))
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
fmt.Printf(" %s: %d bytes\n", fn.Name, fn.Size) gasmCode := img.Code[fn.Offset : fn.Offset+fn.Size]
goCode, ok := gt[fn.Name]
if !ok || bytes.Equal(gasmCode, goCode) {
continue
}
for i := 0; i < len(gasmCode) || i < len(goCode); i += 16 {
var gb, gs string
for j := i; j < i+16 && j < len(gasmCode); j++ {
gb += fmt.Sprintf(" %02x", gasmCode[j])
}
for j := i; j < i+16 && j < len(goCode); j++ {
gs += fmt.Sprintf(" %02x", goCode[j])
}
fmt.Printf(" %04x: gasm:%s\n", i, gb)
fmt.Printf(" %04x: gt: %s\n", i, gs)
}
} }
return 0
} }
func cmdVerify(args []string) int { func cmdVerify(args []string) int {
set := newCommand("verify", "gasm verify [-smoke] [-abi] [-fuzz] [-ground-truth] [-profile] [-call] <file.s>", ` set := newCommand("verify", "gasm verify [-smoke] [-abi] [-fuzz] [-ground-truth] [-profile] [-call] <file.s>", `
Assemble FILE (amd64), map it into executable memory and report the available Assemble FILE (amd64), map it into executable memory and report the available
functions. This confirms the assembled image is self-consistent (no functions. This confirms the assembled image is self-consistent (no
unresolved external symbols) and executable — the prerequisite for dynamic unresolved external symbols) and executable, the prerequisite for dynamic
testing. testing.
With -smoke, each NOSPLIT function is called with a zeroed argument block to With -smoke, each NOSPLIT function is called with a zeroed argument block to
@@ -1052,8 +1019,8 @@ that tolerate nil pointers and zero lengths in their arguments.
With -abi, each function is called with sentinel values in the registers With -abi, each function is called with sentinel values in the registers
the Go ABI fixes across calls (the frame pointer and the goroutine the Go ABI fixes across calls (the frame pointer and the goroutine
pointer) plus a canary below SP; violations are reported. JIT-based pointer) plus a canary below SP; violations are reported. JIT-based
checks run when the host matches the file's architecture (all but checks run when the host matches the file's architecture, on all four
loong64, which is ground-truth only for now). architectures.
With -fuzz, each function with a // func signature is differentially fuzzed With -fuzz, each function with a // func signature is differentially fuzzed
against the go-tool-asm version in a subprocess (so a crash on a partial against the go-tool-asm version in a subprocess (so a crash on a partial
@@ -1095,25 +1062,16 @@ each entry reproduces.
path := set.Arg(0) path := set.Arg(0)
targetArch := arch.FromFilename(path) targetArch := arch.FromFilename(path)
// JIT execution runs when the host CPU matches the kernel's // JIT execution runs when the host CPU matches the kernel's
// architecture, except loong64: its trampoline is implemented but not // architecture; every trampoline is validated end to end under
// yet validated against real hardware (the Go runtime cannot start // qemu-user emulation (the loong64 one included, via the raw-address
// under the available loong64 emulators), so those kernels take the // leave handoff).
// toolchain-comparison path. if targetArch != hostArch() {
if targetArch != hostArch() || targetArch == arch.LOONG64 { // No JIT on this host: ground truth and profile remain available for
// every architecture, because cmdVerifyNonJIT assembles and compares
// against the toolchain without executing anything.
switch targetArch { switch targetArch {
case arch.RISCV: case arch.AMD64, arch.RISCV, arch.LOONG64, arch.ARM64:
// RISC-V: ground-truth only (no JIT on non-RISC-V hosts). return cmdVerifyNonJIT(path, targetArch, *groundTruth, *profile)
return cmdVerifyRISCV(path, *groundTruth, *profile)
case arch.LOONG64:
// LoongArch: ground-truth only (trampoline not yet
// hardware-validated).
return cmdVerifyLOONG64(path, *groundTruth, *profile)
case arch.ARM64:
// AArch64: ground-truth only (no JIT on non-ARM64 hosts).
return cmdVerifyARM64(path, *groundTruth, *profile)
case arch.AMD64:
fmt.Fprintln(os.Stderr, "gasm verify: JIT-based checks need an amd64 host; use --ground-truth here")
return 1
default: default:
fmt.Fprintln(os.Stderr, "gasm verify: unsupported architecture") fmt.Fprintln(os.Stderr, "gasm verify: unsupported architecture")
return 1 return 1
@@ -1224,50 +1182,13 @@ each entry reproduces.
fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err) fmt.Fprintf(os.Stderr, "gasm verify: ground truth: %v\n", err)
return 1 return 1
} }
matched, total := 0, 0 matched, total, diffs := compareGroundTruth(k.Image(), gt)
for _, name := range names { if diffs > 0 {
fl, _ := k.Func(name) printCodeDiff(k.Image(), gt)
gasmCode := k.Image().Code[fl.Offset : fl.Offset+fl.Size] rc = 1
goCode, ok := gt[name]
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", name)
continue
}
total++
// Compare, masking relocation sites (disp32 fields that the
// Go linker fills at link time — gasm resolves them internally).
gasmCmp := make([]byte, len(gasmCode))
goCmp := make([]byte, len(goCode))
copy(gasmCmp, gasmCode)
copy(goCmp, goCode)
for _, r := range fl.Relocs {
for j := r.Off; j < r.Off+4 && j < len(gasmCmp); j++ {
gasmCmp[j] = 0
}
for j := r.Off; j < r.Off+4 && j < len(goCmp); j++ {
goCmp[j] = 0
}
}
if bytes.Equal(gasmCmp, goCmp) {
matched++
if len(fl.Relocs) > 0 {
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", name, fl.Size, len(fl.Relocs))
} else {
fmt.Printf(" %s: MATCH (%d bytes)\n", name, fl.Size)
}
} else {
fmt.Printf(" %s: MISMATCH (gasm %d bytes, go %d bytes)\n", name, fl.Size, len(goCode))
for i := 0; i < len(gasmCmp) && i < len(goCmp); i++ {
if gasmCmp[i] != goCmp[i] {
fmt.Printf(" first diff at byte %d: gasm=%02x go=%02x\n", i, gasmCmp[i], goCmp[i])
break
}
}
rc = 1
}
} }
fmt.Printf("ground truth: %d/%d functions byte-identical\n", matched, total) fmt.Printf("ground truth: %d/%d functions byte-identical\n", matched, total)
if matched < total { if matched < total || diffs > 0 {
rc = 1 rc = 1
} }
} }
@@ -1289,13 +1210,11 @@ each entry reproduces.
sigs := verify.ExtractSignatures(src) sigs := verify.ExtractSignatures(src)
fuzzed := 0 fuzzed := 0
for _, name := range names { for _, name := range names {
sig, ok := sigs[name] if _, ok := sigs[name]; !ok {
if !ok {
fmt.Printf(" %s: SKIP (no // func signature)\n", name) fmt.Printf(" %s: SKIP (no // func signature)\n", name)
continue continue
} }
goCode, ok := gt[name] if _, ok := gt[name]; !ok {
if !ok {
fmt.Printf(" %s: SKIP (not in go tool asm output)\n", name) fmt.Printf(" %s: SKIP (not in go tool asm output)\n", name)
continue continue
} }
@@ -1312,8 +1231,6 @@ each entry reproduces.
rc = 1 rc = 1
} }
} }
_ = sig
_ = goCode
fuzzed++ fuzzed++
} }
fmt.Printf("fuzz: %d functions tested, %d iterations each\n", fuzzed, *fuzzN) fmt.Printf("fuzz: %d functions tested, %d iterations each\n", fuzzed, *fuzzN)
@@ -1338,7 +1255,7 @@ each entry reproduces.
// Run smoke and ABI checks in parallel, each function in its own child // Run smoke and ABI checks in parallel, each function in its own child
// process: the JIT'd code runs with zeroed or fuzzed arguments, and a // process: the JIT'd code runs with zeroed or fuzzed arguments, and a
// function that dereferences them faults — the crash is reported as a // function that dereferences them faults, the crash is reported as a
// CRASH line instead of killing this process (mirrors fuzzInSubprocess). // CRASH line instead of killing this process (mirrors fuzzInSubprocess).
if *smoke || *abi { if *smoke || *abi {
type checkResult struct { type checkResult struct {
@@ -1479,7 +1396,7 @@ func fuzzInSubprocess(path, funcName string, n int, extra ...string) string {
if exitErr, ok := err.(*exec.ExitError); ok { if exitErr, ok := err.(*exec.ExitError); ok {
ws := exitErr.Sys().(syscall.WaitStatus) ws := exitErr.Sys().(syscall.WaitStatus)
if ws.Signaled() { if ws.Signaled() {
return fmt.Sprintf("%s: CRASH (%v — partial function, use --ground-truth)", funcName, ws.Signal()) return fmt.Sprintf("%s: CRASH (%v; partial function, use --ground-truth)", funcName, ws.Signal())
} }
} }
// Non-zero exit without a signal: the fuzz reported mismatches. // Non-zero exit without a signal: the fuzz reported mismatches.
@@ -1504,12 +1421,12 @@ func fuzzInSubprocess(path, funcName string, n int, extra ...string) string {
// sweepInSubprocess runs the smoke/abi checks for a single function in a // sweepInSubprocess runs the smoke/abi checks for a single function in a
// child process. If the child is killed by a signal (e.g. SIGSEGV from a // child process. If the child is killed by a signal (e.g. SIGSEGV from a
// function that dereferences its zeroed or fuzzed arguments), it returns a // function that dereferences its zeroed or fuzzed arguments), it returns a
// CRASH report instead of dying — the same isolation fuzzInSubprocess // CRASH report instead of dying, the same isolation fuzzInSubprocess
// provides for the fuzz sweep. // provides for the fuzz sweep.
func sweepInSubprocess(path, funcName string, smoke, abi bool, abiN int) (string, bool) { func sweepInSubprocess(path, funcName string, smoke, abi bool, abiN int) (string, bool) {
self, err := os.Executable() self, err := os.Executable()
if err != nil { if err != nil {
return fmt.Sprintf(" smoke/abi: FAIL — cannot find self: %v", err), true return fmt.Sprintf(" smoke/abi: FAIL: cannot find self: %v", err), true
} }
args := []string{"verify"} args := []string{"verify"}
if smoke { if smoke {
@@ -1526,7 +1443,7 @@ func sweepInSubprocess(path, funcName string, smoke, abi bool, abiN int) (string
if exitErr, ok := err.(*exec.ExitError); ok { if exitErr, ok := err.(*exec.ExitError); ok {
ws, ok := exitErr.Sys().(syscall.WaitStatus) ws, ok := exitErr.Sys().(syscall.WaitStatus)
if ok && ws.Signaled() { if ok && ws.Signaled() {
return fmt.Sprintf(" smoke/abi: CRASH (%v — the function faults on zeroed or fuzzed\n arguments; verify it with -call and valid buffers)", ws.Signal()), true return fmt.Sprintf(" smoke/abi: CRASH (%v: the function faults on zeroed or fuzzed\n arguments; verify it with -call and valid buffers)", ws.Signal()), true
} }
} }
// Non-zero exit without a signal: the checks themselves failed and // Non-zero exit without a signal: the checks themselves failed and
@@ -1550,7 +1467,7 @@ func sweepCheckLines(out []byte) string {
} }
// runSweepChecks performs the in-process smoke and ABI checks for one // runSweepChecks performs the in-process smoke and ABI checks for one
// function — the child half of sweepInSubprocess. // function, the child half of sweepInSubprocess.
func runSweepChecks(k *verify.Kernel, path, name string, fl asm.FuncLayout, smoke, abi bool, abiN int) ([]string, bool) { func runSweepChecks(k *verify.Kernel, path, name string, fl asm.FuncLayout, smoke, abi bool, abiN int) ([]string, bool) {
var msgs []string var msgs []string
failed := false failed := false
@@ -1559,7 +1476,7 @@ func runSweepChecks(k *verify.Kernel, path, name string, fl asm.FuncLayout, smok
args := make([]byte, fl.Args) args := make([]byte, fl.Args)
_, err := k.CallFunc(name, args) _, err := k.CallFunc(name, args)
if err != nil { if err != nil {
msgs = append(msgs, fmt.Sprintf(" smoke: FAIL — %v", err)) msgs = append(msgs, fmt.Sprintf(" smoke: FAIL: %v", err))
failed = true failed = true
} else { } else {
msgs = append(msgs, " smoke: OK") msgs = append(msgs, " smoke: OK")
@@ -1579,7 +1496,7 @@ func runSweepChecks(k *verify.Kernel, path, name string, fl asm.FuncLayout, smok
args := make([]byte, fl.Args) args := make([]byte, fl.Args)
_, report, err := k.CallFuncChecked(name, args) _, report, err := k.CallFuncChecked(name, args)
if err != nil { if err != nil {
msgs = append(msgs, fmt.Sprintf(" abi: FAIL — %v", err)) msgs = append(msgs, fmt.Sprintf(" abi: FAIL: %v", err))
failed = true failed = true
} else if !report.OK() { } else if !report.OK() {
msgs = append(msgs, fmt.Sprintf(" abi: %s", report)) msgs = append(msgs, fmt.Sprintf(" abi: %s", report))
@@ -1669,7 +1586,7 @@ func cmdVerifyCall(k *verify.Kernel, path, funcName, bufSpec, scalarSpec string,
for i := range repeat { for i := range repeat {
out, err := k.CallFunc(funcName, args) out, err := k.CallFunc(funcName, args)
if err != nil { if err != nil {
fmt.Printf(" call %d: FAIL — %v\n", i+1, err) fmt.Printf(" call %d: FAIL: %v\n", i+1, err)
rc = 1 rc = 1
continue continue
} }
+171 -3
View File
@@ -13,6 +13,9 @@ import (
"strings" "strings"
"syscall" "syscall"
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
) )
const clean = "#include \"textflag.h\"\n" + const clean = "#include \"textflag.h\"\n" +
@@ -219,8 +222,9 @@ func TestCmdVersion(t *testing.T) {
if code != 0 { if code != 0 {
t.Fatalf("code = %d", code) t.Fatalf("code = %d", code)
} }
if !strings.Contains(out, version) { got := version()
t.Errorf("version output %q does not mention %q", out, version) if !strings.Contains(out, got) {
t.Errorf("version output %q does not mention %q", out, got)
} }
} }
@@ -241,6 +245,96 @@ func TestCmdArgErrors(t *testing.T) {
} }
} }
// TestUsageExitCodes pins the exit-code contract for the commands whose main
// dispatches on a returned error: a wrong argument set exits 2, the same as
// the commands that count their arguments themselves, while a runtime
// failure (an unreadable file) keeps exit 1.
func TestUsageExitCodes(t *testing.T) {
for name, err := range map[string]error{
"audit-instructions extra argument": cmdAuditInstructions([]string{"amd64", "extra"}),
"audit-instructions unknown arch": cmdAuditInstructions([]string{"mips"}),
"audit-instructions corpus extra": cmdAuditInstructions([]string{"--corpus", "a", "b"}),
"scaffold no arguments": cmdScaffold(nil),
"scaffold extra arguments": cmdScaffold([]string{"differential", "a.s", "b.s"}),
} {
if err == nil {
t.Errorf("%s: expected an error", name)
continue
}
if code := exitCodeFor(err); code != 2 {
t.Errorf("%s: exit code = %d, want 2 (err: %v)", name, code, err)
}
}
if err := cmdScaffold([]string{"differential", "/nonexistent/file.s"}); err == nil {
t.Error("scaffold on a missing file should fail")
} else if code := exitCodeFor(err); code != 1 {
t.Errorf("scaffold on a missing file: exit code = %d, want 1", code)
}
}
// TestCmdAsmFormatValidation checks that an unknown --format exits 2 with
// and without -o, instead of assembling and silently dumping a raw image.
func TestCmdAsmFormatValidation(t *testing.T) {
path := writeTemp(t, "f_amd64.s", clean)
out := filepath.Join(t.TempDir(), "f.bin")
if _, _, code := capture(func() int { return cmdAsm([]string{"--format", "bogus", path}) }); code != 2 {
t.Errorf("asm --format bogus without -o: code = %d, want 2", code)
}
if _, _, code := capture(func() int { return cmdAsm([]string{"--format", "bogus", "-o", out, path}) }); code != 2 {
t.Errorf("asm --format bogus with -o: code = %d, want 2", code)
}
}
// TestCmdAsmOutputFile pins the documented -o behaviour: the output goes to
// the file and stdout carries no hex dump; without -o the dump is the output.
func TestCmdAsmOutputFile(t *testing.T) {
path := writeTemp(t, "f_amd64.s", clean)
out := filepath.Join(t.TempDir(), "f.bin")
stdout, _, code := capture(func() int { return cmdAsm([]string{"-o", out, path}) })
if code != 0 {
t.Fatalf("code = %d", code)
}
if strings.Contains(stdout, "0000:") {
t.Errorf("stdout carries a hex dump despite -o:\n%s", stdout)
}
if !strings.Contains(stdout, "wrote ") {
t.Errorf("stdout misses the wrote line:\n%s", stdout)
}
b, err := os.ReadFile(out)
if err != nil {
t.Fatal(err)
}
if len(b) == 0 {
t.Error("the output file is empty")
}
stdout, _, code = capture(func() int { return cmdAsm([]string{path}) })
if code != 0 {
t.Fatalf("without -o: code = %d", code)
}
if !strings.Contains(stdout, "0000:") {
t.Errorf("without -o the hex dump is missing:\n%s", stdout)
}
}
// TestVerifyNonJITAMD64GroundTruth drives the cross-architecture
// ground-truth path for an amd64 kernel: the path a host of any other
// architecture takes, which must compare against the toolchain rather than
// refuse to run.
func TestVerifyNonJITAMD64GroundTruth(t *testing.T) {
if testing.Short() {
t.Skip("runs go tool asm")
}
path := writeTemp(t, "f_amd64.s", clean)
out, _, code := capture(func() int { return cmdVerifyNonJIT(path, arch.AMD64, true, false) })
if code != 0 {
t.Fatalf("code = %d (%s)", code, out)
}
if !strings.Contains(out, "1/1 matched") {
t.Errorf("output misses the matched report:\n%s", out)
}
}
// TestVerifySmokeCrashIsolation checks that a function faulting on its // TestVerifySmokeCrashIsolation checks that a function faulting on its
// zeroed smoke arguments is reported as CRASH by a child process instead of // zeroed smoke arguments is reported as CRASH by a child process instead of
// killing `gasm verify` itself. // killing `gasm verify` itself.
@@ -274,7 +368,7 @@ func TestVerifySmokeCrashIsolation(t *testing.T) {
} }
if exitErr, ok := err.(*exec.ExitError); ok { if exitErr, ok := err.(*exec.ExitError); ok {
if ws, ok := exitErr.Sys().(syscall.WaitStatus); ok && ws.Signaled() { if ws, ok := exitErr.Sys().(syscall.WaitStatus); ok && ws.Signaled() {
t.Fatalf("verify died from %v — the crash was not isolated:\n%s", ws.Signal(), out) t.Fatalf("verify died from %v; the crash was not isolated:\n%s", ws.Signal(), out)
} }
} }
if !strings.Contains(string(out), "CRASH") { if !strings.Contains(string(out), "CRASH") {
@@ -292,3 +386,77 @@ func TestSweepCheckLines(t *testing.T) {
t.Errorf("sweepCheckLines = %q, want %q", got, want) t.Errorf("sweepCheckLines = %q, want %q", got, want)
} }
} }
// TestRunCorpusAudit drives the corpus audit over a small fixture tree: one
// suffixed amd64 file, one suffixed arm64 file whose body is not arm64, one
// generic file, and one file that does not parse.
func TestRunCorpusAudit(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("good_amd64.s", "#include \"textflag.h\"\nTEXT ·add(SB), NOSPLIT, $0-0\n\tMOVQ AX, BX\n\tRET\n")
write("bad_arm64.s", "#include \"textflag.h\"\nTEXT ·f(SB), NOSPLIT, $0-0\n\tMOVQ AX, BX\n\tRET\n")
write("generic.s", "#include \"textflag.h\"\nTEXT ·g(SB), NOSPLIT, $0-0\n\tRET\n")
write("broken.s", "#include \"textflag.h\"\nTEXT ·b(SB), NOSPLIT, $0-0\n\tJMP nowhere\n\tRET\n")
stats, err := runCorpusAudit(dir)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
if stats.files != 4 {
t.Errorf("files = %d, want 4", stats.files)
}
if stats.generic != 2 {
t.Errorf("generic = %d, want 2 (generic.s and broken.s)", stats.generic)
}
// good_amd64 and generic.s assemble everywhere they are attempted.
if stats.full != 2 {
t.Errorf("full = %d, want 2", stats.full)
}
get := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
// amd64: good_amd64 + generic.s + broken.s; the broken file fails to parse.
if a := get("amd64"); a.attempted != 3 || a.assembled != 2 {
t.Errorf("amd64 = %d/%d, want 2/3", a.assembled, a.attempted)
}
// arm64: bad_arm64 (MOVQ is not arm64) + generic.s + broken.s.
if a := get("arm64"); a.attempted != 3 || a.assembled != 1 {
t.Errorf("arm64 = %d/%d, want 1/3", a.assembled, a.attempted)
}
if r := get("amd64").reasons["instruction not encodable"]; r != 0 {
t.Errorf("amd64 unexpected unencodable reason: %d", r)
}
if r := get("arm64").reasons["instruction not encodable"]; r != 1 {
t.Errorf("arm64 unencodable reasons = %d, want 1", r)
}
}
// TestCompareGroundTruthPadding pins the padding-aware ground-truth
// comparison: the toolchain pads text symbols to 16-byte boundaries, so
// trailing zeros in the reference must not read as a mismatch, while any
// non-zero tail still must.
func TestCompareGroundTruthPadding(t *testing.T) {
code := []byte{0x48, 0x8b, 0x07, 0xc3} // 4 bytes, not a multiple of 16
img := &asm.Image{Code: code, Funcs: []asm.FuncLayout{{Name: "f", Offset: 0, Size: len(code)}}}
padded := append(append([]byte(nil), code...), 0, 0, 0)
matched, total, diffs := compareGroundTruth(img, map[string][]byte{"f": padded})
if matched != 1 || total != 1 || diffs != 0 {
t.Fatalf("zero padding should match: matched=%d total=%d diffs=%d", matched, total, diffs)
}
dirty := append(append([]byte(nil), code...), 0, 0x90, 0)
matched, _, diffs = compareGroundTruth(img, map[string][]byte{"f": dirty})
if matched != 0 || diffs != 1 {
t.Fatalf("non-zero padding must mismatch: matched=%d diffs=%d", matched, diffs)
}
}
+241
View File
@@ -0,0 +1,241 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"os/exec"
"path/filepath"
"regexp"
"strings"
"testing"
)
// TestManPagesTrackTheCLI builds the binary once, then compares every
// command's live `-h` output with its docs/man/gasm-<command>.1 page: the
// flag sets must agree both ways, and the page's SYNOPSIS line must carry
// the command's usage line. A flag or a usage change that skips the man
// page fails here, so the pages cannot drift from the binary.
func TestManPagesTrackTheCLI(t *testing.T) {
if testing.Short() {
t.Skip("builds the gasm binary")
}
bin := filepath.Join(t.TempDir(), "gasm")
if out, err := exec.Command("go", "build", "-o", bin, ".").CombinedOutput(); err != nil {
t.Fatalf("build gasm: %v\n%s", err, out)
}
for _, cmd := range []string{
"tokens", "parse", "fmt", "lint", "asm", "dis", "verify",
"debug", "diff", "profile", "audit-instructions", "scaffold", "lsp",
} {
t.Run(cmd, func(t *testing.T) {
raw, err := os.ReadFile(filepath.Join("..", "..", "docs", "man", "gasm-"+cmd+".1"))
if err != nil {
t.Fatalf("read man page: %v", err)
}
page := string(raw)
out, _ := exec.Command(bin, cmd, "-h").CombinedOutput()
help := string(out)
binFlags := helpFlags(help)
pageFlags := roffFlags(page)
for f := range binFlags {
if !pageFlags[f] {
t.Errorf("flag -%s is in the binary's help but missing from the man page", f)
}
}
for f := range pageFlags {
if !binFlags[f] {
t.Errorf("flag -%s is in the man page but the binary does not accept it", f)
}
}
want := helpUsage(help)
got := roffSynopsis(page)
if want != "" && got != want {
t.Errorf("SYNOPSIS drift:\n page: %s\nbinary: %s", got, want)
}
})
}
}
// TestManCommandsTrackHelp compares the gasm(1) COMMANDS list with the
// top-level help output, so a subcommand added to the binary cannot miss
// its man entry and a stale entry cannot outlive its command.
func TestManCommandsTrackHelp(t *testing.T) {
if testing.Short() {
t.Skip("builds the gasm binary")
}
bin := filepath.Join(t.TempDir(), "gasm")
if out, err := exec.Command("go", "build", "-o", bin, ".").CombinedOutput(); err != nil {
t.Fatalf("build gasm: %v\n%s", err, out)
}
raw, err := os.ReadFile(filepath.Join("..", "..", "docs", "man", "gasm.1"))
if err != nil {
t.Fatalf("read man page: %v", err)
}
helpOut, err := exec.Command(bin, "--help").Output()
if err != nil {
t.Fatalf("gasm --help: %v", err)
}
binCmds := helpCommands(string(helpOut))
pageCmds := roffCommands(string(raw))
for c := range binCmds {
if !pageCmds[c] {
t.Errorf("command %q is in the binary's help but missing from gasm(1) COMMANDS", c)
}
}
for c := range pageCmds {
if !binCmds[c] {
t.Errorf("command %q is in gasm(1) COMMANDS but the binary does not list it", c)
}
}
}
// helpFlags extracts the flag names from a `gasm <cmd> -h` output.
func helpFlags(help string) map[string]bool {
m := map[string]bool{}
inFlags := false
for line := range strings.SplitSeq(help, "\n") {
if strings.TrimRight(line, " \t") == "Flags:" {
inFlags = true
continue
}
if !inFlags {
continue
}
if !strings.HasPrefix(line, " -") {
continue
}
token := strings.FieldsFunc(strings.TrimLeft(line, " "), func(r rune) bool {
return r == ' ' || r == '\t'
})
if len(token) == 0 {
continue
}
m[strings.TrimLeft(token[0], "-")] = true
}
return m
}
var roffEscape = regexp.MustCompile(`\\f[BIRP]`)
// roffFlags extracts the flag names from a man page's OPTIONS section.
func roffFlags(page string) map[string]bool {
m := map[string]bool{}
inOptions := false
for line := range strings.SplitSeq(page, "\n") {
if strings.HasPrefix(line, ".SH ") {
inOptions = strings.HasPrefix(line, ".SH OPTIONS")
continue
}
if !inOptions {
continue
}
// Flag entries are written as either `.B \-flag` or `\fB\-flag`.
var body string
switch {
case strings.HasPrefix(line, `.B \-`):
body = line[3:]
case strings.HasPrefix(line, `\fB\-`):
body = line[1:]
default:
continue
}
name := roffEscape.ReplaceAllString(body, "")
name = strings.ReplaceAll(name, `\-`, "-")
name = strings.TrimSpace(name)
if i := strings.IndexAny(name, " \t"); i >= 0 {
name = name[:i]
}
m[strings.TrimLeft(name, "-")] = true
}
return m
}
// helpCommands extracts the command names from the top-level help output's
// Commands section.
func helpCommands(help string) map[string]bool {
m := map[string]bool{}
inCmds := false
for line := range strings.SplitSeq(help, "\n") {
if strings.TrimSpace(line) == "Commands:" {
inCmds = true
continue
}
if !inCmds {
continue
}
t := strings.TrimSpace(line)
if t == "" {
break
}
name, _, _ := strings.Cut(t, " ")
m[name] = true
}
return m
}
// roffCommands extracts the command names from gasm(1)'s COMMANDS section,
// where each entry is written as `.B gasm\-<name>(1)` or `.B gasm <name>`.
func roffCommands(page string) map[string]bool {
m := map[string]bool{}
inCmds := false
for line := range strings.SplitSeq(page, "\n") {
if strings.HasPrefix(line, ".SH ") {
inCmds = strings.HasPrefix(line, ".SH COMMANDS")
continue
}
if !inCmds || !strings.HasPrefix(line, ".B gasm") {
continue
}
entry := strings.ReplaceAll(strings.TrimPrefix(line, ".B "), `\-`, "-")
entry = strings.TrimSuffix(entry, "(1)")
switch {
case strings.HasPrefix(entry, "gasm-"):
m[strings.TrimPrefix(entry, "gasm-")] = true
case strings.HasPrefix(entry, "gasm "):
m[strings.TrimPrefix(entry, "gasm ")] = true
}
}
return m
}
// helpUsage returns the command's usage line without the "Usage: " prefix.
func helpUsage(help string) string {
for line := range strings.SplitSeq(help, "\n") {
if strings.HasPrefix(line, "Usage: ") {
return normaliseUsage(line[len("Usage: "):])
}
}
return ""
}
// roffSynopsis returns the page's SYNOPSIS usage line, unescaped.
func roffSynopsis(page string) string {
inSyn := false
for line := range strings.SplitSeq(page, "\n") {
if strings.HasPrefix(line, ".SH ") {
inSyn = strings.HasPrefix(line, ".SH SYNOPSIS")
continue
}
if !inSyn || !strings.HasPrefix(line, ".B ") {
continue
}
return normaliseUsage(strings.ReplaceAll(line[3:], `\-`, "-"))
}
return ""
}
// normaliseUsage flattens whitespace and drops the roff font escapes so that
// the binary's usage line and the page's SYNOPSIS line compare equal.
func normaliseUsage(s string) string {
s = roffEscape.ReplaceAllString(s, "")
return strings.Join(strings.Fields(s), " ")
}
+2 -2
View File
@@ -19,7 +19,7 @@ import (
// file: a Go test that seeds random states, drives both the assembly kernel // file: a Go test that seeds random states, drives both the assembly kernel
// and a caller-provided portable reference, and compares the outputs // and a caller-provided portable reference, and compares the outputs
// byte-for-byte. The lesson this encodes: a pipeline-level fuzz cannot see // byte-for-byte. The lesson this encodes: a pipeline-level fuzz cannot see
// an unwired kernel — only a direct-call differential against the portable // an unwired kernel, only a direct-call differential against the portable
// specification can, so every kernel ships with one. // specification can, so every kernel ships with one.
// //
// The generated file follows two conventions the caller fills in: // The generated file follows two conventions the caller fills in:
@@ -44,7 +44,7 @@ bodies, place the file in the kernel's package, and run it in CI.
rest = rest[1:] rest = rest[1:]
} }
if len(rest) != 1 { if len(rest) != 1 {
return fmt.Errorf("usage: gasm scaffold differential <file.s>") return &usageError{fmt.Errorf("usage: gasm scaffold differential <file.s>")}
} }
path := rest[0] path := rest[0]
src, err := os.ReadFile(path) src, err := os.ReadFile(path)
+140
View File
@@ -0,0 +1,140 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"fmt"
"slices"
"strings"
)
// unifiedDiff renders a unified diff with three lines of context between the
// two line slices, in the form `gofmt -d` prints. An empty result means the
// inputs are identical.
func unifiedDiff(name string, a, b []string) string {
if slices.Equal(a, b) {
return ""
}
var out strings.Builder
fmt.Fprintf(&out, "--- %s\n+++ %s\n", name, name)
// Longest common subsequence over the lines (assembly files are small
// enough for the quadratic table).
n, m := len(a), len(b)
lcs := make([][]int, n+1)
for i := range lcs {
lcs[i] = make([]int, m+1)
}
for i := n - 1; i >= 0; i-- {
for j := m - 1; j >= 0; j-- {
if a[i] == b[j] {
lcs[i][j] = lcs[i+1][j+1] + 1
} else if lcs[i+1][j] >= lcs[i][j+1] {
lcs[i][j] = lcs[i+1][j]
} else {
lcs[i][j] = lcs[i][j+1]
}
}
}
// Walk the LCS once, assigning every op its absolute position in both
// files (1-based, the position an insertion sits before).
type op struct {
kind byte // ' ', '-' or '+'
aLine, bLine int
text string
}
var ops []op
aPos, bPos := 0, 0
emit := func(kind byte, text string) {
ops = append(ops, op{kind: kind, aLine: aPos + 1, bLine: bPos + 1, text: text})
switch kind {
case ' ':
aPos++
bPos++
case '-':
aPos++
case '+':
bPos++
}
}
i, j := 0, 0
for i < n && j < m {
switch {
case a[i] == b[j]:
emit(' ', a[i])
i++
j++
case lcs[i+1][j] >= lcs[i][j+1]:
emit('-', a[i])
i++
default:
emit('+', b[j])
j++
}
}
for ; i < n; i++ {
emit('-', a[i])
}
for ; j < m; j++ {
emit('+', b[j])
}
// Group the edits into hunks: consecutive changes separated by more than
// twice the context lines start a new hunk.
const context = 3
var changes []int
for k, o := range ops {
if o.kind != ' ' {
changes = append(changes, k)
}
}
for g := 0; g < len(changes); {
last := g
for last+1 < len(changes) && changes[last+1]-changes[last]-1 <= 2*context {
last++
}
lo := max(0, changes[g]-context)
hi := min(len(ops), changes[last]+1+context)
// The header numbers are the first line of each side actually shown:
// the first context, deletion or insertion line. A hunk that shows
// no old lines is a pure insertion and reports the position it sits
// before (0 at the top of the file); the mirror rule holds for a
// pure deletion.
aStart := ops[lo].aLine - 1
bStart := ops[lo].bLine - 1
countA, countB := 0, 0
for _, o := range ops[lo:hi] {
switch o.kind {
case ' ':
countA++
countB++
case '-':
countA++
case '+':
countB++
}
}
for _, o := range ops[lo:hi] {
if o.kind != '+' {
aStart = o.aLine
break
}
}
for _, o := range ops[lo:hi] {
if o.kind != '-' {
bStart = o.bLine
break
}
}
fmt.Fprintf(&out, "@@ -%d,%d +%d,%d @@\n", aStart, countA, bStart, countB)
for _, o := range ops[lo:hi] {
out.WriteByte(o.kind)
out.WriteString(o.text)
out.WriteByte('\n')
}
g = last + 1
}
return out.String()
}
+94
View File
@@ -0,0 +1,94 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"slices"
"strings"
"testing"
)
func lines(ss ...string) []string { return ss }
func TestUnifiedDiffIdentical(t *testing.T) {
if got := unifiedDiff("f", lines("a", "b"), lines("a", "b")); got != "" {
t.Errorf("identical inputs produced %q, want empty", got)
}
}
func TestUnifiedDiffSingleChange(t *testing.T) {
a := lines("1", "2", "3", "4", "5", "6", "7", "8")
b := lines("1", "2", "3!", "4", "5", "6", "7", "8")
want := "--- f\n+++ f\n" +
"@@ -1,6 +1,6 @@\n" +
" 1\n 2\n-3\n+3!\n 4\n 5\n 6\n"
if got := unifiedDiff("f", a, b); got != want {
t.Errorf("diff = %q, want %q", got, want)
}
}
func TestUnifiedDiffInsertAtStart(t *testing.T) {
got := unifiedDiff("f", lines("x"), lines("new", "x"))
// The single existing line is shown as trailing context, so the hunk
// covers it.
want := "--- f\n+++ f\n@@ -1,1 +1,2 @@\n+new\n x\n"
if got != want {
t.Errorf("diff = %q, want %q", got, want)
}
}
func TestUnifiedDiffDeleteAtEnd(t *testing.T) {
got := unifiedDiff("f", lines("x", "y"), lines("x"))
want := "--- f\n+++ f\n@@ -1,2 +1,1 @@\n x\n-y\n"
if got != want {
t.Errorf("diff = %q, want %q", got, want)
}
}
func TestUnifiedDiffTwoHunks(t *testing.T) {
var a, b []string
for i := 1; i <= 20; i++ {
a = append(a, itoa(i))
b = append(b, itoa(i))
}
b[1] = "2!"
b[17] = "18!"
got := unifiedDiff("f", a, b)
if !strings.Contains(got, "@@ -1,5 +1,5 @@\n 1\n-2\n+2!\n 3\n 4\n 5\n") {
t.Errorf("first hunk wrong:\n%s", got)
}
if !strings.Contains(got, "@@ -15,6 +15,6 @@\n 15\n 16\n 17\n-18\n+18!\n 19\n 20\n") {
t.Errorf("second hunk wrong:\n%s", got)
}
}
// TestUnifiedDiffAdjacentHunks merges changes separated by exactly twice the
// context into one hunk.
func TestUnifiedDiffAdjacentHunks(t *testing.T) {
a := lines("1", "2", "3", "4", "5", "6", "7", "8")
b := slices.Clone(a)
b[0] = "1!"
b[7] = "8!"
got := unifiedDiff("f", a, b)
want := "--- f\n+++ f\n" +
"@@ -1,8 +1,8 @@\n" +
"-1\n+1!\n 2\n 3\n 4\n 5\n 6\n 7\n-8\n+8!\n"
if got != want {
t.Errorf("diff = %q, want %q", got, want)
}
}
func itoa(n int) string {
if n == 0 {
return "0"
}
var buf [4]byte
i := len(buf)
for n > 0 {
i--
buf[i] = byte('0' + n%10)
n /= 10
}
return string(buf[i:])
}
+101 -47
View File
@@ -9,11 +9,11 @@ import "strings"
import "fmt" import "fmt"
// Breakpoint is one INT3 breakpoint in the debuggee. // Breakpoint is one software breakpoint in the debuggee.
type Breakpoint struct { type Breakpoint struct {
Addr uint64 // absolute address in the debuggee Addr uint64 // absolute address in the debuggee
Label string // source label ("" for raw addresses) Label string // source label ("" for raw addresses)
Orig byte // original byte at Addr (restored on removal) Orig []byte // original bytes at Addr (restored on removal)
Enabled bool Enabled bool
Cond *Condition // optional condition (nil = unconditional) Cond *Condition // optional condition (nil = unconditional)
hits int hits int
@@ -32,11 +32,15 @@ type Condition struct {
MemAddr uint64 // memory address (for register-memory comparison, prefixed with *) MemAddr uint64 // memory address (for register-memory comparison, prefixed with *)
} }
// Eval checks the condition against the current registers. // Eval checks the condition against the current registers. For the
func (c *Condition) Eval(regs *Regs) bool { // register-memory form, mem reads an 8-byte little-endian word from the
// debuggee; it may be nil when no reader is available. Anything that cannot
// be decided (unknown register or operator, unreadable memory) does not
// block the breakpoint.
func (c *Condition) Eval(regs *Regs, mem func(addr uint64) (uint64, bool)) bool {
actual, ok := regs.RegValue(c.Reg) actual, ok := regs.RegValue(c.Reg)
if !ok { if !ok {
return true // unknown register — don't block return true // unknown register, don't block
} }
var expected uint64 var expected uint64
switch { switch {
@@ -48,9 +52,16 @@ func (c *Condition) Eval(regs *Regs) bool {
} }
expected = v expected = v
case c.MemAddr != 0: case c.MemAddr != 0:
// Register-memory comparison — requires a Session, not available here. // Register-memory comparison, resolved in the debuggee at
// Fall back to treating as constant (the caller should resolve). // evaluation time.
expected = c.Value if mem == nil {
return true
}
v, ok := mem(c.MemAddr)
if !ok {
return true
}
expected = v
default: default:
expected = c.Value expected = c.Value
} }
@@ -72,8 +83,19 @@ func (c *Condition) Eval(regs *Regs) bool {
} }
} }
// Breakpoints manages the set of breakpoints for a Session. // String renders the condition for display.
// Breakpoints manages software breakpoints for a debuggee. func (c *Condition) String() string {
switch {
case c.Reg2 != "":
return fmt.Sprintf("%s %s %s", c.Reg, c.Op, c.Reg2)
case c.MemAddr != 0:
return fmt.Sprintf("%s %s *%#x", c.Reg, c.Op, c.MemAddr)
default:
return fmt.Sprintf("%s %s %#x", c.Reg, c.Op, c.Value)
}
}
// Breakpoints manages the software breakpoints of one Session.
type Breakpoints struct { type Breakpoints struct {
t tracer t tracer
bps map[uint64]*Breakpoint bps map[uint64]*Breakpoint
@@ -84,6 +106,18 @@ func NewBreakpoints(t tracer) *Breakpoints {
return &Breakpoints{t: t, bps: make(map[uint64]*Breakpoint)} return &Breakpoints{t: t, bps: make(map[uint64]*Breakpoint)}
} }
// breakpointMask is the byte mask of the breakpoint instruction inside a
// peeked word: the low len(breakpointInsn) bytes, because every supported
// architecture is little-endian and patches the instruction at the lowest
// address of the word.
func breakpointMask() uint64 {
var mask uint64
for range breakpointInsn {
mask = (mask << 8) | 0xFF
}
return mask
}
// Set installs a breakpoint at addr (replaces any existing one). // Set installs a breakpoint at addr (replaces any existing one).
func (bm *Breakpoints) Set(addr uint64, label string) (*Breakpoint, error) { func (bm *Breakpoints) Set(addr uint64, label string) (*Breakpoint, error) {
return bm.SetWithCond(addr, label, nil) return bm.SetWithCond(addr, label, nil)
@@ -101,13 +135,12 @@ func (bm *Breakpoints) SetWithCond(addr uint64, label string, cond *Condition) (
if err != nil { if err != nil {
return nil, err return nil, err
} }
orig := byte(word) orig := make([]byte, len(breakpointInsn))
// Patch with the breakpoint instruction, preserving the rest of the word. for i := range orig {
mask := uint64(0) orig[i] = byte(word >> (8 * i))
for range breakpointInsn {
mask = (mask << 8) | 0xFF
} }
patched := (word &^ mask) | breakpointWord(breakpointInsn) // Patch with the breakpoint instruction, preserving the rest of the word.
patched := (word &^ breakpointMask()) | breakpointWord(breakpointInsn)
if err := bm.t.Poke(addr, patched); err != nil { if err := bm.t.Poke(addr, patched); err != nil {
return nil, err return nil, err
} }
@@ -135,26 +168,40 @@ func (bm *Breakpoints) Info() string {
} }
cond := "" cond := ""
if bp.Cond != nil { if bp.Cond != nil {
cond = fmt.Sprintf(" if %s %s %#x", bp.Cond.Reg, bp.Cond.Op, bp.Cond.Value) cond = " if " + bp.Cond.String()
} }
result.WriteString(fmt.Sprintf(" %d: %s at %#x [%s, %d hits]%s\n", i, label, bp.Addr, status, bp.hits, cond)) result.WriteString(fmt.Sprintf(" %d: %s at %#x [%s, %d hits]%s\n", i, label, bp.Addr, status, bp.hits, cond))
} }
return result.String() return result.String()
} }
// Clear removes the breakpoint at addr, restoring the original byte. // restore writes the saved original bytes back over the breakpoint
// instruction, preserving the rest of the peeked word. It reports whether
// both the peek and the poke succeeded.
func (bm *Breakpoints) restore(addr uint64, bp *Breakpoint) bool {
word, err := bm.t.Peek(addr)
if err != nil {
return false
}
orig := uint64(0)
for i, b := range bp.Orig {
orig |= uint64(b) << (8 * i)
}
return bm.t.Poke(addr, (word&^breakpointMask())|orig) == nil
}
// Clear removes the breakpoint at addr, restoring the original bytes.
func (bm *Breakpoints) Clear(addr uint64) error { func (bm *Breakpoints) Clear(addr uint64) error {
bp, ok := bm.bps[addr] bp, ok := bm.bps[addr]
if !ok { if !ok {
return fmt.Errorf("debug: no breakpoint at %#x", addr) return fmt.Errorf("debug: no breakpoint at %#x", addr)
} }
word, err := bm.t.Peek(addr) if !bm.restore(addr, bp) {
if err != nil { word, err := bm.t.Peek(addr)
return err if err != nil {
} return err
restored := (word &^ 0xFF) | uint64(bp.Orig) }
if err := bm.t.Poke(addr, restored); err != nil { return fmt.Errorf("debug: restore breakpoint at %#x failed, word is %#x", addr, word)
return err
} }
delete(bm.bps, addr) delete(bm.bps, addr)
return nil return nil
@@ -186,43 +233,54 @@ func (bm *Breakpoints) All() []*Breakpoint {
// HandleTrap is called after the debuggee stops on SIGTRAP. It checks // HandleTrap is called after the debuggee stops on SIGTRAP. It checks
// whether the trap was caused by one of our breakpoints (PC-adjust matches // whether the trap was caused by one of our breakpoints (PC-adjust matches
// a breakpoint address), restores the original byte, rewinds PC, and // a breakpoint address), restores the original bytes, rewinds PC, and
// returns the breakpoint that was hit (or nil if it was a single-step). // returns the breakpoint that was hit (or nil if it was a single-step).
// Hits returns how many times the breakpoint has been hit. // Hits returns how many times the breakpoint has been hit.
func (bp *Breakpoint) Hits() int { return bp.hits } func (bp *Breakpoint) Hits() int { return bp.hits }
func (bm *Breakpoints) HandleTrap(regs *Regs) *Breakpoint { func (bm *Breakpoints) HandleTrap(regs *Regs) *Breakpoint {
// After a breakpoint trap, PC points past the breakpoint instruction. // On amd64 the kernel reports the trap with RIP past the INT3; on the
// other supported architectures the PC still stands on the trap
// instruction, which breakpointPCAdjust encodes per architecture.
trapAddr := regs.GetPC() - uint64(breakpointPCAdjust) trapAddr := regs.GetPC() - uint64(breakpointPCAdjust)
bp, ok := bm.bps[trapAddr] bp, ok := bm.bps[trapAddr]
if !ok || !bp.Enabled { if !ok || !bp.Enabled {
return nil // single-step trap or unknown return nil // single-step trap or unknown
} }
// Check the condition (if any). // Check the condition (if any).
if bp.Cond != nil && !bp.Cond.Eval(regs) { if bp.Cond != nil && !bp.Cond.Eval(regs, bm.peekValue) {
// Condition not met — restore the byte but do NOT rewind RIP. // Condition not met: step the original instruction and re-arm the
// The process continues from the next instruction (past the INT3). // breakpoint, leaving the debuggee stopped just past it, ready to
word, err := bm.t.Peek(trapAddr) // resume silently. The PC must be rewound first: on architectures
if err == nil { // that report the trap past the instruction (amd64) it would
restored := (word &^ 0xFF) | uint64(bp.Orig) // otherwise sit on the second byte of the replaced instruction.
bm.t.Poke(trapAddr, restored) if !bm.restore(trapAddr, bp) {
return nil
} }
// RIP is already past the INT3 (trapAddr + 1). Don't rewind. regs.SetPC(trapAddr)
if err := bm.t.SetRegs(regs); err != nil {
return nil
}
if err := bm.t.Step(); err != nil {
return nil
}
bm.Reinsert(trapAddr)
return nil return nil
} }
bp.hits++ bp.hits++
// Restore the original byte. // Restore the original bytes and rewind PC to re-execute them.
word, err := bm.t.Peek(trapAddr) bm.restore(trapAddr, bp)
if err == nil {
restored := (word &^ 0xFF) | uint64(bp.Orig)
bm.t.Poke(trapAddr, restored)
}
// Rewind PC to re-execute the original instruction.
regs.SetPC(trapAddr) regs.SetPC(trapAddr)
bm.t.SetRegs(regs) bm.t.SetRegs(regs)
return bp return bp
} }
// peekValue adapts tracer.Peek to the Condition value reader.
func (bm *Breakpoints) peekValue(addr uint64) (uint64, bool) {
v, err := bm.t.Peek(addr)
return v, err == nil
}
// Reinsert re-inserts the breakpoint at addr after a single-step past it. // Reinsert re-inserts the breakpoint at addr after a single-step past it.
// Called after Step() when we want the breakpoint to fire again on the // Called after Step() when we want the breakpoint to fire again on the
// next Continue(). // next Continue().
@@ -235,11 +293,7 @@ func (bm *Breakpoints) Reinsert(addr uint64) error {
if err != nil { if err != nil {
return err return err
} }
mask := uint64(0) patched := (word &^ breakpointMask()) | breakpointWord(breakpointInsn)
for range breakpointInsn {
mask = (mask << 8) | 0xFF
}
patched := (word &^ mask) | breakpointWord(breakpointInsn)
return bm.t.Poke(addr, patched) return bm.t.Poke(addr, patched)
} }
+265
View File
@@ -0,0 +1,265 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux
package debug
// Architecture-neutral tests: label and line tables, and the breakpoint
// manager against the mock tracer. These do not launch a debuggee, so they
// build on every supported linux architecture.
import (
"strings"
"testing"
)
func TestLineAt(t *testing.T) {
lines := []SourceLine{
{Offset: 0, Line: 5},
{Offset: 5, Line: 6},
{Offset: 10, Line: 7},
{Offset: 15, Line: 8},
}
tests := []struct {
offset int
want int
}{
{0, 5},
{1, 5},
{4, 5},
{5, 6},
{7, 6},
{10, 7},
{12, 7},
{15, 8},
{20, 8},
}
for _, tt := range tests {
got := lineAt(lines, tt.offset)
if got != tt.want {
t.Errorf("lineAt(lines, %d) = %d, want %d", tt.offset, got, tt.want)
}
}
// Empty table.
if lineAt(nil, 5) != 0 {
t.Error("lineAt(nil, 5) should return 0")
}
}
func TestOffsetForLine(t *testing.T) {
lines := []SourceLine{
{Offset: 0, Line: 5},
{Offset: 5, Line: 6},
{Offset: 10, Line: 7},
}
tests := []struct {
line int
want int
}{
{5, 0},
{6, 5},
{7, 10},
{99, -1}, // not found
{0, -1}, // not found
}
for _, tt := range tests {
got := offsetForLine(lines, tt.line)
if got != tt.want {
t.Errorf("offsetForLine(lines, %d) = %d, want %d", tt.line, got, tt.want)
}
}
}
func TestNearestLabel(t *testing.T) {
labels := []Label{
{Name: "start", Offset: 0},
{Name: "loop", Offset: 10},
{Name: "done", Offset: 20},
}
tests := []struct {
offset int
want string
}{
{0, "start"},
{5, "start"},
{10, "loop"},
{15, "loop"},
{20, "done"},
{25, "done"},
}
for _, tt := range tests {
got := nearestLabel(labels, tt.offset)
if got != tt.want {
t.Errorf("nearestLabel(labels, %d) = %q, want %q", tt.offset, got, tt.want)
}
}
}
func TestBreakpointsSetAndClear(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
// Set a breakpoint at address 0x1000.
bp, err := bm.Set(0x1000, "test")
if err != nil {
t.Fatalf("Set: %v", err)
}
if !bp.Enabled {
t.Error("breakpoint not enabled")
}
if bp.Label != "test" {
t.Errorf("label = %q, want test", bp.Label)
}
// Verify Peek was called.
if len(tr.peeks) != 1 || tr.peeks[0] != 0x1000 {
t.Errorf("peeks = %v, want [0x1000]", tr.peeks)
}
// Verify Poke wrote the breakpoint instruction's bytes.
if len(tr.pokes) != 1 || tr.pokes[0].addr != 0x1000 {
t.Errorf("pokes = %v", tr.pokes)
}
if got := tr.pokes[0].val & breakpointMask(); got != breakpointWord(breakpointInsn) {
t.Errorf("patched bytes %#x, want %#x", got, breakpointWord(breakpointInsn))
}
// At should find it.
if bm.At(0x1000) == nil {
t.Error("At(0x1000) returned nil")
}
// All should return it.
all := bm.All()
if len(all) != 1 {
t.Errorf("All() = %d breakpoints, want 1", len(all))
}
// Clear it.
if err := bm.Clear(0x1000); err != nil {
t.Fatalf("Clear: %v", err)
}
if bm.At(0x1000) != nil {
t.Error("At(0x1000) after Clear should be nil")
}
}
// TestBreakpointRestoreWidth proves the restore path writes back every
// byte of the breakpoint instruction's width, not just the first byte: on
// arm64, riscv64 and loong64 the instruction is four bytes, and restoring
// one byte would leave three bytes of the trap instruction in place.
func TestBreakpointRestoreWidth(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
tr.mem[0x3000] = 0x11
tr.mem[0x3001] = 0x22
tr.mem[0x3002] = 0x33
tr.mem[0x3003] = 0x44
if _, err := bm.Set(0x3000, "width"); err != nil {
t.Fatalf("Set: %v", err)
}
for i, b := range breakpointInsn {
if tr.mem[0x3000+uint64(i)] != b {
t.Fatalf("byte %d after Set = %#x, want the breakpoint byte %#x", i, tr.mem[0x3000+uint64(i)], b)
}
}
if len(bm.At(0x3000).Orig) != len(breakpointInsn) {
t.Fatalf("Orig holds %d bytes, want %d", len(bm.At(0x3000).Orig), len(breakpointInsn))
}
if err := bm.Clear(0x3000); err != nil {
t.Fatalf("Clear: %v", err)
}
want := []byte{0x11, 0x22, 0x33, 0x44}
for i, b := range want {
if tr.mem[0x3000+uint64(i)] != b {
t.Errorf("byte %d after Clear = %#x, want %#x (restore must cover the full instruction width)", i, tr.mem[0x3000+uint64(i)], b)
}
}
}
func TestBreakpointsSetWithCond(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
cond := &Condition{Reg: "rax", Op: "==", Value: 42}
bp, err := bm.SetWithCond(0x2000, "cond_test", cond)
if err != nil {
t.Fatalf("SetWithCond: %v", err)
}
if bp.Cond == nil || bp.Cond.Value != 42 {
t.Error("condition not set")
}
// Re-setting the same address should update the condition.
cond2 := &Condition{Reg: "rbx", Op: "<", Value: 100}
bp2, err := bm.SetWithCond(0x2000, "cond_test2", cond2)
if err != nil {
t.Fatalf("SetWithCond (update): %v", err)
}
if bp2.Cond.Value != 100 {
t.Error("condition not updated")
}
// Should have only 1 Peek (first Set), second is update (no Peek needed).
if len(tr.peeks) != 1 {
t.Errorf("expected 1 Peek, got %d", len(tr.peeks))
}
}
func TestBreakpointsClearAll(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
bm.Set(0x1000, "a")
bm.Set(0x2000, "b")
bm.Set(0x3000, "c")
if len(bm.All()) != 3 {
t.Fatalf("expected 3 breakpoints, got %d", len(bm.All()))
}
bm.ClearAll()
if len(bm.All()) != 0 {
t.Errorf("ClearAll: expected 0 breakpoints, got %d", len(bm.All()))
}
}
func TestBreakpointInfo(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
bm.Set(0x4000, "info_test")
info := bm.Info()
if info == "" {
t.Error("Info returned empty string")
}
if !strings.Contains(info, "info_test") {
t.Errorf("Info %q does not contain label", info)
}
}
// TestConditionString covers the display of all three condition forms.
func TestConditionString(t *testing.T) {
tests := []struct {
cond Condition
want string
}{
{Condition{Reg: "rax", Op: "==", Value: 42}, "rax == 0x2a"},
{Condition{Reg: "rax", Op: "!=", Reg2: "rbx"}, "rax != rbx"},
{Condition{Reg: "rax", Op: "<", MemAddr: 0x5000}, "rax < *0x5000"},
}
for _, tt := range tests {
if got := tt.cond.String(); got != tt.want {
t.Errorf("Condition.String() = %q, want %q", got, tt.want)
}
}
}
+42 -193
View File
@@ -6,7 +6,6 @@
package debug package debug
import ( import (
"strings"
"testing" "testing"
) )
@@ -48,7 +47,7 @@ func TestConditionEval(t *testing.T) {
} }
for _, tt := range tests { for _, tt := range tests {
got := tt.cond.Eval(regs) got := tt.cond.Eval(regs, nil)
if got != tt.want { if got != tt.want {
t.Errorf("Condition{%q %q %d}.Eval() = %v, want %v", t.Errorf("Condition{%q %q %d}.Eval() = %v, want %v",
tt.cond.Reg, tt.cond.Op, tt.cond.Value, got, tt.want) tt.cond.Reg, tt.cond.Op, tt.cond.Value, got, tt.want)
@@ -56,65 +55,33 @@ func TestConditionEval(t *testing.T) {
} }
} }
func TestLineAt(t *testing.T) { // TestConditionEvalMem covers the register-memory form: the value is read
lines := []SourceLine{ // through the supplied reader, and a missing or failing reader must not
{Offset: 0, Line: 5}, // block the breakpoint.
{Offset: 5, Line: 6}, func TestConditionEvalMem(t *testing.T) {
{Offset: 10, Line: 7}, regs := &Regs{RAX: 7}
{Offset: 15, Line: 8}, mem := func(addr uint64) (uint64, bool) {
} if addr == 0x5000 {
return 7, true
tests := []struct {
offset int
want int
}{
{0, 5},
{1, 5},
{4, 5},
{5, 6},
{7, 6},
{10, 7},
{12, 7},
{15, 8},
{20, 8},
}
for _, tt := range tests {
got := lineAt(lines, tt.offset)
if got != tt.want {
t.Errorf("lineAt(lines, %d) = %d, want %d", tt.offset, got, tt.want)
} }
return 0, false
} }
// Empty table. eq := Condition{Reg: "rax", Op: "==", MemAddr: 0x5000}
if lineAt(nil, 5) != 0 { if !eq.Eval(regs, mem) {
t.Error("lineAt(nil, 5) should return 0") t.Error("register-memory comparison with matching word should hold")
} }
} ne := Condition{Reg: "rax", Op: "!=", MemAddr: 0x5000}
if ne.Eval(regs, mem) {
func TestOffsetForLine(t *testing.T) { t.Error("register-memory comparison with mismatching word should not hold")
lines := []SourceLine{
{Offset: 0, Line: 5},
{Offset: 5, Line: 6},
{Offset: 10, Line: 7},
} }
bad := Condition{Reg: "rax", Op: "==", MemAddr: 0x6000}
tests := []struct { if !bad.Eval(regs, mem) {
line int t.Error("unreadable memory must not block the breakpoint")
want int
}{
{5, 0},
{6, 5},
{7, 10},
{99, -1}, // not found
{0, -1}, // not found
} }
noReader := Condition{Reg: "rax", Op: "==", MemAddr: 0x5000}
for _, tt := range tests { if !noReader.Eval(regs, nil) {
got := offsetForLine(lines, tt.line) t.Error("missing memory reader must not block the breakpoint")
if got != tt.want {
t.Errorf("offsetForLine(lines, %d) = %d, want %d", tt.line, got, tt.want)
}
} }
} }
@@ -141,142 +108,8 @@ func TestDecodeRflags(t *testing.T) {
} }
} }
func TestNearestLabel(t *testing.T) {
labels := []Label{
{Name: "start", Offset: 0},
{Name: "loop", Offset: 10},
{Name: "done", Offset: 20},
}
tests := []struct {
offset int
want string
}{
{0, "start"},
{5, "start"},
{10, "loop"},
{15, "loop"},
{20, "done"},
{25, "done"},
}
for _, tt := range tests {
got := nearestLabel(labels, tt.offset)
if got != tt.want {
t.Errorf("nearestLabel(labels, %d) = %q, want %q", tt.offset, got, tt.want)
}
}
}
func TestBreakpointsSetAndClear(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
// Set a breakpoint at address 0x1000.
bp, err := bm.Set(0x1000, "test")
if err != nil {
t.Fatalf("Set: %v", err)
}
if !bp.Enabled {
t.Error("breakpoint not enabled")
}
if bp.Label != "test" {
t.Errorf("label = %q, want test", bp.Label)
}
// Verify Peek was called.
if len(tr.peeks) != 1 || tr.peeks[0] != 0x1000 {
t.Errorf("peeks = %v, want [0x1000]", tr.peeks)
}
// Verify Poke wrote INT3.
if len(tr.pokes) != 1 || tr.pokes[0].addr != 0x1000 {
t.Errorf("pokes = %v", tr.pokes)
}
// At should find it.
if bm.At(0x1000) == nil {
t.Error("At(0x1000) returned nil")
}
// All should return it.
all := bm.All()
if len(all) != 1 {
t.Errorf("All() = %d breakpoints, want 1", len(all))
}
// Clear it.
if err := bm.Clear(0x1000); err != nil {
t.Fatalf("Clear: %v", err)
}
if bm.At(0x1000) != nil {
t.Error("At(0x1000) after Clear should be nil")
}
}
func TestBreakpointsSetWithCond(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
cond := &Condition{Reg: "rax", Op: "==", Value: 42}
bp, err := bm.SetWithCond(0x2000, "cond_test", cond)
if err != nil {
t.Fatalf("SetWithCond: %v", err)
}
if bp.Cond == nil || bp.Cond.Value != 42 {
t.Error("condition not set")
}
// Re-setting the same address should update the condition.
cond2 := &Condition{Reg: "rbx", Op: "<", Value: 100}
bp2, err := bm.SetWithCond(0x2000, "cond_test2", cond2)
if err != nil {
t.Fatalf("SetWithCond (update): %v", err)
}
if bp2.Cond.Value != 100 {
t.Error("condition not updated")
}
// Should have only 1 Peek (first Set), second is update (no Peek needed).
if len(tr.peeks) != 1 {
t.Errorf("expected 1 Peek, got %d", len(tr.peeks))
}
}
func TestBreakpointsClearAll(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
bm.Set(0x1000, "a")
bm.Set(0x2000, "b")
bm.Set(0x3000, "c")
if len(bm.All()) != 3 {
t.Fatalf("expected 3 breakpoints, got %d", len(bm.All()))
}
bm.ClearAll()
if len(bm.All()) != 0 {
t.Errorf("ClearAll: expected 0 breakpoints, got %d", len(bm.All()))
}
}
func TestBreakpointInfo(t *testing.T) {
tr := newMockTracer()
bm := NewBreakpoints(tr)
bm.Set(0x4000, "info_test")
info := bm.Info()
if info == "" {
t.Error("Info returned empty string")
}
if !strings.Contains(info, "info_test") {
t.Errorf("Info %q does not contain label", info)
}
}
func TestWatchpointSlotTracking(t *testing.T) { func TestWatchpointSlotTracking(t *testing.T) {
wpSlots = [4]bool{} // reset s := &Session{} // per-session slots start free
s := &Session{}
// All four slots are free initially. // All four slots are free initially.
for i := range 4 { for i := range 4 {
@@ -289,8 +122,8 @@ func TestWatchpointSlotTracking(t *testing.T) {
} }
// Manually mark slots 0 and 2 as used (simulating successful SetWatchpoint). // Manually mark slots 0 and 2 as used (simulating successful SetWatchpoint).
wpSlots[0] = true s.wpSlots[0] = true
wpSlots[2] = true s.wpSlots[2] = true
if !s.IsWatchpointSlotUsed(0) { if !s.IsWatchpointSlotUsed(0) {
t.Error("slot 0 should be in use") t.Error("slot 0 should be in use")
@@ -318,9 +151,25 @@ func TestWatchpointSlotTracking(t *testing.T) {
// Mark all slots used: FindFreeWatchpointSlot returns -1. // Mark all slots used: FindFreeWatchpointSlot returns -1.
for i := range 4 { for i := range 4 {
wpSlots[i] = true s.wpSlots[i] = true
} }
if got := s.FindFreeWatchpointSlot(); got != -1 { if got := s.FindFreeWatchpointSlot(); got != -1 {
t.Errorf("FindFreeWatchpointSlot() with all slots used = %d, want -1", got) t.Errorf("FindFreeWatchpointSlot() with all slots used = %d, want -1", got)
} }
} }
// TestUnwatchSlotBound checks the bound the REPL parses against: it must
// cover the architecture's whole slot range, not a hardcoded 0-3.
func TestUnwatchSlotBound(t *testing.T) {
max := maxWatchpoints()
if max < 4 {
t.Fatalf("maxWatchpoints() = %d, want at least 4", max)
}
s := &Session{}
if s.IsWatchpointSlotUsed(max - 1) {
t.Errorf("slot %d should be free initially", max-1)
}
if s.IsWatchpointSlotUsed(max) {
t.Errorf("slot %d must be out of range", max)
}
}
+14 -11
View File
@@ -9,27 +9,22 @@ import (
"fmt" "fmt"
"strings" "strings"
"golang.org/x/arch/x86/x86asm" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
) )
// Disassemble decodes the instruction at the given address in the debuggee's // Disassemble decodes the instruction at the given address in the debuggee's
// memory and returns its text representation and length in bytes. // memory and returns its text representation and length in bytes.
func (s *Session) Disassemble(addr uint64) (string, int, error) { func (s *Session) Disassemble(addr uint64) (string, int, error) {
// Read up to 15 bytes (max x86 instruction length).
mem, err := s.ReadMemory(addr, 15) mem, err := s.ReadMemory(addr, 15)
if err != nil { if err != nil {
// Try a shorter read if we're near a page boundary. return "", 0, err
mem, err = s.ReadMemory(addr, 1)
if err != nil {
return "", 0, err
}
} }
inst, err := x86asm.Decode(mem, 64) ins, err := disasm.Decode(arch.AMD64, mem, addr)
if err != nil { if err != nil {
return "???", 1, nil return "", 0, err
} }
text := x86asm.IntelSyntax(inst, addr, nil) return ins.Text, ins.Len, nil
return text, inst.Len, nil
} }
// DisassembleN decodes up to n instructions starting at addr and returns // DisassembleN decodes up to n instructions starting at addr and returns
@@ -51,3 +46,11 @@ func (s *Session) DisassembleN(addr uint64, n int) string {
} }
return result.String() return result.String()
} }
// isCallInsn reports whether disassembled text (x86asm.IntelSyntax) is a
// call. The first token must match exactly: a prefix test would also catch
// unrelated mnemonics.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
return strings.ToLower(m) == "call"
}
+18 -5
View File
@@ -7,8 +7,10 @@ package debug
import ( import (
"fmt" "fmt"
"strings"
"golang.org/x/arch/arm64/arm64asm" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
) )
// Disassemble decodes the instruction at the given address in the debuggee's // Disassemble decodes the instruction at the given address in the debuggee's
@@ -18,12 +20,11 @@ func (s *Session) Disassemble(addr uint64) (string, int, error) {
if err != nil { if err != nil {
return "", 0, err return "", 0, err
} }
inst, err := arm64asm.Decode(mem) ins, err := disasm.Decode(arch.ARM64, mem, addr)
if err != nil { if err != nil {
return "???", 4, nil return "", 0, err
} }
text := arm64asm.GoSyntax(inst, addr, nil, nil) return ins.Text, ins.Len, nil
return text, 4, nil
} }
// DisassembleN decodes up to n instructions starting at addr. // DisassembleN decodes up to n instructions starting at addr.
@@ -44,3 +45,15 @@ func (s *Session) DisassembleN(addr uint64, n int) string {
} }
return result return result
} }
// isCallInsn reports whether disassembled text (arm64asm.GoSyntax) is a
// call. GoSyntax renders bl as CALL; the native mnemonic is accepted too.
// The first token must match exactly so branches never match.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
switch strings.ToLower(m) {
case "call", "bl":
return true
}
return false
}
+19 -5
View File
@@ -7,8 +7,10 @@ package debug
import ( import (
"fmt" "fmt"
"strings"
"golang.org/x/arch/loong64/loong64asm" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
) )
// Disassemble decodes the instruction at the given address in the debuggee's // Disassemble decodes the instruction at the given address in the debuggee's
@@ -18,12 +20,11 @@ func (s *Session) Disassemble(addr uint64) (string, int, error) {
if err != nil { if err != nil {
return "", 0, err return "", 0, err
} }
inst, err := loong64asm.Decode(mem) ins, err := disasm.Decode(arch.LOONG64, mem, addr)
if err != nil { if err != nil {
return "???", 4, nil return "", 0, err
} }
text := loong64asm.GoSyntax(inst, addr, nil) return ins.Text, ins.Len, nil
return text, 4, nil
} }
// DisassembleN decodes up to n instructions starting at addr. // DisassembleN decodes up to n instructions starting at addr.
@@ -44,3 +45,16 @@ func (s *Session) DisassembleN(addr uint64, n int) string {
} }
return result return result
} }
// isCallInsn reports whether disassembled text (loong64asm.GoSyntax) is a
// call. GoSyntax renders bl and jirl calls as CALL (jirl returns print
// RET); the native mnemonics are accepted too. The first token must match
// exactly: a "bl" prefix would catch bltz and other branches.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
switch strings.ToLower(m) {
case "call", "bl", "jirl":
return true
}
return false
}
+20 -5
View File
@@ -7,8 +7,10 @@ package debug
import ( import (
"fmt" "fmt"
"strings"
"golang.org/x/arch/riscv64/riscv64asm" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
) )
// Disassemble decodes the instruction at the given address in the debuggee's // Disassemble decodes the instruction at the given address in the debuggee's
@@ -18,12 +20,11 @@ func (s *Session) Disassemble(addr uint64) (string, int, error) {
if err != nil { if err != nil {
return "", 0, err return "", 0, err
} }
inst, err := riscv64asm.Decode(mem) ins, err := disasm.Decode(arch.RISCV, mem, addr)
if err != nil { if err != nil {
return "???", 4, nil return "", 0, err
} }
text := riscv64asm.GoSyntax(inst, addr, nil, nil) return ins.Text, ins.Len, nil
return text, inst.Len, nil
} }
// DisassembleN decodes up to n instructions starting at addr. // DisassembleN decodes up to n instructions starting at addr.
@@ -44,3 +45,17 @@ func (s *Session) DisassembleN(addr uint64, n int) string {
} }
return result return result
} }
// isCallInsn reports whether disassembled text (riscv64asm.GoSyntax) is a
// call. GoSyntax renders jal and jalr calls as CALL; the native mnemonics
// are accepted too. The first token must match exactly: a prefix test on
// "bl" would catch branches on other architectures, and jalr as ret prints
// RET, which must not be stepped over.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
switch strings.ToLower(m) {
case "call", "jal", "jalr":
return true
}
return false
}
+22 -2
View File
@@ -73,9 +73,29 @@ func decodeRflags(f uint64) string {
return flags[:len(flags)-1] return flags[:len(flags)-1]
} }
// archReturnAddr reads the return address from the stack (amd64 ABI0 convention). // archReturnAddr reads the return address of the current frame (amd64
// ABI0 convention). A function that contains a CALL (or has a frame) is
// assembled with the prologue PUSHQ BP; MOVQ SP, BP, so mid-function the
// word at SP is the saved caller BP, a stack address, and the return
// address sits further up. Walk the stack from SP and take the first word
// that lies in an executable mapping: stack and data words never do, a
// return address always does.
func archReturnAddr(s *Session, regs *Regs) (uint64, error) { func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
return s.Peek(regs.GetSP()) ranges := execRanges(s.pid)
for off := uint64(0); off < 512; off += 8 {
word, err := s.Peek(regs.RSP + off)
if err != nil {
break
}
for _, r := range ranges {
if word >= r.lo && word < r.hi {
return word, nil
}
}
}
// No mapping available or nothing code-like on the stack: fall back to
// the raw entry convention, [SP] before any push.
return s.Peek(regs.RSP)
} }
// archSPLabel returns the SP register name for display. // archSPLabel returns the SP register name for display.
+6 -3
View File
@@ -5,7 +5,10 @@
package debug package debug
import "fmt" import (
"encoding/binary"
"fmt"
)
func printRegs(regs *Regs, codeBase, funcOff uint64) { func printRegs(regs *Regs, codeBase, funcOff uint64) {
fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff) fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff)
@@ -31,8 +34,8 @@ func printRegs(regs *Regs, codeBase, funcOff uint64) {
func printVectorRegs(v *VectorRegs) { func printVectorRegs(v *VectorRegs) {
fmt.Println("\n Vector registers (V0-V31):") fmt.Println("\n Vector registers (V0-V31):")
for i := 0; i < 32; i += 2 { for i := 0; i < 32; i += 2 {
fmt.Printf(" V%-2d = %016x%016x\n", i, v.V[i][8], v.V[i][0]) fmt.Printf(" V%-2d = %016x%016x\n", i, binary.LittleEndian.Uint64(v.V[i][8:16]), binary.LittleEndian.Uint64(v.V[i][0:8]))
fmt.Printf(" V%-2d = %016x%016x\n", i+1, v.V[i+1][8], v.V[i+1][0]) fmt.Printf(" V%-2d = %016x%016x\n", i+1, binary.LittleEndian.Uint64(v.V[i+1][8:16]), binary.LittleEndian.Uint64(v.V[i+1][0:8]))
} }
} }
+427
View File
@@ -0,0 +1,427 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux && amd64
package debug
import (
"bytes"
"fmt"
"io"
"os"
"path/filepath"
"runtime"
"strings"
"testing"
"time"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
)
// Integration tests beyond the basic entry breakpoint: hardware watchpoints,
// conditional breakpoints, next/finish over a CALL, faulting kernels and the
// xstate vector-register readout. All drive a real ptrace session, so they
// run on amd64 hosts only.
// writeKernel writes an assembly source to a temporary file with the
// architecture suffix the assembler dispatcher expects.
func writeKernel(t *testing.T, src string) string {
t.Helper()
path := filepath.Join(t.TempDir(), "kernel_amd64.s")
if err := os.WriteFile(path, []byte(src), 0o644); err != nil {
t.Fatalf("write kernel: %v", err)
}
return path
}
// launchKernel launches a session for the kernel source and returns the
// session, its breakpoint manager and the function layout.
func launchKernel(t *testing.T, bin, path, funcName string, args []byte) (*Session, *Breakpoints, asm.FuncLayout) {
t.Helper()
k, err := verify.Load(path)
if err != nil {
t.Fatalf("Load: %v", err)
}
t.Cleanup(k.Close)
fl, err := k.Func(funcName)
if err != nil {
t.Fatalf("Func: %v", err)
}
if len(args) < fl.Args {
padded := make([]byte, fl.Args)
copy(padded, args)
args = padded
}
sess, err := Launch(bin, path, funcName, args)
if err != nil {
t.Fatalf("Launch: %v", err)
}
t.Cleanup(sess.Kill)
bm := NewBreakpoints(sess)
return sess, bm, fl
}
// runToEntry resumes the freshly launched debuggee until the breakpoint at
// the function entry traps, mirroring the REPL continue loop: the debuggee
// SIGSTOPs twice (launch barrier and entry barrier) before entering the JIT
// call.
func runToEntry(t *testing.T, sess *Session, bm *Breakpoints, entry uint64) {
t.Helper()
for range 50 {
for _, bp := range bm.All() {
bm.Reinsert(bp.Addr)
}
if err := sess.Continue(); err != nil {
t.Fatalf("Continue: %v", err)
}
if sess.Exited() {
t.Fatal("debuggee exited before the entry breakpoint trapped")
}
regs, err := sess.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
if bm.HandleTrap(&regs) != nil {
return
}
}
t.Fatal("no entry breakpoint trap after 50 resumes")
}
// captureStdout runs fn with os.Stdout redirected to a pipe and returns
// what it printed (the REPL writes its reports to stdout).
func captureStdout(t *testing.T, fn func()) string {
t.Helper()
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
old := os.Stdout
os.Stdout = w
done := make(chan string, 1)
go func() {
b, _ := io.ReadAll(r)
done <- string(b)
}()
defer func() { os.Stdout = old }()
fn()
w.Close()
return <-done
}
// TestWatchpointArmRunHit proves the debug-register offsets: the watchpoint
// must fire on the store, with si_addr naming the watched address. The
// kernel writes its return value to ret+0(FP), which is the 8-byte word
// right above the stack pointer at entry.
func TestWatchpointArmRunHit(t *testing.T) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := buildGasm(t)
const kernel = `#include "textflag.h"
// func wpret() int64
TEXT ·wpret(SB), NOSPLIT, $0-8
MOVQ $0x5a5a5a5a5a5a5a5a, AX
MOVQ AX, ret+0(FP)
RET
`
path := writeKernel(t, kernel)
sess, bm, fl := launchKernel(t, bin, path, "wpret", nil)
entry := sess.CodeBase() + uint64(fl.Offset)
if _, err := bm.Set(entry, "entry"); err != nil {
t.Fatalf("Set: %v", err)
}
runToEntry(t, sess, bm, entry)
regs, err := sess.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
watched := regs.RSP + 8 // ret+0(FP): the store target
slot := sess.FindFreeWatchpointSlot()
if slot < 0 {
t.Fatal("no free watchpoint slot")
}
if err := sess.SetWatchpoint(slot, watched, WatchWrite, 8); err != nil {
t.Fatalf("SetWatchpoint: %v (wrong debug-register offsets?)", err)
}
if err := sess.Continue(); err != nil {
t.Fatalf("Continue: %v", err)
}
reason, addr := sess.StopInfo()
if reason != StopWatchpoint {
t.Fatalf("stop reason = %v, want StopWatchpoint (DR0-DR3/DR7 offsets are wrong)", reason)
}
if addr != watched {
t.Fatalf("watchpoint address = %#x, want %#x", addr, watched)
}
// The watched word holds the stored value: x86 data breakpoints are
// reported with the access complete.
if word, err := sess.Peek(watched); err != nil || word != 0x5a5a5a5a5a5a5a5a {
t.Errorf("watched word = %#x (err %v), want 0x5a5a5a5a5a5a5a5a", word, err)
}
if err := sess.ClearWatchpoint(slot); err != nil {
t.Fatalf("ClearWatchpoint: %v", err)
}
}
// TestConditionalBreakpointFalseThenTrue proves the false-condition path:
// the breakpoint steps over the original instruction, re-arms itself and
// keeps running silently, and the true condition stops exactly once with the
// register in the expected state.
func TestConditionalBreakpointFalseThenTrue(t *testing.T) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := buildGasm(t)
const kernel = `#include "textflag.h"
// func countdown(n int64) int64
TEXT ·countdown(SB), NOSPLIT, $0-16
MOVQ n+0(FP), CX
loop:
DECQ CX
CMPQ CX, $0
JNE loop
MOVQ CX, ret+8(FP)
RET
`
path := writeKernel(t, kernel)
sess, bm, fl := launchKernel(t, bin, path, "countdown", []byte{8})
loopAddr := sess.CodeBase() + uint64(fl.Offset) + uint64(fl.Labels["loop"])
// The length of the breakpointed instruction, from a disassembly taken
// before the INT3 is patched in.
_, insnLen, err := sess.Disassemble(loopAddr)
if err != nil || insnLen <= 0 {
t.Fatalf("Disassemble at %#x: len=%d err=%v", loopAddr, insnLen, err)
}
cond := &Condition{Reg: "rcx", Op: "==", Value: 1}
bp, err := bm.SetWithCond(loopAddr, "loop", cond)
if err != nil {
t.Fatalf("SetWithCond: %v", err)
}
hits := 0
exited := false
for range 200 {
for _, b := range bm.All() {
bm.Reinsert(b.Addr)
}
if err := sess.Continue(); err != nil {
exited = true
break // the debuggee finished
}
if sess.Exited() {
exited = true
break
}
if sig := sess.LastSignal(); sig != 0 {
t.Fatalf("unexpected signal stop %v", sig)
}
regs, err := sess.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
if hit := bm.HandleTrap(&regs); hit != nil {
hits++
if regs.RCX != 1 {
t.Fatalf("hit with RCX=%d, want 1", regs.RCX)
}
// Park after the instruction, as the REPL does.
if err := sess.Step(); err != nil {
t.Fatalf("Step: %v", err)
}
} else {
// A false evaluation must leave the debuggee past the whole
// original instruction: a PC inside it (trapAddr+1 on amd64)
// means the resume happens mid-instruction.
fresh, err := sess.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
if fresh.RIP > loopAddr && fresh.RIP < loopAddr+uint64(insnLen) {
t.Fatalf("false evaluation left the PC at %#x, inside the %d-byte instruction at %#x",
fresh.RIP, insnLen, loopAddr)
}
}
}
if hits != 1 {
t.Fatalf("conditional breakpoint hit %d times, want exactly 1 (false evaluations must run through silently)", hits)
}
if bp.Hits() != 1 {
t.Errorf("bp.Hits() = %d, want 1", bp.Hits())
}
if !exited || !sess.Exited() {
t.Fatal("debuggee did not run to completion after the conditional hit")
}
}
// TestNextAndFinishOverCall proves next and finish evaluate the trap with
// registers fetched after the stop: next lands exactly on the instruction
// after the CALL, and finish stops exactly on the return address.
func TestNextAndFinishOverCall(t *testing.T) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := buildGasm(t)
const kernel = `#include "textflag.h"
// func caller(x int64) int64
// The argument travels in AX: FP argument slots of CALL-bearing functions
// are an assembler concern outside this test's scope.
TEXT ·caller(SB), NOSPLIT, $0-16
MOVQ $5, AX
CALL ·bump(SB)
aftercall:
MOVQ AX, ret+8(FP)
RET
// func bump(x int64) int64
TEXT ·bump(SB), NOSPLIT, $0-0
ADDQ $3, AX
RET
`
path := writeKernel(t, kernel)
// next: step the prologue and the constant load (3 instructions), then
// step over the CALL and check the landing address and RAX.
sess, bm, fl := launchKernel(t, bin, path, "caller", nil)
entry := sess.CodeBase() + uint64(fl.Offset)
if _, err := bm.Set(entry, "entry"); err != nil {
t.Fatalf("Set: %v", err)
}
runToEntry(t, sess, bm, entry)
afterOff := uint64(fl.Labels["aftercall"])
out := captureStdout(t, func() {
REPL(sess, bm, sess.CodeBase(), fl.Offset, fl.Size, fl.Args, nil, nil,
strings.NewReader("step 3\nnext\nregs\nquit\n"))
})
if !strings.Contains(out, fmt.Sprintf("func+%#x", afterOff)) {
t.Errorf("next did not land on the instruction after the CALL (func+%#x); output:\n%s", afterOff, out)
}
if !strings.Contains(out, "RAX = 0x0000000000000008") {
t.Errorf("callee did not run exactly once under next (want RAX=8); output:\n%s", out)
}
// finish: run to the return address read off the stack at entry.
sess2, bm2, fl2 := launchKernel(t, bin, path, "caller", nil)
entry2 := sess2.CodeBase() + uint64(fl2.Offset)
if _, err := bm2.Set(entry2, "entry"); err != nil {
t.Fatalf("Set: %v", err)
}
runToEntry(t, sess2, bm2, entry2)
regs, err := sess2.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
retAddr, err := sess2.Peek(regs.RSP)
if err != nil {
t.Fatalf("Peek return address: %v", err)
}
out2 := captureStdout(t, func() {
REPL(sess2, bm2, sess2.CodeBase(), fl2.Offset, fl2.Size, fl2.Args, nil, nil,
strings.NewReader("step 1\nfinish\nquit\n"))
})
want := fmt.Sprintf("finished, now at %#x\n", retAddr)
if !strings.Contains(out2, want) {
t.Errorf("finish stopped at the wrong PC; want %q in output:\n%s", want, out2)
}
}
// TestSignalStopSurfaced proves a faulting kernel surfaces as a reported
// stop instead of an infinite fault loop. A regression here hangs, so a
// watchdog fails the run rather than letting CI stall.
func TestSignalStopSurfaced(t *testing.T) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := buildGasm(t)
const kernel = `#include "textflag.h"
// func crash() int64
TEXT ·crash(SB), NOSPLIT, $0-8
XORQ AX, AX
MOVQ (AX), AX
MOVQ AX, ret+0(FP)
RET
`
path := writeKernel(t, kernel)
sess, bm, _ := launchKernel(t, bin, path, "crash", nil)
timer := time.AfterFunc(time.Minute, func() {
panic("watchdog: the debugger hung on the faulting kernel instead of reporting the signal stop")
})
defer timer.Stop()
out := captureStdout(t, func() {
REPL(sess, bm, sess.CodeBase(), 0, 0, 0, nil, nil,
strings.NewReader("continue\nquit\n"))
})
if !strings.Contains(out, "stopped on signal") {
t.Errorf("SIGSEGV did not surface as a reported stop; output:\n%s", out)
}
if !sess.Exited() {
t.Error("debuggee should be killed by quit after the signal stop")
}
}
// TestGetVectorRegsXState proves the NT_X86_XSTATE readout: the request
// succeeds on a normal process and the XMM halves agree with
// PTRACE_GETFPREGS.
func TestGetVectorRegsXState(t *testing.T) {
// The FPRegs layout must mirror the kernel's user_fpregs_struct
// exactly: PTRACE_GETFPREGS fills all 512 bytes, so a short struct
// overflows the caller's memory.
if got := unsafe.Sizeof(FPRegs{}); got != 512 {
t.Fatalf("sizeof(FPRegs) = %d, want 512", got)
}
if got := unsafe.Offsetof(FPRegs{}.XMM); got != 160 {
t.Fatalf("offsetof(FPRegs.XMM) = %d, want 160", got)
}
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := buildGasm(t)
const kernel = `#include "textflag.h"
// func vprobe() int64
TEXT ·vprobe(SB), NOSPLIT, $0-8
MOVQ $1, AX
MOVQ AX, ret+0(FP)
RET
`
path := writeKernel(t, kernel)
sess, bm, fl := launchKernel(t, bin, path, "vprobe", nil)
entry := sess.CodeBase() + uint64(fl.Offset)
if _, err := bm.Set(entry, "entry"); err != nil {
t.Fatalf("Set: %v", err)
}
runToEntry(t, sess, bm, entry)
v, err := sess.GetVectorRegs()
if err != nil {
t.Fatalf("GetVectorRegs: %v", err)
}
fp, err := sess.GetFPRegs()
if err != nil {
t.Fatalf("GetFPRegs: %v", err)
}
for i := range 16 {
if !bytes.Equal(v.YMM[i][:16], fp.XMM[i][:]) {
t.Errorf("YMM%d low half %x, want the FPRegs XMM half %x", i, v.YMM[i][:16], fp.XMM[i][:])
}
}
}
+67 -28
View File
@@ -22,7 +22,14 @@ type Session struct {
cmd *exec.Cmd cmd *exec.Cmd
stopped bool stopped bool
exited bool exited bool
codeBase uint64 // base address of the JIT code in the debuggee codeBase uint64 // base address of the JIT code in the debuggee
tmpDir string // scratch directory of the session, removed on Kill
wpSlots [16]bool // hardware watchpoint slots in use (DR0-DR3, arm64 DBGWVR0-15)
// lastSignal holds the signal of the most recent stop when that stop
// was a genuine signal-delivery-stop the caller must see (a fault such
// as SIGSEGV, SIGBUS, SIGFPE or SIGILL); 0 for breakpoint traps,
// single-steps, SIGSTOP and suppressed runtime signals.
lastSignal syscall.Signal
} }
// Launch starts the debuggee subprocess (gasm debug --target ...) and // Launch starts the debuggee subprocess (gasm debug --target ...) and
@@ -77,7 +84,7 @@ func LaunchWithBuffers(gasmBin, asmPath, funcName string, args []byte, bufSpec s
return nil, nil, fmt.Errorf("debug: start debuggee: %w", err) return nil, nil, fmt.Errorf("debug: start debuggee: %w", err)
} }
s := &Session{pid: cmd.Process.Pid, cmd: cmd} s := &Session{pid: cmd.Process.Pid, cmd: cmd, tmpDir: tmpDir}
readyFile := filepath.Join(tmpDir, "ready") readyFile := filepath.Join(tmpDir, "ready")
for range 500 { for range 500 {
@@ -124,29 +131,17 @@ func LaunchWithBuffers(gasmBin, asmPath, funcName string, args []byte, bufSpec s
return s, bufAddrs, nil return s, bufAddrs, nil
} }
// wait waits for the debuggee to stop and returns the wait status.
func (s *Session) wait() error {
var ws syscall.WaitStatus
_, err := syscall.Wait4(s.pid, &ws, 0, nil)
if err != nil {
return err
}
if ws.Exited() {
s.exited = true
return fmt.Errorf("debuggee exited with status %d", ws.ExitStatus())
}
s.stopped = true
return nil
}
// waitStopped consumes ptrace-stop events until one the debugger cares // waitStopped consumes ptrace-stop events until one the debugger cares
// about arrives: SIGTRAP (a breakpoint or a completed single-step) or the // about arrives: SIGTRAP (a breakpoint or a completed single-step), the
// debuggee's own SIGSTOP. A Go tracee's runtime raises SIGURG for // debuggee's own SIGSTOP, or a genuine signal-delivery-stop. A Go tracee's
// asynchronous preemption, and every signal on a traced thread surfaces as // runtime raises SIGURG for asynchronous preemption, and every signal on a
// a signal-delivery-stop, so those are suppressed and the tracee resumed // traced thread surfaces as a signal-delivery-stop, so SIGURG is suppressed
// without them. Runtime noise is why a single wait can return in the // and the tracee resumed without it. Every other signal (SIGSEGV, SIGBUS,
// middle of runtime code and a resume can then fail: the event stream must // SIGFPE, SIGILL, ...) is returned to the caller: resuming with signal 0
// be drained by the tracer. // would restart the faulting instruction and fault forever, so a faulting
// kernel must surface as a stop the caller reports. Runtime noise is also
// why a single wait can return in the middle of runtime code and a resume
// can then fail: the event stream must be drained by the tracer.
func (s *Session) waitStopped() (syscall.Signal, error) { func (s *Session) waitStopped() (syscall.Signal, error) {
for { for {
var ws syscall.WaitStatus var ws syscall.WaitStatus
@@ -164,10 +159,12 @@ func (s *Session) waitStopped() (syscall.Signal, error) {
switch sig := ws.StopSignal(); sig { switch sig := ws.StopSignal(); sig {
case syscall.SIGTRAP, syscall.SIGSTOP: case syscall.SIGTRAP, syscall.SIGSTOP:
s.stopped = true s.stopped = true
s.lastSignal = 0
return sig, nil return sig, nil
default: case syscall.SIGURG:
// Runtime noise (SIGURG preemption and friends): resume the // Go runtime asynchronous preemption: resume the tracee
// tracee without delivering the signal. // without delivering the signal.
s.lastSignal = 0
if _, _, errno := syscall.Syscall6( if _, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE, syscall.SYS_PTRACE,
uintptr(syscall.PTRACE_CONT), uintptr(syscall.PTRACE_CONT),
@@ -176,10 +173,22 @@ func (s *Session) waitStopped() (syscall.Signal, error) {
); errno != 0 { ); errno != 0 {
return 0, fmt.Errorf("debug: PTRACE_CONT: %w", errno) return 0, fmt.Errorf("debug: PTRACE_CONT: %w", errno)
} }
default:
// A genuine signal-delivery-stop. Report it; the caller
// decides how to proceed.
s.stopped = true
s.lastSignal = sig
return sig, nil
} }
} }
} }
// LastSignal returns the signal of the most recent stop when that stop was
// a genuine signal-delivery-stop (a fault such as SIGSEGV, SIGFPE, SIGILL
// or SIGBUS), and 0 for breakpoint traps, single-steps, SIGSTOP and
// suppressed runtime signals.
func (s *Session) LastSignal() syscall.Signal { return s.lastSignal }
// Peek reads a word (8 bytes) from the debuggee's memory at addr. // Peek reads a word (8 bytes) from the debuggee's memory at addr.
func (s *Session) Peek(addr uint64) (uint64, error) { func (s *Session) Peek(addr uint64) (uint64, error) {
mem, err := os.OpenFile(fmt.Sprintf("/proc/%d/mem", s.pid), os.O_RDONLY, 0) mem, err := os.OpenFile(fmt.Sprintf("/proc/%d/mem", s.pid), os.O_RDONLY, 0)
@@ -293,7 +302,8 @@ func (s *Session) Pid() int { return s.pid }
// CodeBase returns the base address of the JIT code in the debuggee. // CodeBase returns the base address of the JIT code in the debuggee.
func (s *Session) CodeBase() uint64 { return s.codeBase } func (s *Session) CodeBase() uint64 { return s.codeBase }
// Kill terminates the debuggee. // Kill terminates the debuggee and removes the session's scratch
// directory, so a successful session leaves no gasm-debug-* debris behind.
func (s *Session) Kill() { func (s *Session) Kill() {
if !s.exited { if !s.exited {
syscall.Kill(s.pid, syscall.SIGKILL) syscall.Kill(s.pid, syscall.SIGKILL)
@@ -303,6 +313,35 @@ func (s *Session) Kill() {
if s.cmd != nil && s.cmd.Process != nil { if s.cmd != nil && s.cmd.Process != nil {
s.cmd.Wait() s.cmd.Wait()
} }
if s.tmpDir != "" {
os.RemoveAll(s.tmpDir)
s.tmpDir = ""
}
}
// execRange is one executable mapping of the debuggee.
type execRange struct {
lo, hi uint64
}
// execRanges parses the debuggee's executable mappings from /proc/pid/maps.
func execRanges(pid int) []execRange {
data, err := os.ReadFile(fmt.Sprintf("/proc/%d/maps", pid))
if err != nil {
return nil
}
var out []execRange
for line := range strings.SplitSeq(string(data), "\n") {
fields := strings.Fields(line)
if len(fields) < 2 || !strings.Contains(fields[1], "x") {
continue
}
var lo, hi uint64
if _, err := fmt.Sscanf(fields[0], "%x-%x", &lo, &hi); err == nil {
out = append(out, execRange{lo, hi})
}
}
return out
} }
// findRWXMapping reads /proc/pid/maps and returns the base address of the // findRWXMapping reads /proc/pid/maps and returns the base address of the
+68 -11
View File
@@ -6,6 +6,7 @@
package debug package debug
import ( import (
"encoding/binary"
"fmt" "fmt"
"syscall" "syscall"
"unsafe" "unsafe"
@@ -44,20 +45,24 @@ func (s *Session) SetRegs(regs *Regs) error {
return nil return nil
} }
// FPRegs holds the x87 FPU and SSE (XMM) register state from PTRACE_GETFPREGS. // FPRegs holds the x87 FPU and SSE (XMM) register state from
// PTRACE_GETFPREGS. The layout is the kernel's struct user_fpregs_struct
// (sys/user.h), the FXSAVE image: 512 bytes with XMM0-15 at offset 160.
// The i387 fcs/ds segment fields do not exist in the 64-bit layout. The
// size matters: the copy fills all 512 bytes, so a short or misaligned
// struct makes PTRACE_GETFPREGS overflow the caller's memory.
type FPRegs struct { type FPRegs struct {
FCW uint16 FCW uint16
FSW uint16 FSW uint16
FTW byte FTW uint16
FOP uint16 FOP uint16
FIP uint64 FIP uint64
FCS uint16
FDP uint64 FDP uint64
FDS uint16
MXCSR uint32 MXCSR uint32
MXCSRMask uint32 MXCSRMask uint32
ST [8][16]byte // x87 stack (10 bytes per reg, padded to 16) ST [8][16]byte // x87 stack (10 bytes per reg, padded to 16)
XMM [16][16]byte // XMM0-15 XMM [16][16]byte // XMM0-15, struct offset 160
Reserved [96]byte // FXSAVE padding, to the full 512 bytes
} }
// GetFPRegs retrieves the FPU/SSE register state of the stopped debuggee. // GetFPRegs retrieves the FPU/SSE register state of the stopped debuggee.
@@ -82,16 +87,68 @@ type VectorRegs struct {
YMM [16][32]byte // YMM0-15 (full 256-bit values) YMM [16][32]byte // YMM0-15 (full 256-bit values)
} }
// GetVectorRegs retrieves the YMM registers via PTRACE_GETREGSET + XSAVE. // NT_X86_XSTATE (0x202), the xsave extended-state regset
// (include/uapi/linux/elf.h).
const ntX86XState = 0x202
// Layout of the buffer PTRACE_GETREGSET returns for NT_X86_XSTATE: the
// 512-byte legacy fxsave image (x87 state in 0-159, XMM0-15 in 160-511),
// then the 64-byte xsave header whose first 8 bytes are xstate_bv, then one
// component per set feature bit, each 64-byte aligned. The YMM high halves
// are the first extended component, at offset 576; that offset is fixed by
// the ISA on AVX-capable x86-64. XFEATURE_MASK_YMM is bit 2 of xstate_bv
// (arch/x86/include/asm/fpu/types.h); the high halves are zero when the bit
// is clear.
const (
xsaveXMMOffset = 160
xsaveXMMSize = 256
xsaveHeaderOffset = 512
xsaveBVOffset = xsaveHeaderOffset
ymmOffset = xsaveHeaderOffset + 64 // 576
ymmSize = 256 // 16 registers, 16 bytes each
xfeatureMaskYMM = 1 << 2
xstateMaxBuffer = 4096 // CPUID(0xD).xsave_size is far below this
)
// GetVectorRegs retrieves the YMM registers via PTRACE_GETREGSET on
// NT_X86_XSTATE. The low (XMM) halves always come from the legacy image;
// the high halves are copied only when xstate_bv reports the YMM feature,
// and read as zero otherwise. When the regset request fails the FP image
// still provides correct XMM halves, so that is the fallback.
func (s *Session) GetVectorRegs() (VectorRegs, error) { func (s *Session) GetVectorRegs() (VectorRegs, error) {
var v VectorRegs var v VectorRegs
fp, err := s.GetFPRegs() buf := make([]byte, xstateMaxBuffer)
if err != nil { iovec := syscall.Iovec{
return v, err Base: &buf[0],
Len: uint64(len(buf)),
} }
_, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE,
uintptr(syscall.PTRACE_GETREGSET),
uintptr(s.pid),
uintptr(ntX86XState),
uintptr(unsafe.Pointer(&iovec)),
0, 0,
)
if errno != 0 {
fp, err := s.GetFPRegs()
if err != nil {
return v, err
}
for i := range 16 {
copy(v.YMM[i][:16], fp.XMM[i][:])
}
return v, nil
}
n := int(iovec.Len)
for i := range 16 { for i := range 16 {
for j := range 16 { copy(v.YMM[i][:16], buf[xsaveXMMOffset+16*i:xsaveXMMOffset+16*i+16])
v.YMM[i][j] = fp.XMM[i][j] }
if n >= ymmOffset+ymmSize {
if binary.LittleEndian.Uint64(buf[xsaveBVOffset:xsaveBVOffset+8])&xfeatureMaskYMM != 0 {
for i := range 16 {
copy(v.YMM[i][16:], buf[ymmOffset+16*i:ymmOffset+16*i+16])
}
} }
} }
return v, nil return v, nil
+5 -1
View File
@@ -91,5 +91,9 @@ func (r *Regs) RegValue(name string) (uint64, bool) {
// breakpointInsn is the software breakpoint instruction. // breakpointInsn is the software breakpoint instruction.
var breakpointInsn = []byte{0xCC} // INT3 var breakpointInsn = []byte{0xCC} // INT3
// breakpointPCAdjust is how far PC is past the breakpoint instruction after a trap. // breakpointPCAdjust is how far PC is past the breakpoint instruction after
// a trap. x86-64 reports the #DB for INT3 with RIP on the byte after the
// INT3 (Intel SDM vol 3, "Debug Exceptions"), so the trap address is
// PC-1. The other supported architectures leave the PC on the trap
// instruction and use 0 there.
const breakpointPCAdjust = 1 const breakpointPCAdjust = 1
+8 -2
View File
@@ -130,5 +130,11 @@ func (r *Regs) RegValue(name string) (uint64, bool) {
// breakpointInsn is the software breakpoint instruction (BRK #0). // breakpointInsn is the software breakpoint instruction (BRK #0).
var breakpointInsn = []byte{0x00, 0x00, 0x20, 0xD4} // BRK #0 var breakpointInsn = []byte{0x00, 0x00, 0x20, 0xD4} // BRK #0
// breakpointPCAdjust is how far PC is past the breakpoint instruction after a trap. // breakpointPCAdjust is how far PC is past the breakpoint instruction after
const breakpointPCAdjust = 4 // a trap: 0, because the arm64 kernel delivers the BRK SIGTRAP with the PC
// still on the BRK. do_el0_brk64 calls send_user_sigtrap, which uses
// instruction_pointer(regs) unmodified (arch/arm64/kernel/debug-monitors.c);
// only the kernel-internal skip paths advance the PC. GDB history agrees:
// decr_pc_after_break on aarch64 Linux is 0 (the +4 variant was a QEMU bug,
// sourceware PR 17280).
const breakpointPCAdjust = 0
+6 -2
View File
@@ -126,5 +126,9 @@ func (r *Regs) RegValue(name string) (uint64, bool) {
// breakpointInsn is the software breakpoint instruction (BRK $0). // breakpointInsn is the software breakpoint instruction (BRK $0).
var breakpointInsn = []byte{0x05, 0x00, 0x2a, 0x00} // break 0 var breakpointInsn = []byte{0x05, 0x00, 0x2a, 0x00} // break 0
// breakpointPCAdjust is how far PC is past the breakpoint instruction after a trap. // breakpointPCAdjust is how far PC is past the breakpoint instruction after
const breakpointPCAdjust = 4 // a trap: 0, because the kernel delivers the break SIGTRAP with csr_era
// still on the break instruction. do_bp passes regs->csr_era straight to
// force_sig_fault(SIGTRAP, TRAP_BRKPT, ...) and never adjusts era on the
// signal path (arch/loongarch/kernel/traps.c).
const breakpointPCAdjust = 0
+6 -2
View File
@@ -126,5 +126,9 @@ func (r *Regs) RegValue(name string) (uint64, bool) {
// breakpointInsn is the software breakpoint instruction (EBREAK). // breakpointInsn is the software breakpoint instruction (EBREAK).
var breakpointInsn = []byte{0x73, 0x00, 0x10, 0x00} // ebreak var breakpointInsn = []byte{0x73, 0x00, 0x10, 0x00} // ebreak
// breakpointPCAdjust is how far PC is past the breakpoint instruction after a trap. // breakpointPCAdjust is how far PC is past the breakpoint instruction after
const breakpointPCAdjust = 4 // a trap: 0, because the kernel delivers the EBREAK SIGTRAP with sepc still
// on the ebreak. handle_break passes regs->epc straight to
// force_sig_fault(SIGTRAP, TRAP_BRKPT, ...) and only the kernel-internal
// WARN/CFI paths advance epc (arch/riscv/kernel/traps.c).
const breakpointPCAdjust = 0
+74 -18
View File
@@ -7,9 +7,10 @@ package debug
import ( import (
"bufio" "bufio"
"cmp"
"fmt" "fmt"
"io" "io"
"sort" "slices"
"strconv" "strconv"
"strings" "strings"
) )
@@ -93,7 +94,7 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
regs, _ := s.GetRegs() regs, _ := s.GetRegs()
pc := regs.GetPC() pc := regs.GetPC()
text, instLen, _ := s.Disassemble(pc) text, instLen, _ := s.Disassemble(pc)
if strings.HasPrefix(strings.ToLower(text), "call") || strings.HasPrefix(strings.ToLower(text), "bl") { if isCallInsn(text) {
afterAddr := pc + uint64(instLen) afterAddr := pc + uint64(instLen)
_, err := bm.Set(afterAddr, "(next)") _, err := bm.Set(afterAddr, "(next)")
if err != nil { if err != nil {
@@ -108,6 +109,21 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
bm.Clear(afterAddr) bm.Clear(afterAddr)
continue continue
} }
if s.Exited() {
bm.Clear(afterAddr)
fmt.Println("debuggee exited")
continue
}
if sig := s.LastSignal(); sig != 0 {
bm.Clear(afterAddr)
regs, _ := s.GetRegs()
fmt.Printf("stopped on signal %v at %#x\n", sig, regs.GetPC())
continue
}
// Fetch the registers after the stop: the trap must be
// evaluated against the real PC, not the pre-Continue
// snapshot, and a stale SetRegs would clobber live state.
regs, _ = s.GetRegs()
bm.HandleTrap(&regs) bm.HandleTrap(&regs)
bm.Clear(afterAddr) bm.Clear(afterAddr)
} else { } else {
@@ -143,9 +159,21 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
bm.Clear(retAddr) bm.Clear(retAddr)
continue continue
} }
if !s.Exited() { if s.Exited() {
bm.HandleTrap(&regs) bm.Clear(retAddr)
fmt.Println("debuggee exited")
continue
} }
if sig := s.LastSignal(); sig != 0 {
bm.Clear(retAddr)
regs, _ := s.GetRegs()
fmt.Printf("stopped on signal %v at %#x\n", sig, regs.GetPC())
continue
}
// Fetch the registers after the stop, as the continue case
// does: HandleTrap must see the PC the trap left behind.
regs, _ = s.GetRegs()
bm.HandleTrap(&regs)
bm.Clear(retAddr) bm.Clear(retAddr)
if s.Exited() { if s.Exited() {
fmt.Println("debuggee exited") fmt.Println("debuggee exited")
@@ -171,6 +199,15 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
fmt.Println("debuggee exited") fmt.Println("debuggee exited")
break break
} }
if sig := s.LastSignal(); sig != 0 {
// A genuine signal-delivery-stop (a fault): report it
// and return to the prompt. Continuing would restart
// the faulting instruction and fault forever.
regs, _ := s.GetRegs()
fmt.Printf("stopped on signal %v at %#x (func+%#x)\n",
sig, regs.GetPC(), regs.GetPC()-codeBase-uint64(funcOffset))
break
}
reason, wpAddr := s.StopInfo() reason, wpAddr := s.StopInfo()
if reason == StopWatchpoint { if reason == StopWatchpoint {
fmt.Printf("watchpoint hit at %#x\n", wpAddr) fmt.Printf("watchpoint hit at %#x\n", wpAddr)
@@ -221,13 +258,27 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
reg := strings.ToLower(parts[3]) reg := strings.ToLower(parts[3])
op := parts[4] op := parts[4]
operand := parts[5] operand := parts[5]
if val, err := strconv.ParseUint(operand, 0, 64); err == nil { switch {
cond = &Condition{Reg: reg, Op: op, Value: val} case strings.HasPrefix(operand, "*"):
} else { // Memory operand: compare against the 8-byte word at
cond = &Condition{Reg: reg, Op: op, Reg2: strings.ToLower(operand)} // the address, resolved in the debuggee when the
// breakpoint is evaluated.
addr, err := strconv.ParseUint(strings.TrimPrefix(operand, "*"), 0, 64)
if err != nil {
fmt.Printf("invalid memory operand: %s\n", operand)
continue
}
cond = &Condition{Reg: reg, Op: op, MemAddr: addr}
default:
val, err := strconv.ParseUint(operand, 0, 64)
if err == nil {
cond = &Condition{Reg: reg, Op: op, Value: val}
} else {
cond = &Condition{Reg: reg, Op: op, Reg2: strings.ToLower(operand)}
}
} }
} else if len(parts) >= 4 && parts[2] == "if" { } else if len(parts) >= 4 && parts[2] == "if" {
fmt.Println("usage: break <label|addr> if <reg> <op> <value|reg>") fmt.Println("usage: break <label|addr> if <reg> <op> <value|reg|*addr>")
continue continue
} }
bp, err := bm.SetWithCond(addr, label, cond) bp, err := bm.SetWithCond(addr, label, cond)
@@ -237,7 +288,7 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
} }
condStr := "" condStr := ""
if cond != nil { if cond != nil {
condStr = fmt.Sprintf(" if %s %s %#x", cond.Reg, cond.Op, cond.Value) condStr = " if " + cond.String()
} }
fmt.Printf("breakpoint set: %s at %#x (func+%#x)%s\n", bp.Label, bp.Addr, bp.Addr-codeBase-uint64(funcOffset), condStr) fmt.Printf("breakpoint set: %s at %#x (func+%#x)%s\n", bp.Label, bp.Addr, bp.Addr-codeBase-uint64(funcOffset), condStr)
@@ -277,7 +328,11 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
addr, _ = resolveAddr(parts[1], codeBase, uint64(funcOffset), labels) addr, _ = resolveAddr(parts[1], codeBase, uint64(funcOffset), labels)
} }
if len(parts) > 2 { if len(parts) > 2 {
length, _ = strconv.Atoi(parts[2]) // A malformed or non-positive length would panic
// ReadMemory's make; fall back to the default instead.
if n, err := strconv.Atoi(parts[2]); err == nil && n > 0 {
length = n
}
} }
mem, err := s.ReadMemory(addr, length) mem, err := s.ReadMemory(addr, length)
if err != nil { if err != nil {
@@ -336,10 +391,8 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
} }
case "labels", "l": case "labels", "l":
sorted := make([]Label, len(labels)) slices.SortFunc(labels, func(a, b Label) int { return cmp.Compare(a.Offset, b.Offset) })
copy(sorted, labels) for _, l := range labels {
sort.Slice(sorted, func(i, j int) bool { return sorted[i].Offset < sorted[j].Offset })
for _, l := range sorted {
fmt.Printf(" func+%#04x %s\n", l.Offset, l.Name) fmt.Printf(" func+%#04x %s\n", l.Offset, l.Name)
} }
@@ -369,7 +422,10 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
fmt.Println() fmt.Println()
case "help", "h", "?": case "help", "h", "?":
fmt.Printf(` break <label|addr> [if <reg> <op> <val>] set a breakpoint fmt.Printf(` break <label|addr> [if <reg> <op> <val|reg|*addr>]
set a breakpoint, optionally conditional on a
register compared to a constant, a register, or the
8-byte word at *addr
delete <label|addr> remove a breakpoint delete <label|addr> remove a breakpoint
info break list all breakpoints info break list all breakpoints
watch <addr> [r|w] [size] set a hardware watchpoint (write by default) watch <addr> [r|w] [size] set a hardware watchpoint (write by default)
@@ -463,8 +519,8 @@ func REPL(s *Session, bm *Breakpoints, codeBase uint64, funcOffset, funcSize, ar
case "unwatch": case "unwatch":
if len(parts) >= 2 { if len(parts) >= 2 {
slot, err := strconv.Atoi(parts[1]) slot, err := strconv.Atoi(parts[1])
if err != nil || slot < 0 || slot > 3 { if err != nil || slot < 0 || slot >= maxWatchpoints() {
fmt.Println("usage: unwatch [<slot>]") fmt.Printf("usage: unwatch [<slot 0-%d>]\n", maxWatchpoints()-1)
continue continue
} }
if err := s.ClearWatchpoint(slot); err != nil { if err := s.ClearWatchpoint(slot); err != nil {
+10 -2
View File
@@ -6,6 +6,7 @@
package debug package debug
import ( import (
"encoding/binary"
"syscall" "syscall"
"unsafe" "unsafe"
) )
@@ -61,8 +62,15 @@ func (s *Session) StopInfo() (StopReason, uint64) {
case trapBRKPT: case trapBRKPT:
return StopBreakpoint, 0 return StopBreakpoint, 0
case trapHWBRKPT: case trapHWBRKPT:
addr := *(*uint64)(unsafe.Add(unsafe.Pointer(&info), 16)) // si_addr sits at struct offset 16 (12 bytes of signo/errno/code
return StopWatchpoint, addr // plus 4 bytes of union alignment). The siginfo buffer is only
// 4-byte aligned, so the address is read byte-wise to keep the
// load aligned on riscv64 and loong64. What si_addr names is
// architecture-specific (the data address on arm64, the
// instruction pointer on x86), so the per-architecture
// archWatchpointAddr resolves it to the watched address.
addr := binary.LittleEndian.Uint64(info._pad[4:12])
return StopWatchpoint, archWatchpointAddr(s, addr)
default: default:
return StopSingleStep, 0 return StopSingleStep, 0
} }
+1 -9
View File
@@ -28,15 +28,7 @@ func RunTarget(asmPath, funcName, argsFile, tmpDir string) error {
return fmt.Errorf("debug target: parse: %v", errs[0]) return fmt.Errorf("debug target: parse: %v", errs[0])
} }
var img *asm.Image img, err := asm.AssembleFileARM64(file)
switch "arm64" {
case "arm64":
img, err = asm.AssembleFileARM64(file)
case "riscv64":
img, err = asm.AssembleFileRISCV(file)
case "loong64":
img, err = asm.AssembleFileLOONG64(file)
}
if err != nil { if err != nil {
return fmt.Errorf("debug target: assemble: %w", err) return fmt.Errorf("debug target: assemble: %w", err)
} }
+1 -9
View File
@@ -28,15 +28,7 @@ func RunTarget(asmPath, funcName, argsFile, tmpDir string) error {
return fmt.Errorf("debug target: parse: %v", errs[0]) return fmt.Errorf("debug target: parse: %v", errs[0])
} }
var img *asm.Image img, err := asm.AssembleFileLOONG64(file)
switch "loong64" {
case "arm64":
img, err = asm.AssembleFileARM64(file)
case "riscv64":
img, err = asm.AssembleFileRISCV(file)
case "loong64":
img, err = asm.AssembleFileLOONG64(file)
}
if err != nil { if err != nil {
return fmt.Errorf("debug target: assemble: %w", err) return fmt.Errorf("debug target: assemble: %w", err)
} }
+8 -2
View File
@@ -12,10 +12,11 @@ type tracer interface {
Peek(addr uint64) (uint64, error) Peek(addr uint64) (uint64, error)
Poke(addr uint64, val uint64) error Poke(addr uint64, val uint64) error
SetRegs(regs *Regs) error SetRegs(regs *Regs) error
Step() error
Pid() int Pid() int
} }
// mockTracer records Peek/Poke calls and provides fake register state. // mockTracer records Peek/Poke/Step calls and provides fake register state.
type mockTracer struct { type mockTracer struct {
mem map[uint64]byte mem map[uint64]byte
peeks []uint64 peeks []uint64
@@ -23,7 +24,8 @@ type mockTracer struct {
addr uint64 addr uint64
val uint64 val uint64
} }
regs *Regs steps int
regs *Regs
} }
func newMockTracer() *mockTracer { func newMockTracer() *mockTracer {
@@ -57,4 +59,8 @@ func (m *mockTracer) SetRegs(regs *Regs) error {
m.regs = regs m.regs = regs
return nil return nil
} }
func (m *mockTracer) Step() error {
m.steps++
return nil
}
func (m *mockTracer) Pid() int { return 42 } func (m *mockTracer) Pid() int { return 42 }
+59 -31
View File
@@ -8,10 +8,48 @@ package debug
import ( import (
"fmt" "fmt"
"syscall" "syscall"
"unsafe"
) )
// Hardware watchpoint support via x86-64 debug registers (DR0-DR3, DR7). // Hardware watchpoint support via x86-64 debug registers (DR0-DR3, DR7).
// The kernel translates PTRACE_POKEUSER/PEEKUSER offsets inside
// [offsetof(struct user, u_debugreg[0]), u_debugreg[7]] to DR0-DR7
// (arch/x86/kernel/ptrace.c, arch_ptrace). sys/user.h places u_debugreg at
// 0x350: DR0-DR3 are 0x350/0x358/0x360/0x368, DR6 (status) is 0x380 and
// DR7 (control) is 0x388. Offsets below 0x350 write user_regs_struct
// fields (r15 at 0x0, r10 at 0x38), not debug registers.
const (
drOffset = 0x350 // offsetof(struct user, u_debugreg[0]), DR0
dr6Off = 0x380 // offsetof(struct user, u_debugreg[6]), DR6
dr7Off = 0x388 // offsetof(struct user, u_debugreg[7]), DR7
)
// archWatchpointAddr resolves the address of the watchpoint that fired.
// x86 delivers si_addr = the instruction pointer of the trapping access
// (arch/x86/kernel/ptrace.c send_sigtrap passes regs->ip), so the watched
// data address is recovered from DR6's slot bits (B0-B3, positive polarity
// through PEEKUSER) and the matching DR0-DR3.
func archWatchpointAddr(s *Session, siAddr uint64) uint64 {
dr6, err := ptracePeekUser(s.pid, dr6Off)
if err != nil {
return siAddr
}
for slot := range 4 {
if dr6&(1<<slot) != 0 {
addr, err := ptracePeekUser(s.pid, drOffset+uintptr(slot*8))
if err == nil && addr != 0 {
return addr
}
}
}
return siAddr
}
// maxWatchpoints reports the number of hardware watchpoint slots the
// architecture provides: four address registers, DR0-DR3.
func maxWatchpoints() int { return 4 }
// WatchpointType selects what triggers the watchpoint. // WatchpointType selects what triggers the watchpoint.
type WatchpointType int type WatchpointType int
@@ -20,14 +58,11 @@ const (
WatchRead WatchpointType = 3 // trigger on read or write WatchRead WatchpointType = 3 // trigger on read or write
) )
// wpSlots tracks watchpoint slot occupancy (DR0-DR3).
var wpSlots [4]bool
// FindFreeWatchpointSlot returns the index of the first free watchpoint slot // FindFreeWatchpointSlot returns the index of the first free watchpoint slot
// (0-3), or -1 if all four hardware watchpoints are in use. // (0-3), or -1 if all four hardware watchpoints are in use.
func (s *Session) FindFreeWatchpointSlot() int { func (s *Session) FindFreeWatchpointSlot() int {
for i := range 4 { for i := range 4 {
if !wpSlots[i] { if !s.wpSlots[i] {
return i return i
} }
} }
@@ -39,7 +74,7 @@ func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot > 3 { if slot < 0 || slot > 3 {
return false return false
} }
return wpSlots[slot] return s.wpSlots[slot]
} }
// SetWatchpoint installs a hardware watchpoint on the given address. // SetWatchpoint installs a hardware watchpoint on the given address.
@@ -47,7 +82,7 @@ func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size
if slot < 0 || slot > 3 { if slot < 0 || slot > 3 {
return fmt.Errorf("debug: watchpoint slot must be 0-3") return fmt.Errorf("debug: watchpoint slot must be 0-3")
} }
if wpSlots[slot] { if s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot) return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
} }
@@ -65,23 +100,11 @@ func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8") return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
} }
var drAddr uintptr if err := ptracePokeUser(s.pid, drOffset+uintptr(slot*8), addr); err != nil {
switch slot {
case 0:
drAddr = 0x0
case 1:
drAddr = 0x8
case 2:
drAddr = 0x10
case 3:
drAddr = 0x18
}
if err := ptracePokeUser(s.pid, drAddr, addr); err != nil {
return fmt.Errorf("debug: set DR%d: %w", slot, err) return fmt.Errorf("debug: set DR%d: %w", slot, err)
} }
dr7, err := ptracePeekUser(s.pid, 0x38) dr7, err := ptracePeekUser(s.pid, dr7Off)
if err != nil { if err != nil {
return fmt.Errorf("debug: read DR7: %w", err) return fmt.Errorf("debug: read DR7: %w", err)
} }
@@ -93,10 +116,10 @@ func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size
mask := ^((uint64(1) << (2 * slot)) | (uint64(3) << (16 + 4*slot)) | (uint64(3) << (18 + 4*slot))) mask := ^((uint64(1) << (2 * slot)) | (uint64(3) << (16 + 4*slot)) | (uint64(3) << (18 + 4*slot)))
dr7 = (dr7 & mask) | enableBit | rwBits | lenField dr7 = (dr7 & mask) | enableBit | rwBits | lenField
if err := ptracePokeUser(s.pid, 0x38, dr7); err != nil { if err := ptracePokeUser(s.pid, dr7Off, dr7); err != nil {
return fmt.Errorf("debug: set DR7: %w", err) return fmt.Errorf("debug: set DR7: %w", err)
} }
wpSlots[slot] = true s.wpSlots[slot] = true
return nil return nil
} }
@@ -105,25 +128,25 @@ func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot > 3 { if slot < 0 || slot > 3 {
return fmt.Errorf("debug: watchpoint slot must be 0-3") return fmt.Errorf("debug: watchpoint slot must be 0-3")
} }
if !wpSlots[slot] { if !s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot) return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
} }
dr7, err := ptracePeekUser(s.pid, 0x38) dr7, err := ptracePeekUser(s.pid, dr7Off)
if err != nil { if err != nil {
return err return err
} }
dr7 &^= uint64(1) << (2 * slot) dr7 &^= uint64(1) << (2 * slot)
if err := ptracePokeUser(s.pid, 0x38, dr7); err != nil { if err := ptracePokeUser(s.pid, dr7Off, dr7); err != nil {
return err return err
} }
wpSlots[slot] = false s.wpSlots[slot] = false
return nil return nil
} }
// ClearAllWatchpoints removes all hardware watchpoints. // ClearAllWatchpoints removes all hardware watchpoints.
func (s *Session) ClearAllWatchpoints() error { func (s *Session) ClearAllWatchpoints() error {
for slot := range 4 { for slot := range maxWatchpoints() {
if wpSlots[slot] { if s.wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil { if err := s.ClearWatchpoint(slot); err != nil {
return err return err
} }
@@ -149,16 +172,21 @@ func ptracePokeUser(pid int, offset uintptr, val uint64) error {
} }
func ptracePeekUser(pid int, offset uintptr) (uint64, error) { func ptracePeekUser(pid int, offset uintptr) (uint64, error) {
// x86 PEEKUSR writes the word to the user-space pointer in data
// (arch/x86/kernel/ptrace.c uses put_user); passing 0 there fails with
// EFAULT, so the word is read through a real address.
const ptracePeekuser = 3 const ptracePeekuser = 3
val, _, errno := syscall.Syscall6( var word uint64
_, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE, syscall.SYS_PTRACE,
uintptr(ptracePeekuser), uintptr(ptracePeekuser),
uintptr(pid), uintptr(pid),
offset, offset,
0, 0, 0, uintptr(unsafe.Pointer(&word)),
0, 0,
) )
if errno != 0 { if errno != 0 {
return 0, errno return 0, errno
} }
return uint64(val), nil return word, nil
} }
+48 -38
View File
@@ -12,7 +12,7 @@ import (
) )
// Hardware watchpoint support via arm64 debug registers (DBGWVR/DBGWCR). // Hardware watchpoint support via arm64 debug registers (DBGWVR/DBGWCR).
// Accessed via PTRACE_SETREGSET with NT_ARM_HW_BREAK. // Accessed via PTRACE_GETREGSET/SETREGSET with NT_ARM_HW_WATCH.
// WatchpointType selects what triggers the watchpoint. // WatchpointType selects what triggers the watchpoint.
type WatchpointType int type WatchpointType int
@@ -22,30 +22,36 @@ const (
WatchRead WatchpointType = 3 WatchRead WatchpointType = 3
) )
// wpSlots tracks watchpoint slot occupancy. // maxWatchpoints reports the number of hardware watchpoint slots the
var wpSlots [16]bool // arm64 supports up to 16 watchpoints // architecture provides: DBGWVR0-DBGWCR15.
func maxWatchpoints() int { return 16 }
const maxWatchpoints = 16 // hwWatchState mirrors the kernel's struct user_hwdebug_state.
type hwWatchState struct {
// hwBreakState mirrors the kernel's struct user_hwdebug_state.
type hwBreakState struct {
DbgInfo uint32 DbgInfo uint32
_pad [4]byte _pad [4]byte
DbgRegs [16]hwBreakReg DbgRegs [16]hwWatchReg
} }
type hwBreakReg struct { type hwWatchReg struct {
Addr uint64 Addr uint64
Ctrl uint64 Ctrl uint64
} }
const ( // ntArmHWWatch is NT_ARM_HW_WATCH (0x403), the watchpoint regset
ntArmHWBreak = 0x403 // NT_ARM_HW_BREAK // (include/uapi/linux/elf.h; 0x402 is NT_ARM_HW_BREAK). Watchpoints and
) // breakpoints live in different regsets with the same struct shape, so the
// constant is named for what it arms to keep a future edit from arming
// breakpoints instead.
const ntArmHWWatch = 0x403
// archWatchpointAddr resolves the address of the watchpoint that fired:
// the arm64 kernel already reports the watched data address as si_addr.
func archWatchpointAddr(s *Session, siAddr uint64) uint64 { return siAddr }
func (s *Session) FindFreeWatchpointSlot() int { func (s *Session) FindFreeWatchpointSlot() int {
for i := range maxWatchpoints { for i := range maxWatchpoints() {
if !wpSlots[i] { if !s.wpSlots[i] {
return i return i
} }
} }
@@ -53,35 +59,39 @@ func (s *Session) FindFreeWatchpointSlot() int {
} }
func (s *Session) IsWatchpointSlotUsed(slot int) bool { func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return false return false
} }
return wpSlots[slot] return s.wpSlots[slot]
} }
// SetWatchpoint installs a hardware watchpoint on the given address. // SetWatchpoint installs a hardware watchpoint on the given address.
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error { func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1) return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
} }
if wpSlots[slot] { if s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot) return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
} }
state, err := s.getHWBreakState() state, err := s.getHWWatchState()
if err != nil { if err != nil {
return fmt.Errorf("debug: read watchpoint state: %w", err) return fmt.Errorf("debug: read watchpoint state: %w", err)
} }
if uint32(slot) >= state.DbgInfo { // MDSCR_EL1 packs (debug_arch << 8) | num_slots into dbg_info, so only
return fmt.Errorf("debug: slot %d exceeds available watchpoints (%d)", slot, state.DbgInfo) // the low byte counts slots.
if uint32(slot) >= state.DbgInfo&0xff {
return fmt.Errorf("debug: slot %d exceeds available watchpoints (%d)", slot, state.DbgInfo&0xff)
} }
state.DbgRegs[slot].Addr = addr state.DbgRegs[slot].Addr = addr
// DBGWCR bits 3-4 select the access type: 01 load, 10 store, 11 either
// (ARM DDI 0487, DBGWCR<n>_EL1 watchpoint type field).
ctrl := uint64(1) // enable ctrl := uint64(1) // enable
switch typ { switch typ {
case WatchWrite: case WatchWrite:
ctrl |= 1 << 3 // store only ctrl |= 2 << 3 // store only
case WatchRead: case WatchRead:
ctrl |= 3 << 3 // load+store ctrl |= 3 << 3 // load+store
} }
@@ -101,38 +111,38 @@ func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size
ctrl |= bas << 5 ctrl |= bas << 5
state.DbgRegs[slot].Ctrl = ctrl state.DbgRegs[slot].Ctrl = ctrl
if err := s.setHWBreakState(state); err != nil { if err := s.setHWWatchState(state); err != nil {
return fmt.Errorf("debug: set watchpoint: %w", err) return fmt.Errorf("debug: set watchpoint: %w", err)
} }
wpSlots[slot] = true s.wpSlots[slot] = true
return nil return nil
} }
func (s *Session) ClearWatchpoint(slot int) error { func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1) return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
} }
if !wpSlots[slot] { if !s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot) return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
} }
state, err := s.getHWBreakState() state, err := s.getHWWatchState()
if err != nil { if err != nil {
return err return err
} }
state.DbgRegs[slot].Addr = 0 state.DbgRegs[slot].Addr = 0
state.DbgRegs[slot].Ctrl = 0 state.DbgRegs[slot].Ctrl = 0
if err := s.setHWBreakState(state); err != nil { if err := s.setHWWatchState(state); err != nil {
return err return err
} }
wpSlots[slot] = false s.wpSlots[slot] = false
return nil return nil
} }
func (s *Session) ClearAllWatchpoints() error { func (s *Session) ClearAllWatchpoints() error {
for slot := 0; slot < maxWatchpoints; slot++ { for slot := range maxWatchpoints() {
if wpSlots[slot] { if s.wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil { if err := s.ClearWatchpoint(slot); err != nil {
return err return err
} }
@@ -141,8 +151,8 @@ func (s *Session) ClearAllWatchpoints() error {
return nil return nil
} }
func (s *Session) getHWBreakState() (*hwBreakState, error) { func (s *Session) getHWWatchState() (*hwWatchState, error) {
var state hwBreakState var state hwWatchState
iovec := syscall.Iovec{ iovec := syscall.Iovec{
Base: (*byte)(unsafe.Pointer(&state)), Base: (*byte)(unsafe.Pointer(&state)),
Len: uint64(unsafe.Sizeof(state)), Len: uint64(unsafe.Sizeof(state)),
@@ -151,7 +161,7 @@ func (s *Session) getHWBreakState() (*hwBreakState, error) {
syscall.SYS_PTRACE, syscall.SYS_PTRACE,
uintptr(syscall.PTRACE_GETREGSET), uintptr(syscall.PTRACE_GETREGSET),
uintptr(s.pid), uintptr(s.pid),
uintptr(ntArmHWBreak), uintptr(ntArmHWWatch),
uintptr(unsafe.Pointer(&iovec)), uintptr(unsafe.Pointer(&iovec)),
0, 0, 0, 0,
) )
@@ -161,7 +171,7 @@ func (s *Session) getHWBreakState() (*hwBreakState, error) {
return &state, nil return &state, nil
} }
func (s *Session) setHWBreakState(state *hwBreakState) error { func (s *Session) setHWWatchState(state *hwWatchState) error {
iovec := syscall.Iovec{ iovec := syscall.Iovec{
Base: (*byte)(unsafe.Pointer(state)), Base: (*byte)(unsafe.Pointer(state)),
Len: uint64(unsafe.Sizeof(*state)), Len: uint64(unsafe.Sizeof(*state)),
@@ -170,7 +180,7 @@ func (s *Session) setHWBreakState(state *hwBreakState) error {
syscall.SYS_PTRACE, syscall.SYS_PTRACE,
uintptr(syscall.PTRACE_SETREGSET), uintptr(syscall.PTRACE_SETREGSET),
uintptr(s.pid), uintptr(s.pid),
uintptr(ntArmHWBreak), uintptr(ntArmHWWatch),
uintptr(unsafe.Pointer(&iovec)), uintptr(unsafe.Pointer(&iovec)),
0, 0, 0, 0,
) )
+127 -69
View File
@@ -8,10 +8,47 @@ package debug
import ( import (
"fmt" "fmt"
"syscall" "syscall"
"unsafe"
) )
// Hardware watchpoint support for LoongArch via debug registers. // Hardware watchpoint support via the NT_LOONGARCH_HW_WATCH regset.
// Uses PTRACE_POKEUSER/PEEKUSER to access HW watchpoint registers. //
// The kernel's PTRACE_POKEUSER on loong64 accepts only the user_pt_regs
// indices 0-34 (GPRs, orig_a0, era, badv, per
// arch/loongarch/include/uapi/asm/ptrace.h), so there is no debug-register
// window to poke. The real interface is PTRACE_GETREGSET/SETREGSET on
// NT_LOONGARCH_HW_WATCH (0xa06, include/uapi/linux/elf.h) with struct
// user_watch_state_v2 (arch/loongarch/include/uapi/asm/ptrace.h): a dbg_info
// word followed by 14 slots of {addr u64, mask u64, ctrl u32, pad u32}.
// hw_break_get puts the slot count in the low byte of dbg_info
// (arch/loongarch/kernel/ptrace.c, ptrace_hbp_get_resource_info) and
// hw_break_set ignores dbg_info, reading addr, mask and ctrl per slot.
const ntLoongHWWatch = 0xa06
// loongWatchState mirrors the kernel's struct user_watch_state_v2.
type loongWatchState struct {
DbgInfo uint64
DbgRegs [14]loongWatchReg
}
type loongWatchReg struct {
Addr uint64
Mask uint64
Ctrl uint32
Pad uint32
}
// Control word bit layout (arch/loongarch/include/asm/hw_breakpoint.h):
// bits 1-4 privilege enables (CTRL_PLV3_ENABLE, 0x10, covers user mode),
// bits 8-9 access type (LOAD 1<<0, STORE 1<<1), bits 10-11 length
// (0=8 bytes, 1=4, 2=2, 3=1, inverted like the hardware FWP cfg).
const (
loongCtrlPLV3Enable = 0x10
loongTypeLoad = 1 << 8
loongTypeStore = 2 << 8
loongLenShift = 10
)
// WatchpointType selects what triggers the watchpoint. // WatchpointType selects what triggers the watchpoint.
type WatchpointType int type WatchpointType int
@@ -21,14 +58,17 @@ const (
WatchRead WatchpointType = 3 WatchRead WatchpointType = 3
) )
// wpSlots tracks watchpoint slot occupancy. // maxWatchpoints reports the slot capacity of the regset struct; the number
var wpSlots [4]bool // the hardware actually provides is read from dbg_info at arm time.
func maxWatchpoints() int { return len(loongWatchState{}.DbgRegs) }
const maxWatchpoints = 4 // archWatchpointAddr resolves the address of the watchpoint that fired:
// the loongarch kernel already reports the accessed address as si_addr.
func archWatchpointAddr(s *Session, siAddr uint64) uint64 { return siAddr }
func (s *Session) FindFreeWatchpointSlot() int { func (s *Session) FindFreeWatchpointSlot() int {
for i := range maxWatchpoints { for i := range maxWatchpoints() {
if !wpSlots[i] { if !s.wpSlots[i] {
return i return i
} }
} }
@@ -36,77 +76,87 @@ func (s *Session) FindFreeWatchpointSlot() int {
} }
func (s *Session) IsWatchpointSlotUsed(slot int) bool { func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return false return false
} }
return wpSlots[slot] return s.wpSlots[slot]
} }
// SetWatchpoint installs a hardware watchpoint. // SetWatchpoint installs a hardware watchpoint on the given address.
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error { func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1) return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
} }
if wpSlots[slot] { if s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot) return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
} }
if size != 1 && size != 2 && size != 4 && size != 8 {
var ctrlType uint32
switch typ {
case WatchWrite:
ctrlType = loongTypeStore
case WatchRead:
ctrlType = loongTypeLoad | loongTypeStore
}
var lenBits uint32
switch size {
case 1:
lenBits = 3
case 2:
lenBits = 2
case 4:
lenBits = 1
case 8:
lenBits = 0
default:
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8") return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
} }
// LoongArch debug registers: DBGWVR (watchpoint value) and DBGWCR (watchpoint control). state, err := s.getLoongWatchState()
// Accessed via PTRACE_POKEUSER at architecture-specific offsets. if err != nil {
if err := ptracePokeUser(s.pid, uintptr(0x1000+slot*8), addr); err != nil { return fmt.Errorf("debug: read watchpoint state: %w", err)
return fmt.Errorf("debug: set watchpoint address: %w", err) }
if uint64(slot) >= state.DbgInfo&0xff {
return fmt.Errorf("debug: slot %d exceeds available watchpoints (%d)", slot, state.DbgInfo&0xff)
} }
// DBGWCR: enable + type + size. state.DbgRegs[slot].Addr = addr
var wcr uint64 = 1 // enable state.DbgRegs[slot].Mask = 0
switch typ { state.DbgRegs[slot].Ctrl = loongCtrlPLV3Enable | ctrlType | lenBits<<loongLenShift
case WatchWrite:
wcr |= 1 << 3 // store
case WatchRead:
wcr |= 3 << 3 // load+store
}
var sizeBits uint64
switch size {
case 1:
sizeBits = 0
case 2:
sizeBits = 1
case 4:
sizeBits = 2
case 8:
sizeBits = 3
}
wcr |= sizeBits << 5
if err := ptracePokeUser(s.pid, uintptr(0x1001+slot*8), wcr); err != nil { if err := s.setLoongWatchState(state); err != nil {
return fmt.Errorf("debug: set watchpoint control: %w", err) return fmt.Errorf("debug: set watchpoint: %w", err)
} }
wpSlots[slot] = true s.wpSlots[slot] = true
return nil return nil
} }
func (s *Session) ClearWatchpoint(slot int) error { func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1) return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
} }
if !wpSlots[slot] { if !s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot) return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
} }
if err := ptracePokeUser(s.pid, uintptr(0x1001+slot*8), 0); err != nil { state, err := s.getLoongWatchState()
if err != nil {
return err return err
} }
wpSlots[slot] = false state.DbgRegs[slot].Addr = 0
state.DbgRegs[slot].Mask = 0
state.DbgRegs[slot].Ctrl = 0
if err := s.setLoongWatchState(state); err != nil {
return err
}
s.wpSlots[slot] = false
return nil return nil
} }
func (s *Session) ClearAllWatchpoints() error { func (s *Session) ClearAllWatchpoints() error {
for slot := 0; slot < maxWatchpoints; slot++ { for slot := range maxWatchpoints() {
if wpSlots[slot] { if s.wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil { if err := s.ClearWatchpoint(slot); err != nil {
return err return err
} }
@@ -115,14 +165,37 @@ func (s *Session) ClearAllWatchpoints() error {
return nil return nil
} }
func ptracePokeUser(pid int, offset uintptr, val uint64) error { func (s *Session) getLoongWatchState() (*loongWatchState, error) {
const ptracePokeuser = 6 var state loongWatchState
iovec := syscall.Iovec{
Base: (*byte)(unsafe.Pointer(&state)),
Len: uint64(unsafe.Sizeof(state)),
}
_, _, errno := syscall.Syscall6( _, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE, syscall.SYS_PTRACE,
uintptr(ptracePokeuser), uintptr(syscall.PTRACE_GETREGSET),
uintptr(pid), uintptr(s.pid),
offset, uintptr(ntLoongHWWatch),
uintptr(val), uintptr(unsafe.Pointer(&iovec)),
0, 0,
)
if errno != 0 {
return nil, errno
}
return &state, nil
}
func (s *Session) setLoongWatchState(state *loongWatchState) error {
iovec := syscall.Iovec{
Base: (*byte)(unsafe.Pointer(state)),
Len: uint64(unsafe.Sizeof(*state)),
}
_, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE,
uintptr(syscall.PTRACE_SETREGSET),
uintptr(s.pid),
uintptr(ntLoongHWWatch),
uintptr(unsafe.Pointer(&iovec)),
0, 0, 0, 0,
) )
if errno != 0 { if errno != 0 {
@@ -130,18 +203,3 @@ func ptracePokeUser(pid int, offset uintptr, val uint64) error {
} }
return nil return nil
} }
func ptracePeekUser(pid int, offset uintptr) (uint64, error) {
const ptracePeekuser = 3
val, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE,
uintptr(ptracePeekuser),
uintptr(pid),
offset,
0, 0, 0,
)
if errno != 0 {
return 0, errno
}
return uint64(val), nil
}
+27 -106
View File
@@ -7,11 +7,16 @@ package debug
import ( import (
"fmt" "fmt"
"syscall"
) )
// Hardware watchpoint support for RISC-V via Sdtrig trigger registers. // Hardware watchpoints are not reachable through the riscv64 kernel ptrace
// Uses PTRACE_POKEUSER/PEEKUSER to access debug registers. // interface. arch/riscv/kernel/ptrace.c forwards every POKEUSER/PEEKUSER to
// the generic ptrace_request, and the riscv user_regset view contains only
// the GPR, FP and vector regsets: there is no debug-register or trigger
// regset, and offsets outside the view fail with EIO. The Sdtrig CSRs
// (tselect/tdata1/tdata2) are not exposed to ptrace either. Until the
// kernel grows a trigger regset, SetWatchpoint reports the fact instead of
// poking a window that does not exist.
// WatchpointType selects what triggers the watchpoint. // WatchpointType selects what triggers the watchpoint.
type WatchpointType int type WatchpointType int
@@ -21,14 +26,19 @@ const (
WatchRead WatchpointType = 3 WatchRead WatchpointType = 3
) )
// wpSlots tracks watchpoint slot occupancy. // maxWatchpoints reports the number of hardware watchpoint slots the
var wpSlots [4]bool // architecture provides. riscv64 exposes none via ptrace; the bound exists
// so the slot bookkeeping stays consistent.
func maxWatchpoints() int { return 4 }
const maxWatchpoints = 4 // archWatchpointAddr resolves the address of the watchpoint that fired.
// Unreachable in practice (watchpoints cannot be armed), but si_addr names
// the accessed address where the kernel does report one.
func archWatchpointAddr(s *Session, siAddr uint64) uint64 { return siAddr }
func (s *Session) FindFreeWatchpointSlot() int { func (s *Session) FindFreeWatchpointSlot() int {
for i := range maxWatchpoints { for i := range maxWatchpoints() {
if !wpSlots[i] { if !s.wpSlots[i] {
return i return i
} }
} }
@@ -36,115 +46,26 @@ func (s *Session) FindFreeWatchpointSlot() int {
} }
func (s *Session) IsWatchpointSlotUsed(slot int) bool { func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return false return false
} }
return wpSlots[slot] return s.wpSlots[slot]
} }
// SetWatchpoint installs a hardware watchpoint. // SetWatchpoint always fails: the riscv64 kernel ptrace interface has no
// hardware-watchpoint access.
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error { func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
if slot < 0 || slot >= maxWatchpoints { return fmt.Errorf("debug: hardware watchpoints are not supported by the riscv64 kernel ptrace interface")
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1)
}
if wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
}
if size != 1 && size != 2 && size != 4 && size != 8 {
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
}
// RISC-V trigger registers: tdata1 encodes type/control, tdata2 holds address.
// The exact encoding depends on the trigger implementation (Sdtrig).
// Use PTRACE_POKEUSER to write to the trigger CSRs via the kernel's
// debug register interface.
if err := ptracePokeUser(s.pid, uintptr(0x1000+slot*8), addr); err != nil {
return fmt.Errorf("debug: set watchpoint address: %w", err)
}
// tdata1: set match control. Mode=2 (data match), select=0, action=1 (debug exception).
var tdata1 uint64 = 2 << 60 // type = match (2)
tdata1 |= 1 << 0 // action = enter debug mode
tdata1 |= 1 << 7 // store (write) trigger
if typ == WatchRead {
tdata1 |= 1 << 6 // load trigger
}
// Size encoding: 0=1byte, 1=2byte, 2=4byte, 3=8byte.
var sizeBits uint64
switch size {
case 1:
sizeBits = 0
case 2:
sizeBits = 1
case 4:
sizeBits = 2
case 8:
sizeBits = 3
}
tdata1 |= sizeBits << 16 // size field
if err := ptracePokeUser(s.pid, uintptr(0x1001+slot*8), tdata1); err != nil {
return fmt.Errorf("debug: set watchpoint control: %w", err)
}
wpSlots[slot] = true
return nil
} }
// ClearWatchpoint always fails: no watchpoint can ever be armed.
func (s *Session) ClearWatchpoint(slot int) error { func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot >= maxWatchpoints { if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints-1) return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
} }
if !wpSlots[slot] { return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
}
// Disable by clearing tdata1.
if err := ptracePokeUser(s.pid, uintptr(0x1001+slot*8), 0); err != nil {
return err
}
wpSlots[slot] = false
return nil
} }
func (s *Session) ClearAllWatchpoints() error { func (s *Session) ClearAllWatchpoints() error {
for slot := 0; slot < maxWatchpoints; slot++ {
if wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil {
return err
}
}
}
return nil return nil
} }
func ptracePokeUser(pid int, offset uintptr, val uint64) error {
const ptracePokeuser = 6
_, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE,
uintptr(ptracePokeuser),
uintptr(pid),
offset,
uintptr(val),
0, 0,
)
if errno != 0 {
return errno
}
return nil
}
func ptracePeekUser(pid int, offset uintptr) (uint64, error) {
const ptracePeekuser = 3
val, _, errno := syscall.Syscall6(
syscall.SYS_PTRACE,
uintptr(ptracePeekuser),
uintptr(pid),
offset,
0, 0, 0,
)
if errno != 0 {
return 0, errno
}
return uint64(val), nil
}
+95
View File
@@ -0,0 +1,95 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Package disasm decodes machine code back to instruction text for the four
// architectures gasm assembles. It is a thin, platform-independent wrapper
// over golang.org/x/arch and backs both the `gasm dis` command and the live
// debugger views.
package disasm
import (
"fmt"
"golang.org/x/arch/arm64/arm64asm"
"golang.org/x/arch/loong64/loong64asm"
"golang.org/x/arch/riscv64/riscv64asm"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
)
// Instruction is one decoded instruction: its text form, its length in bytes
// and the address it was decoded at.
type Instruction struct {
Addr uint64
Text string
Len int
}
// Decode decodes the instruction at the start of code, located at addr.
// code needs to hold at least the one instruction being decoded (amd64 may
// consume up to 15 bytes). Undecodable bytes yield the placeholder text "???"
// and a length of one word (four bytes, one on amd64) so that a listing can
// keep making progress, mirroring the debugger's behaviour.
func Decode(a arch.Arch, code []byte, addr uint64) (Instruction, error) {
if len(code) == 0 {
return Instruction{}, fmt.Errorf("disasm: empty input")
}
switch a {
case arch.ARM64:
if len(code) < 4 {
return Instruction{}, fmt.Errorf("disasm: need 4 bytes, have %d", len(code))
}
inst, err := arm64asm.Decode(code)
if err != nil {
return Instruction{Addr: addr, Text: "???", Len: 4}, nil
}
return Instruction{Addr: addr, Text: arm64asm.GoSyntax(inst, addr, nil, nil), Len: 4}, nil
case arch.RISCV:
// The compressed extensions are decoded transparently; a 16-bit
// instruction only needs its two bytes.
inst, err := riscv64asm.Decode(code)
if err != nil {
return Instruction{Addr: addr, Text: "???", Len: 2}, nil
}
return Instruction{Addr: addr, Text: riscv64asm.GoSyntax(inst, addr, nil, nil), Len: inst.Len}, nil
case arch.LOONG64:
if len(code) < 4 {
return Instruction{}, fmt.Errorf("disasm: need 4 bytes, have %d", len(code))
}
inst, err := loong64asm.Decode(code)
if err != nil {
return Instruction{Addr: addr, Text: "???", Len: 4}, nil
}
return Instruction{Addr: addr, Text: loong64asm.GoSyntax(inst, addr, nil), Len: 4}, nil
default: // amd64
inst, err := x86asm.Decode(code, 64)
if err != nil {
return Instruction{Addr: addr, Text: "???", Len: 1}, nil
}
return Instruction{Addr: addr, Text: x86asm.IntelSyntax(inst, addr, nil), Len: inst.Len}, nil
}
}
// Block decodes up to max instructions from code starting at addr and returns
// them in order. Decoding stops at the end of code or once an instruction
// would run past it.
func Block(a arch.Arch, code []byte, addr uint64, max int) []Instruction {
var out []Instruction
pc := 0
for len(out) < max && pc < len(code) {
ins, err := Decode(a, code[pc:], addr+uint64(pc))
if err != nil {
break
}
if ins.Len <= 0 || pc+ins.Len > len(code) {
break
}
out = append(out, ins)
pc += ins.Len
}
return out
}
+141
View File
@@ -0,0 +1,141 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package disasm
import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
func TestDecodeKnownBytes(t *testing.T) {
for _, tt := range []struct {
a arch.Arch
code []byte
text string
want int
}{
{arch.AMD64, []byte{0x55}, "push rbp", 1},
{arch.AMD64, []byte{0x48, 0x89, 0xE5}, "mov rbp, rsp", 3},
{arch.ARM64, []byte{0xc0, 0x03, 0x5f, 0xd6}, "RET", 4},
{arch.RISCV, []byte{0x67, 0x80, 0x00, 0x00}, "RET", 4},
{arch.LOONG64, []byte{0x20, 0x00, 0x00, 0x4c}, "RET", 4},
} {
ins, err := Decode(tt.a, tt.code, 0)
if err != nil {
t.Errorf("%s: %v", tt.a, err)
continue
}
if ins.Text != tt.text || ins.Len != tt.want {
t.Errorf("%s: % x decoded to %q (%d bytes), want %q (%d)",
tt.a, tt.code, ins.Text, ins.Len, tt.text, tt.want)
}
}
}
func TestDecodeUndecodable(t *testing.T) {
// Zero words do not encode a usable instruction on arm64 and loong64; the
// placeholder keeps a listing going. RISC-V is the exception: an all-zero
// word is the defined UNIMP instruction.
for _, a := range []arch.Arch{arch.ARM64, arch.LOONG64} {
ins, err := Decode(a, []byte{0, 0, 0, 0}, 0)
if err != nil {
t.Fatalf("%s: %v", a, err)
}
if ins.Text != "???" {
t.Errorf("%s: text = %q, want ???", a, ins.Text)
}
}
// The compressed quadrant claims the zero halfword first, so the zero
// word decodes as the 2-byte compressed UNIMP.
if ins, err := Decode(arch.RISCV, []byte{0, 0, 0, 0}, 0); err != nil || ins.Text != "UNIMP" || ins.Len != 2 {
t.Errorf("riscv zero word: %q len %d err %v, want UNIMP with 2 bytes", ins.Text, ins.Len, err)
}
if _, err := Decode(arch.ARM64, []byte{0, 0}, 0); err == nil {
t.Error("short input: expected an error")
}
if _, err := Decode(arch.AMD64, nil, 0); err == nil {
t.Error("empty input: expected an error")
}
}
// TestBlockRoundTrip assembles a small kernel with the gasm encoder for every
// architecture and disassembles it back: the listing must cover the whole
// function and end in RET.
func TestBlockRoundTrip(t *testing.T) {
for _, tt := range []struct {
a arch.Arch
name string
}{
{arch.AMD64, "k_amd64.s"},
{arch.ARM64, "k_arm64.s"},
{arch.RISCV, "k_riscv64.s"},
{arch.LOONG64, "k_loong64.s"},
} {
src := "TEXT \u00b7k(SB), NOSPLIT, $0\n\tMOVQ AX, CX\n\tRET\n"
if tt.a != arch.AMD64 {
src = "TEXT \u00b7k(SB), NOSPLIT, $0\n\tRET\n"
}
f, errs := parser.Parse(tt.name, src)
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", tt.a, errs)
}
img, err := assemble(t, tt.a, f)
if err != nil {
t.Fatalf("%s: assemble: %v", tt.a, err)
}
fn := img.Funcs[0]
code := img.Code[fn.Offset : fn.Offset+fn.Size]
ins := Block(tt.a, code, 0, 100)
if len(ins) == 0 {
t.Fatalf("%s: empty listing", tt.a)
}
consumed := 0
for _, in := range ins {
if in.Text == "" || in.Text == "???" {
t.Errorf("%s: undecoded instruction at %#x: %q", tt.a, in.Addr, in.Text)
}
consumed += in.Len
}
if consumed != len(code) {
t.Errorf("%s: listing consumed %d of %d bytes", tt.a, consumed, len(code))
}
if last := ins[len(ins)-1]; !strings.Contains(strings.ToLower(last.Text), "ret") {
t.Errorf("%s: last instruction = %q, want RET", tt.a, last.Text)
}
}
}
func TestBlockLimits(t *testing.T) {
code := []byte{0x55, 0x55, 0x55, 0x55, 0x55}
if got := Block(arch.AMD64, code, 0, 3); len(got) != 3 {
t.Errorf("max=3 produced %d instructions, want 3", len(got))
}
if got := Block(arch.AMD64, code, 0, 100); len(got) != 5 {
t.Errorf("code end produced %d instructions, want 5", len(got))
}
if got := Block(arch.AMD64, nil, 0, 3); len(got) != 0 {
t.Errorf("empty code produced %d instructions, want 0", len(got))
}
}
// assemble assembles the parsed file with the encoder for a.
func assemble(t *testing.T, a arch.Arch, f *ast.File) (*asm.Image, error) {
t.Helper()
switch a {
case arch.ARM64:
return asm.AssembleFileARM64(f)
case arch.RISCV:
return asm.AssembleFileRISCV(f)
case arch.LOONG64:
return asm.AssembleFileLOONG64(f)
default:
return asm.AssembleFile(f)
}
}
+121 -22
View File
@@ -4,7 +4,9 @@ How gasm-devkit is put together and why.
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit) Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
## Design goals ## Overview
Three design goals shape everything below.
1. **A real AST, not a grammar hack.** The linter, analyser, assembler and 1. **A real AST, not a grammar hack.** The linter, analyser, assembler and
language server all need to *reason* about assembly, not just colour it. language server all need to *reason* about assembly, not just colour it.
@@ -20,10 +22,10 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
through two vendor-neutral interfaces: a CLI and an LSP server. No editor through two vendor-neutral interfaces: a CLI and an LSP server. No editor
owns the toolkit; the toolkit is offered to editors on standard terms. owns the toolkit; the toolkit is offered to editors on standard terms.
## Pipeline The components, and how data moves between them:
```mermaid ```mermaid
graph TD flowchart TD
SRC["source .s"] --> LEX["lexer<br/>token stream"] SRC["source .s"] --> LEX["lexer<br/>token stream"]
LEX --> PAR["parser<br/>AST + diagnostics"] LEX --> PAR["parser<br/>AST + diagnostics"]
LEX --> FMT["format<br/>re-space tokens"] LEX --> FMT["format<br/>re-space tokens"]
@@ -42,9 +44,35 @@ graph TD
The lexer is the shared foundation: the parser builds the AST from it, the The lexer is the shared foundation: the parser builds the AST from it, the
formatter re-spaces its tokens directly, and the language server uses it for formatter re-spaces its tokens directly, and the language server uses it for
semantic highlighting. semantic highlighting. The packages follow a dependency chain: static analysis
builds only on the AST, the standalone assembler emits object code, and both
the dynamic analysis and the debugger consume the execution substrate the
assembler provides.
## Components ## Packages
| Package | Responsibility |
|---|---|
| `token` | token kinds and positions |
| `lexer` | hand-written scanner; permissive, and it never panics |
| `ast` | the typed syntax tree: declarations, lines, operands |
| `parser` | line-oriented parser producing the AST and its diagnostics |
| `arch` | register and instruction tables for the four architectures |
| `lint` | static checks over the AST |
| `format` | canonical formatter over the token stream |
| `lsp` | the language server |
| `asm` | standalone assembler: encoders, image layout, object emitters |
| `disasm` | disassembly backend over golang.org/x/arch |
| `verify` | JIT execution, ABI checks, differential fuzzing |
| `debug` | interactive ptrace debugger |
| `cmd/gasm` | the CLI |
| `_gen` | rebuilds the `arch` tables from the Go toolchain source |
The boundaries matter as much as the responsibilities: `ast` records syntax
only, and whether a name is a register or a label is left to `arch`, so the
parser stays architecture-agnostic. `asm` and `verify` are the only packages
that touch machine code and executable memory, and `cmd/gasm` owns no logic
beyond flags and output.
### `token` and `lexer` ### `token` and `lexer`
@@ -134,7 +162,8 @@ Two deeper analyses sit on top of the AST:
- **`unreachable-code`.** Code after a `RET` and before the next label is - **`unreachable-code`.** Code after a `RET` and before the next label is
dead. The check is suppressed for any function whose reachability cannot be dead. The check is suppressed for any function whose reachability cannot be
decided statically: those using PC-relative jumps (`JMP 2(PC)`), decided statically: those using PC-relative jumps (`JMP 2(PC)`),
register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`, or a `JMP`/`CALL`
through a register or memory operand), or living in a
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator: file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator:
code after it is occasionally intentional metadata. code after it is occasionally intentional metadata.
- **`register-clobber` (register liveness).** The linter builds the function's - **`register-clobber` (register liveness).** The linter builds the function's
@@ -202,20 +231,20 @@ them from the standard LSP legend, so no editor-specific grammar is needed.
### `asm` ### `asm`
The standalone assembler (Phase 2). Its core is an amd64 instruction encoder: The standalone assembler. Its core is an amd64 instruction encoder:
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set, a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
with the Plan 9 operand order (source first) mapped onto the x86 encoding. with the Plan 9 operand order (source first) mapped onto the x86 encoding.
Every encoding is validated by decoding it again with `golang.org/x/arch`, the Every encoding is validated by decoding it again with `golang.org/x/arch`, the
one module dependency, used in tests only and never linked into the binary. one module dependency, which also backs the `gasm dis` listings.
A **RISC-V encoder** (Phase 5, RV64IMAFDC + RVC compression) encodes the full A **RISC-V encoder** (RV64IMAFDC + RVC compression) encodes the full
integer, atomic, float/double, FMA and CSR instruction sets with the MOV integer, atomic, float/double, FMA and CSR instruction sets with the MOV
pseudo-instruction and SB/global symbol references (AUIPC pairs with pseudo-instruction and SB/global symbol references (AUIPC pairs with
R_RISCV_PCREL_HI20/LO12 relocations). The encoder compresses eligible R_RISCV_PCREL_HI20/LO12 relocations). The encoder compresses eligible
instructions to 16-bit RVC forms and is validated byte-for-byte against instructions to 16-bit RVC forms and is validated byte-for-byte against
`GOARCH=riscv64 go tool asm`. `GOARCH=riscv64 go tool asm`.
A **LoongArch encoder** (Phase 5, LoongArch64) encodes the integer and A **LoongArch encoder** (LoongArch64) encodes the integer and
floating-point instruction sets with the dual-form arithmetic mnemonics (3R floating-point instruction sets with the dual-form arithmetic mnemonics (3R
vs 2RI12), the 16/21-bit branch families, the MOV pseudo-instruction and its vs 2RI12), the 16/21-bit branch families, the MOV pseudo-instruction and its
constant materialisation (the dcon classification driving lu12i.w/ori/lu32i.d/ constant materialisation (the dcon classification driving lu12i.w/ori/lu32i.d/
@@ -226,7 +255,7 @@ relocations). Like the RISC-V encoder it is validated byte-for-byte against
`GOARCH=loong64 go tool asm`, and its GOOBJ output is proven end-to-end by `GOARCH=loong64 go tool asm`, and its GOOBJ output is proven end-to-end by
substituting it into a cross-compiled `go build` and linking with `cmd/link`. substituting it into a cross-compiled `go build` and linking with `cmd/link`.
An **AArch64 encoder** (Phase 5, arm64) encodes the integer instruction set An **AArch64 encoder** (arm64) encodes the integer instruction set
with the data-processing (shifted register and immediate forms), load/store with the data-processing (shifted register and immediate forms), load/store
(scaled unsigned immediate and unscaled9-bit immediate), conditional and (scaled unsigned immediate and unscaled9-bit immediate), conditional and
unconditional branches, the MOV pseudo-instruction and its constant unconditional branches, the MOV pseudo-instruction and its constant
@@ -237,6 +266,18 @@ STP+SUB for large frames) and SB/global symbol references (ADRP+ADD pairs with
R_ADDRARM64 relocations). Like the other encoders it is validated R_ADDRARM64 relocations). Like the other encoders it is validated
byte-for-byte against `GOARCH=arm64 go tool asm`. byte-for-byte against `GOARCH=arm64 go tool asm`.
On top of the per-architecture encoders, every framed function carries the
**stack-split guard**: the prologue check against `g.stackguard0` (small,
medium and large frame classes, the medium and large classes materialising
their offset through the architecture's temporary register and the large
class adding the SP-underflow branch) and the trailing morestack block
(save the link register, `CALL runtime.morestack_noctxt`, jump back to the
function entry). The auto-NOSPLIT rule, the frame classes, the large-frame
prologue and epilogue forms and the tail calls match the toolchain's
`stacksplit` and `preprocess` output byte for byte; a parity suite
assembles kernel files with gasm and the installed `go tool asm` and diffs
the bytes on all four architectures.
On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each
operand to an encoder operand, and lays the instructions out so local labels operand to an encoder operand, and lays the instructions out so local labels
resolve to relative jump offsets: jumps start in the short (rel8) form and resolve to relative jump offsets: jumps start in the short (rel8) form and
@@ -349,7 +390,7 @@ compiled packages it references.
### `verify` ### `verify`
The dynamic-analysis substrate (Phase 3). It JIT-loads assembled images into The dynamic-analysis substrate. It JIT-loads assembled images into
executable memory and invokes them directly, enabling differential testing, executable memory and invokes them directly, enabling differential testing,
runtime ABI checks and coverage profiling. runtime ABI checks and coverage profiling.
@@ -369,10 +410,10 @@ every architecture too: `enterJITChecked` plants sentinels in the registers
the Go ABI fixes across calls (amd64 `BP`/`R14`, arm64 `R29`/`R28`, riscv64 the Go ABI fixes across calls (amd64 `BP`/`R14`, arm64 `R29`/`R28`, riscv64
`X27`, loong64 `R22`; the latter two keep no hardware frame pointer) and the `X27`, loong64 `R22`; the latter two keep no hardware frame pointer) and the
raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the
saved registers before Go code resumes. riscv64 is validated end to saved registers before Go code resumes. All three non-amd64 trampolines
end under qemu-user emulation; arm64 shares the same stack convention and are validated end to end under qemu-user emulation, the loong64 one
fix; loong64 stays ground-truth-only until hardware validation (see through its raw-address leave handoff.
docs/DECISIONS.md). `gasm verify` runs the JIT checks when the host `gasm verify` runs the JIT checks when the host
matches the kernel's architecture and the toolchain comparisons matches the kernel's architecture and the toolchain comparisons
elsewhere. elsewhere.
@@ -422,16 +463,74 @@ For non-interactive use, `--script` runs REPL commands from a file (or
stdin) and exits, `--timeout` kills the debuggee when a run hangs (the stdin) and exits, `--timeout` kills the debuggee when a run hangs (the
watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot
block it), and `--cover` runs to completion with a breakpoint on every block it), and `--cover` runs to completion with a breakpoint on every
label and reports which blocks executed. instruction and reports which instructions executed and how often.
## Extension points ### Extending the toolkit
- **New architecture:** add an entry to the generator in `_gen`, run - **New architecture:** add an entry to the generator in `_gen`, run
`just gen`, and add a `buildXXX()` register file plus a case in `ForArch`. `just gen`, and add a `buildXXX()` register file plus a case in `ForArch`.
- **New lint rule:** add a function in `lint` and a rule-code constant. - **New lint rule:** add a function in `lint` and a rule-code constant.
- **New LSP feature:** add a method case in `dispatch` and a handler. - **New LSP feature:** add a method case in `dispatch` and a handler.
The phases follow a dependency chain. Phase 1 (static analysis) builds only on ## Data flow
the AST; Phase 2 (the standalone assembler) emits object code; Phases 3
(dynamic analysis) and 4 (the debugger) both consume the execution substrate The main operation, assembling one file:
that the assembler provides.
```mermaid
sequenceDiagram
participant User
participant CLI as gasm CLI
participant Parser as parser
participant Asm as asm
participant Go as go toolchain
User->>CLI: gasm asm --format goobj -p pkg -o k.o k_amd64.s
CLI->>Parser: Parse(path, src)
Parser-->>CLI: AST, diagnostics
CLI->>Asm: AssembleFile(AST)
Asm->>Asm: encode operands, settle label offsets, lay out data
Asm-->>CLI: Image, code and data and relocations
CLI->>Asm: GOObject(pkg, path)
Asm->>Go: go list -json -export, externals only
Go-->>Asm: package and symbol indices
Asm-->>CLI: Go object bytes
CLI-->>User: wrote N bytes to k.o
```
Errors are produced where the parse or the encoding fails and become values at
the CLI boundary: the parser returns a diagnostic list and never aborts a file,
`AssembleFile` returns an error, and `cmd/gasm` prints what it has to stderr
and returns a non-zero exit code. The formatter and the linter take the same
AST by a different route: `gasm fmt` re-spaces the token stream and `gasm lint`
walks the parsed file, so neither depends on an encoding.
## State and lifetime
- The analysis packages (`lexer`, `parser`, `format`, `lint`, `arch`) hold only
read-only lookup tables and no mutable state: every call allocates its own
tokens and AST, and any number of goroutines may read the `arch` tables.
- A `verify.Kernel` owns one executable mapping, which `Close` releases. The
JIT trampolines keep the Go stack pointer and the checked-call sentinels in
package globals, so a call is a process-wide, one-at-a-time operation. The
`gasm verify` sweeps therefore run each function in a child process, which
contains a crash and keeps the globals unshared.
- `lsp.Server` is long-lived: it runs a single read and dispatch loop over the
stream and touches its document store only from that loop, so one server
serves one connection.
- A `debug.Session` owns a traced child process and pins its goroutine to the
forking OS thread, because ptrace requests must stay on that thread.
## Dependencies
- **`golang.org/x/arch`** (v0.30.0) is the one module dependency: it is the
disassembler backend (`gasm dis` and the debugger's listings) and the source
of the register metadata the encoder consults (`asm/reg.go`, `asm/vex.go`).
The tests additionally decode through it to validate the encodings.
- **The Go toolchain**, as an oracle and never as a library: `go tool asm`
supplies the object preamble and the ground truth for `gasm verify
--ground-truth`, `go list -json -export` locates the archives of the packages
a GOOBJ object references, and `_gen` parses
`$GOROOT/src/cmd/internal/obj/<arch>/anames.go` to rebuild the tables.
- **Linux process interfaces** for the dynamic work: `mmap` and `mprotect` for
the JIT mapping, ptrace with `/proc/pid/mem` for the debugger. That is why
`verify` runs a JIT check only when the host architecture matches the
kernel's, and why `debug` is Linux-only.

Some files were not shown because too many files have changed in this diff Show More