Compare commits

...
25 Commits
Author SHA1 Message Date
petrbalvin 93c47a312a feat(docs): man pages for gasm and every command, guarded against CLI drift
Test / test (push) Successful in 2m4s
Assisted-by: GLM 5.3 Flash
2026-09-19 21:18:43 +02:00
petrbalvin 708d0a0a5e docs: trim the changelog entries to user-visible deltas
Assisted-by: GLM 5.3 Flash
2026-09-19 20:54:56 +02:00
petrbalvin 3c8f7cb411 test(format): pin the fuzz-found crashers as regression seeds
Assisted-by: GLM 5.3 Flash
2026-09-19 20:48:51 +02:00
petrbalvin 7c5b7a1419 docs: add the changelog entries and the corpus number to the readme
Assisted-by: GLM 5.3 Flash
2026-09-19 20:48:51 +02:00
petrbalvin bc3f448738 feat(format): fuzz targets for the parser and formatter
Assisted-by: GLM 5.3 Flash
2026-09-19 20:41:43 +02:00
petrbalvin f37f183577 feat(riscv64): GOROOT instruction shapes, DATA order and offset expressions
Assisted-by: GLM 5.3 Flash
2026-09-19 19:58:43 +02:00
petrbalvin 1e77e58250 feat(gasm): audit a .s corpus with audit-instructions --corpus
Assisted-by: GLM 5.3 Flash
2026-09-19 19:27:30 +02:00
petrbalvin 1d0969ed64 feat(gasm): select the asm and diff architecture with -GOARCH
Assisted-by: GLM 5.3 Flash
2026-09-19 19:20:47 +02:00
petrbalvin 23c001be51 feat(asm): encode indirect JMP and CALL on all four architectures
Assisted-by: GLM 5.3 Flash
2026-09-19 19:17:07 +02:00
petrbalvin 96e81cc98d docs: add the Plan 9 assembly case and real-use note to the README
Test / test (push) Successful in 2m6s
2026-09-19 18:06:18 +02:00
petrbalvin c834d98210 docs: bring the document set into the standard shape
Test / test (push) Successful in 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 03d6d4da54 style: put the repository assembly in gasm fmt canonical form
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 0b42ce7952 style: use one spelling for colour across the CLI
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 288a64ccd2 ci: align the pipelines with the hand-written templates
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin 5fddfa704b build: declare the exact toolchain and the canonical recipes
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:18 +02:00
petrbalvin a2bb5eeb4e chore: drop the stale comment from the ignore list
Assisted-by: GLM 5.3 Flash
2026-09-17 20:33:14 +02:00
petrbalvin 48449b7a7f build: declare the go1.27.1 toolchain
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 3de043c494 docs: add SECURITY.md and record the round in the CHANGELOG
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 0078f7be5c style: purge em dashes from the produced text
Assisted-by: GLM 5.3 Flash
2026-09-16 23:12:31 +02:00
petrbalvin 6a7317d141 chore: trim the ignore list to the convention
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin d08523caa5 docs: move the recipe and version descriptions with the behaviour
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 20e4b8d9c4 ci: align the pipelines with the hand-written templates
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 61f4247cef refactor(gasm): report the toolchain-recorded version
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 049872ddff build: restore the canonical justfile recipe set
Assisted-by: GLM 5.3 Flash
2026-09-16 22:53:01 +02:00
petrbalvin 3669f64ff6 build: install the gasm binary into the user-local bin directory
Test / vet (push) Successful in 46s
Test / test (push) Successful in 2m44s
Test / build (push) Successful in 42s
2026-09-14 23:41:21 +02:00
115 changed files with 4463 additions and 1565 deletions
+37
View File
@@ -0,0 +1,37 @@
# Race, Go. Dispatched by hand, and never a gate on a push or a tag: the release tag is
# cut only after `just gates` has already raced the tree, so this workflow is the
# explicit second opinion, not a step of the release.
#
# The race detector roughly doubles both time and memory, which the shared runner box
# cannot afford on every push. Locally it belongs to `just gates`, which runs it once per
# task; here it is a decision rather than a routine.
#
# Every step is one command, so the step that fails is the gate that failed.
name: Race
on:
workflow_dispatch:
env:
# One core: parallelism buys no speed here and costs memory the box does not have.
GOFLAGS: -p=1
GOMAXPROCS: "2"
jobs:
race:
runs-on: fedora
timeout-minutes: 20
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version-file: go.mod
cache: true
- name: Install gcc
# The race detector needs cgo and the runner image carries no C compiler.
run: dnf install -y gcc
- name: Race
run: go test -race -count=1 -timeout 10m ./...
+255 -86
View File
@@ -1,73 +1,198 @@
# Release — gasm binaries. Runs on version tags (v0.28.0) pushed to main. # Release, Go binaries. Runs on version tags (v1.2.3) pushed to main.
#
# The module sits at the repository root: the toolchain records a version only for a root
# module, measured on go1.27.1, so a build of a module in a subdirectory reports (devel)
# even at its own <module>/vX.Y.Z tag and this workflow's smoke test can never pass for
# it. A Go repository is one module at the root.
#
# The version contract these steps implement is in the `release` skill, and its point is
# that nothing is injected: the toolchain records the tag into the binary's build
# information, so the build simply has to happen at the tag, which the trigger guarantees.
#
# The gates run in their own job, once, before the matrix, minus the race detector: race
# never runs on a push path or a tag, and the local gate raced this tree before the tag
# was cut. Putting the gates inside the matrix would run the whole suite once per target
# on the box that also hosts the forge. Each job validates the tag for itself rather than
# passing a value between jobs, so no workflow feature has to be trusted for the version
# to reach the file name.
name: Release name: Release
on: on:
push: push:
tags: ["v*"] tags: ["v*"]
env:
# The box is shared with the forge, so parallelism is bounded on purpose. The gates job
# needs it most; the build jobs inherit it for their parallel compilation.
GOFLAGS: -p=1
GOMAXPROCS: "2"
jobs: jobs:
gates:
runs-on: fedora
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version-file: go.mod
cache: true
- name: Install Perl
# Perl for the steps below. The install is a no-op where the package
# is already present.
run: dnf install -y perl
- name: Validate the tag
env:
VERSION: ${{ gitea.ref_name }}
run: |
perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ m{^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$}
or die qq{ERROR: expected a semver tag like v1.2.3, got: $v\n};
print qq{tag $v\n};
'
- name: Build
run: go build ./...
- name: Format
run: |
perl -e '
open(my $g, q{-|}, q{gofmt}, q{-l}, q{.}) or die qq{gofmt: $!};
my @bad = <$g>;
close($g);
print @bad;
exit(@bad ? 1 : 0);
'
- name: Vet
run: go vet ./...
- name: Modernise
run: go fix -diff ./...
- name: Tests
# The same command as in test.yml, so the floor is the same number everywhere.
run: go test -count=1 -timeout 10m -coverprofile=coverage.out ./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... ./lint/... ./lsp/... ./parser/... ./token/... ./verify/...
- name: Coverage floor
run: |
perl -e '
open(my $c, q{-|}, q{go}, q{tool}, q{cover}, q{-func=coverage.out}) or die qq{cover: $!};
my $total;
while (my $l = <$c>) { $total = $1 if $l =~ m{^total:\s+\S+\s+([0-9.]+)%} }
close($c);
die qq{no total line in coverage.out\n} unless defined $total;
printf qq{Total coverage: %s%%\n}, $total;
exit($total < 80 ? 1 : 0);
'
build: build:
runs-on: fedora runs-on: fedora
timeout-minutes: 25
needs: gates
strategy: strategy:
fail-fast: false fail-fast: false
matrix: matrix:
# Portable targets: amd64, arm64, loong64 and riscv64 on Linux, at the toolchain
# default level. No 32-bit, no wasm, no macOS, no Windows. FreeBSD stays out until
# verify/jit.go ports off syscall.Mprotect: the Go syscall package defines no
# Mprotect for freebsd, and verify/jit.go:50 calls it to drop the write bit from
# the JIT mapping, so every freebsd target fails to build with "undefined:
# syscall.Mprotect" (verified for amd64, arm64 and riscv64 on go1.27.1).
include: include:
- goos: linux - goos: linux
goarch: amd64 goarch: amd64
- goos: linux - goos: linux
goarch: arm64 goarch: arm64
- goos: linux
goarch: riscv64
- goos: linux - goos: linux
goarch: loong64 goarch: loong64
- goos: linux
goarch: riscv64
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
- uses: actions/setup-go@v6 - uses: actions/setup-go@v6
with: with:
go-version: "1.27" go-version-file: go.mod
cache: true
- name: Download dependencies - name: Install Perl
run: go mod download run: dnf install -y perl
- name: Validate tag and build - name: Validate the tag
id: build id: version
env: env:
VERSION: ${{ gitea.ref_name }} VERSION: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ m{^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$}
or die qq{ERROR: expected a semver tag like v1.2.3, got: $v\n};
(my $nv = $v) =~ s{^v}{};
open(my $o, q{>>}, $ENV{GITEA_OUTPUT}) or die qq{GITEA_OUTPUT: $!};
print $o qq{version_no_v=$nv\n};
close($o);
print qq{version $nv\n};
'
if ! echo "$VERSION" | grep -qE '^v[0-9]+(\.[0-9]+){0,2}([-+].*)?$'; then - name: Build
echo "ERROR: expected a semver tag like v1.2.3, got: '$VERSION'" env:
exit 1 VERSION_NO_V: ${{ steps.version.outputs.version_no_v }}
fi GOOS: ${{ matrix.goos }}
GOARCH: ${{ matrix.goarch }}
VERSION_NO_V="${VERSION#v}" CGO_ENABLED: "0"
echo "version_no_v=${VERSION_NO_V}" >> "$GITEA_OUTPUT" run: |
# Nothing is injected. The toolchain records the tag into the binary's build
mkdir -p bin # information, so the version is right because this build happens at the tag, and
GOOS=${{ matrix.goos }} GOARCH=${{ matrix.goarch }} CGO_ENABLED=0 \ # there is no path for anyone to get wrong. -s -w only strips symbols.
go build -ldflags "-s -w -X main.version=${VERSION_NO_V}" \ go build -ldflags "-s -w" -o "bin/gasm-${VERSION_NO_V}-${GOOS}-${GOARCH}" ./cmd/gasm
-o "bin/gasm-${VERSION_NO_V}-${{ matrix.goos }}-${{ matrix.goarch }}" \
./cmd/gasm
# Artifacts stay on v3: v4 and later detect Gitea as GHES and abort.
- name: Upload artifact - name: Upload artifact
uses: actions/upload-artifact@v3 uses: actions/upload-artifact@v3
with: with:
name: gasm-${{ matrix.goos }}-${{ matrix.goarch }} name: gasm-${{ matrix.goos }}-${{ matrix.goarch }}
path: bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} path: bin/gasm-${{ steps.version.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }}
if-no-files-found: error if-no-files-found: error
- name: Smoke test - name: Smoke test
# Only a binary matching the runner can be run here. The check is not that --version
# exits cleanly but that it reports the tag and nothing more: a build outside version
# control reports (devel), and a build whose tree was dirty reports +dirty, and both
# would otherwise be published.
if: matrix.goos == 'linux' && matrix.goarch == 'amd64' if: matrix.goos == 'linux' && matrix.goarch == 'amd64'
env:
TAG: ${{ gitea.ref_name }}
BIN: bin/gasm-${{ steps.version.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }}
run: | run: |
chmod +x bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} perl -e '
./bin/gasm-${{ steps.build.outputs.version_no_v }}-${{ matrix.goos }}-${{ matrix.goarch }} --version my $want = $ENV{TAG} // die qq{ERROR: no tag\n};
open(my $bin, q{-|}, $ENV{BIN}, q{--version}) or die qq{$ENV{BIN}: $!};
my $got = <$bin>;
close($bin);
$got = defined $got ? $got : q{};
chomp $got;
index($got, $want) >= 0
or die qq{ERROR: the binary printed "$got", which does not contain $want. Version control was disabled, so there is no recorded version.\n};
index($got, q{+dirty}) < 0
or die qq{ERROR: the binary printed "$got". The tree was dirty at build time, which means the checkout was not the tag, or the build artefacts are not ignored.\n};
print qq{$ENV{BIN} reports $got\n};
'
release: release:
runs-on: fedora runs-on: fedora
timeout-minutes: 15
needs: build needs: build
permissions: permissions:
# contents: read is required for the checkout: a job that declares any
# permissions gets a token scoped to exactly those, and releases: write
# alone leaves the fetch with no read access, which Gitea answers with
# a 404 "Repository not found". Verified on the instance 2026-09-16.
contents: read
releases: write releases: write
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
@@ -77,81 +202,125 @@ jobs:
with: with:
path: dist path: dist
- name: Extract CHANGELOG section - name: Install Perl
run: dnf install -y perl
- name: Extract the CHANGELOG section
env: env:
VERSION: ${{ gitea.ref_name }} VERSION: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail # Each step derives what it needs from the tag, so no value has to travel between
VERSION_NO_V="${VERSION#v}" # jobs.
perl -e '
my $v = $ENV{VERSION} // q{};
$v =~ s{^v}{};
open(my $vout, q{>}, q{version-no-v.txt}) or die qq{version-no-v.txt: $!};
print $vout $v;
close($vout);
open(my $in, q{<}, q{CHANGELOG.md}) or die qq{CHANGELOG.md: $!};
my @lines = <$in>;
close($in);
my ($start, $end) = (-1, scalar @lines);
for my $i (0 .. $#lines) {
if ($start < 0) { $start = $i if $lines[$i] =~ m{^##\s+\[\Q$v\E\]} }
elsif ($lines[$i] =~ m{^##\s+\[}) { $end = $i; last }
}
$start >= 0 or die qq{ERROR: no CHANGELOG section for $v, expected a heading like: ## [$v] - YYYY-MM-DD\n};
my @body = grep { m{\S} } @lines[$start + 1 .. $end - 1];
@body or die qq{ERROR: the CHANGELOG section for $v is empty\n};
open(my $out, q{>}, q{release-body.md}) or die qq{release-body.md: $!};
print $out @body;
close($out);
printf qq{notes for %s: %d lines\n}, $v, scalar @body;
'
sed -n "/^## \[${VERSION_NO_V}\] /,/^## \[/p" CHANGELOG.md \ - name: Build the release request
| sed '$d' \ run: |
| tail -n +2 \ perl -e '
> release-body.md open(my $vin, q{<}, q{version-no-v.txt}) or die qq{version-no-v.txt: $!};
my $v = <$vin>;
close($vin);
chomp $v;
open(my $in, q{<:raw}, q{release-body.md}) or die qq{release-body.md: $!};
my $body = do { local $/; <$in> };
close($in);
# Byte-oriented escaping: JSON is UTF-8, so non-ASCII passes through and only the
# characters JSON forbids are rewritten.
$body =~ s/([\\"])/\\$1/g;
$body =~ s/\t/\\t/g;
$body =~ s/\r//g;
$body =~ s/\n/\\n/g;
$body =~ s/([\x00-\x08\x0b\x0c\x0e-\x1f])/sprintf(q{\u%04x}, ord($1))/ge;
my $json = sprintf(qq{{"tag_name":"v%s","name":"v%s","body":"%s","draft":false,"prerelease":false}}, $v, $v, $body);
open(my $out, q{>}, q{release.json}) or die qq{release.json: $!};
print $out $json;
close($out);
print qq{release.json written for v$v\n};
'
if [ ! -s release-body.md ]; then - name: Create the release
echo "ERROR: no CHANGELOG section found for ${VERSION_NO_V}"
echo "Expected a heading like: ## [${VERSION_NO_V}] — YYYY-MM-DD"
exit 1
fi
- name: Create release
env: env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }} GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
GITEA_SERVER_URL: ${{ gitea.server_url }} GITEA_SERVER_URL: ${{ gitea.server_url }}
GITEA_REPOSITORY: ${{ gitea.repository }} GITEA_REPOSITORY: ${{ gitea.repository }}
GITEA_REF_NAME: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
my @cmd = (q{curl}, q{-sS}, q{-o}, q{response.json}, q{-w}, q{%{http_code}},
BODY=$(sed -e 's/\\/\\\\/g' -e 's/"/\\"/g' -e 's/\t/\\t/g' -e 's/\r//g' release-body.md | sed ':a;N;$!ba;s/\n/\\n/g') q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
BODY="\"${BODY}\"" q{-H}, q{Content-Type: application/json},
q{-X}, q{POST},
response=$(curl -sS -w '\n%{http_code}' \ qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases},
-H "Authorization: token ${GITEA_TOKEN}" \ q{--data-binary}, q{@release.json});
-H "Content-Type: application/json" \ open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
-X POST \ my $code = <$curl>;
"${GITEA_SERVER_URL}/api/v1/repos/${GITEA_REPOSITORY}/releases" \ my $ok = close($curl);
-d "{\"tag_name\":\"${GITEA_REF_NAME}\",\"name\":\"${GITEA_REF_NAME}\",\"body\":${BODY},\"draft\":false,\"prerelease\":false}") my $exit = $? >> 8;
$code = defined $code ? $code : q{};
http_code=$(echo "$response" | tail -1) $ok or die qq{ERROR: curl failed (exit $exit) calling $ENV{GITEA_SERVER_URL}\n};
payload=$(echo "$response" | sed '$d') open(my $r, q{<:raw}, q{response.json}) or die qq{response.json: $!};
my $body = do { local $/; <$r> };
echo "HTTP ${http_code}" close($r);
if [ "$http_code" != "201" ]; then $code eq q{201} or die qq{ERROR: the release was not created, HTTP $code: $body\n};
echo "Failed to create release: ${payload}" $body =~ m{"id"\s*:\s*([0-9]+)} or die qq{ERROR: no release id in the response: $body\n};
exit 1 open(my $o, q{>}, q{release-id.txt}) or die qq{release-id.txt: $!};
fi print $o $1;
close($o);
RELEASE_ID=$(echo "$payload" | grep -oE '"id"[[:space:]]*:[[:space:]]*[0-9]+' | head -1 | grep -oE '[0-9]+') print qq{release id $1\n};
echo "Created release ID=${RELEASE_ID}" '
printf '%s' "${RELEASE_ID}" > release-id.txt
- name: Upload assets - name: Upload assets
env: env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }} GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
GITEA_SERVER_URL: ${{ gitea.server_url }} GITEA_SERVER_URL: ${{ gitea.server_url }}
GITEA_REPOSITORY: ${{ gitea.repository }} GITEA_REPOSITORY: ${{ gitea.repository }}
GITEA_REF_NAME: ${{ gitea.ref_name }}
run: | run: |
set -euo pipefail perl -e '
RELEASE_ID=$(cat release-id.txt) open(my $f, q{<}, q{release-id.txt}) or die qq{release-id.txt: $!};
my $id = <$f>;
for binary in dist/gasm-*/gasm-*; do close($f);
[ -f "$binary" ] || continue chomp $id;
fname=$(basename "$binary") my @files = grep { -f $_ } glob(q{dist/*/*});
echo "Uploading ${fname}..." @files or die qq{ERROR: no assets under dist/\n};
http_code=$(curl -sS -o /dev/null -w '%{http_code}' \ my $bad = 0;
-H "Authorization: token ${GITEA_TOKEN}" \ for my $path (@files) {
-H "Content-Type: application/octet-stream" \ (my $name = $path) =~ s{.*/}{};
-X POST \ my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
--data-binary "@${binary}" \ q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
"${GITEA_SERVER_URL}/api/v1/repos/${GITEA_REPOSITORY}/releases/${RELEASE_ID}/assets?name=${fname}") q{-H}, q{Content-Type: application/octet-stream},
echo " HTTP ${http_code}" q{-X}, q{POST}, q{--data-binary}, qq{@$path},
if [ "$http_code" != "201" ]; then qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
echo "Failed to upload ${fname}" open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
exit 1 my $code = <$curl>;
fi my $ok = close($curl);
done my $exit = $? >> 8;
$code = defined $code ? $code : q{};
echo "Release ${GITEA_REF_NAME} is live." unless ($ok) {
printf qq{%s: curl failed (exit %d)\n}, $name, $exit;
$bad = 1;
next;
}
printf qq{%s: HTTP %s\n}, $name, $code;
$bad = 1 if $code ne q{201};
}
exit($bad ? 1 : 0);
'
+82 -76
View File
@@ -1,4 +1,17 @@
# Test — gasm-devkit. Runs on push and pull request to development. # Test, Go. Push and pull request to development. Never on main.
#
# The gates are the ones the justfile's `gates` recipe runs, minus race: the shared
# runner box cannot afford the race detector on every push, so it lives in race.yml.
# The box is one core and 2 GB beside Gitea, so parallelism is bounded on purpose and
# everything runs in one job. Extra jobs would duplicate the checkout, the Go setup and
# the dependency download three times without buying any parallelism.
#
# Every step is one command, so the step that fails is the gate that failed, and no shell
# option has to be trusted for the run to stop. The scripted steps are Perl, not shell and
# not Python: Perl behaves the same on both runner images, there is no bashism to trip over
# on ash, and it is one language instead of two. The Perl uses builtins only, because
# Fedora packages the Perl modules separately and nothing beyond `perl` itself may be
# assumed present.
name: Test name: Test
on: on:
@@ -7,90 +20,83 @@ on:
pull_request: pull_request:
branches: [development] branches: [development]
env:
# One core: parallelism buys no speed here and costs memory the box does not have.
GOFLAGS: -p=1
GOMAXPROCS: "2"
# A superseded run of the same ref is cancelled instead of queueing behind one that
# no longer matters. Verified on Gitea 1.27.1 on 2026-09-17: a queued run whose ref
# moved on is cancelled before it ever reaches the runner, while a run already
# dispatched there runs to completion.
concurrency:
group: ${{ gitea.workflow }}-${{ gitea.ref }}
cancel-in-progress: true
jobs: jobs:
vet:
runs-on: fedora
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version: "1.27"
- name: Download dependencies
run: go mod download
- name: gofmt
run: |
set -euo pipefail
unformatted=$(gofmt -l .)
if [ -n "$unformatted" ]; then
echo "These files need gofmt:"
echo "$unformatted"
exit 1
fi
- name: go vet
run: go vet ./...
test: test:
runs-on: fedora runs-on: fedora
needs: vet timeout-minutes: 10
steps: steps:
- uses: actions/checkout@v7 - uses: actions/checkout@v7
- uses: actions/setup-go@v6 - uses: actions/setup-go@v6
with: with:
go-version: "1.27" # The module is the source of truth for the version, so it cannot drift.
go-version-file: go.mod
cache: true
- name: Download dependencies - name: Install Perl
run: go mod download # The runner images are minimal and Perl is not guaranteed. The install is a
# no-op where it is already present; drop this step once verified on the box.
- name: Install gcc run: dnf install -y perl
run: dnf install -y gcc
- name: go test -race
run: go test -race -count=1 ./...
- name: Coverage gate — 80 % minimum
run: |
set -euo pipefail
# Exclude packages inherently untestable without hardware:
# debug — interactive ptrace, requires a live process
# cmd/gasm — CLI glue, covered by integration tests
go test -coverprofile=coverage.out \
sourcedock.dev/petrbalvin/gasm-devkit/arch \
sourcedock.dev/petrbalvin/gasm-devkit/asm \
sourcedock.dev/petrbalvin/gasm-devkit/ast \
sourcedock.dev/petrbalvin/gasm-devkit/format \
sourcedock.dev/petrbalvin/gasm-devkit/lexer \
sourcedock.dev/petrbalvin/gasm-devkit/lint \
sourcedock.dev/petrbalvin/gasm-devkit/lsp \
sourcedock.dev/petrbalvin/gasm-devkit/parser \
sourcedock.dev/petrbalvin/gasm-devkit/token \
sourcedock.dev/petrbalvin/gasm-devkit/verify
coverage=$(go tool cover -func=coverage.out | awk '/^total:/ { gsub("%", "", $3); print $3 }')
echo "Total coverage: ${coverage}%"
if awk -v c="$coverage" 'BEGIN { exit !(c+0 < 80) }'; then
echo "ERROR: coverage ${coverage}% is below the 80% threshold"
exit 1
fi
build:
runs-on: fedora
needs: test
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
go-version: "1.27"
- name: Download dependencies
run: go mod download
# The steps follow the `gates` order of the justfile contract: build, format,
# vet, test. The vet gate is go vet and go fix -diff, two steps here.
- name: Build - name: Build
run: go build -ldflags="-s -w" -o bin/gasm ./cmd/gasm run: go build ./...
- name: Smoke test - name: Format
run: ./bin/gasm --version run: |
perl -e '
open(my $g, q{-|}, q{gofmt}, q{-l}, q{.}) or die qq{gofmt: $!};
my @bad = <$g>;
close($g);
print @bad;
exit(@bad ? 1 : 0);
'
- name: Vet
run: go vet ./...
- name: Modernise
# Exits non-zero when it has something to rewrite, so it needs no output capture.
run: go fix -diff ./...
- name: Tests
# The suite must be fast: a push pipeline that cannot finish in a few minutes moves
# its heavy part behind a dispatch. The inner timeout matches the job's, so a
# hanging test reports its own goroutine dump rather than a silent job kill.
# The pattern is `packages` in the project's justfile: the logic packages, since a
# thin cmd/ would drag the total under the floor. release.yml runs the same
# command, so the floor is the same number everywhere. ./verify/... carries the
# live oracle-parity comparison against `go tool asm` (the TestGroundTruth
# suites); the runner's Go setup provides both the tool and GOROOT.
run: go test -count=1 -timeout 10m -coverprofile=coverage.out ./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... ./lint/... ./lsp/... ./parser/... ./token/... ./verify/...
- name: Oracle parity
# Re-run the live go-tool-asm comparison as its own step so that a parity
# regression names the gate that failed instead of hiding inside the suite.
run: go test -count=1 -timeout 10m -run 'TestGroundTruth' ./verify/...
- name: Coverage floor
run: |
perl -e '
open(my $c, q{-|}, q{go}, q{tool}, q{cover}, q{-func=coverage.out}) or die qq{cover: $!};
my $total;
while (my $l = <$c>) { $total = $1 if $l =~ m{^total:\s+\S+\s+([0-9.]+)%} }
close($c);
die qq{no total line in coverage.out\n} unless defined $total;
printf qq{Total coverage: %s%%\n}, $total;
exit($total < 80 ? 1 : 0);
'
+3 -15
View File
@@ -1,25 +1,13 @@
# Metadata (always first, per repo convention)
.idea/ .idea/
.zcode/ .zcode/
.qwen/
.mimocode/
# Binaries # Build output
/gasm
/bin/ /bin/
*.exe /gasm
# Test and coverage artefacts
coverage.out coverage.out
*.test *.test
# Crash dumps # Crash dumps from the emulator runs
core core
core.* core.*
*.core *.core
# Scratch / temporary work
_scratch/
# ZCode workspace
.zcode
+239 -140
View File
@@ -3,15 +3,114 @@
All notable changes to gasm-devkit are documented here. All notable changes to gasm-devkit are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Conventional Commits](https://www.conventionalcommits.org/). and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [development] ## [development]
### Added ### Added
- - **Indirect JMP and CALL on all four architectures.** `JMP AX`,
`CALL AX`, `JMP (BX)` and the memory forms encode at byte parity with
the toolchain (FF /2 and FF /4 on amd64); arm64 lowers `JMP (R0)` to
BR and accepts the raw BR/BLR spellings; riscv64 lowers `JMP (X5)` to
JALR; loong64 accepts the raw `JIRL rd, rj, off` spelling the Go
assembler cannot express. A frameless amd64 function containing a
CALL now receives the toolchain's forced base-pointer frame. The
verify trampolines join the ground-truth lists, and a lint check for
control flow through registers and memory extends to the new forms.
- **`gasm asm -GOARCH` and `gasm diff -GOARCH`.** The target
architecture can be named explicitly instead of inferred from the
file-name suffix, which is how the suffix-less majority of GOROOT's
`.s` files (cpu_x86.s, stub.s, ...) become assemblable.
- **`gasm audit-instructions --corpus [dir]`.** Assembles every `.s`
file under a directory (default GOROOT/src) with the gasm encoder
only: suffixed files for their architecture, suffix-less files for
all four, as a GOARCH build would. Reports the headline number (108
of 627 GOROOT files, 17.2 %, assemble for every target architecture,
against 23 in the previous release), the per-architecture pass rates
and the most common failure reasons with a representative file each,
which drive the encodability backlog by frequency.
- **Fuzz targets for the parser and the formatter.** FuzzParse (no
panic, always a usable file) and FuzzFormatIdempotency (formatting
twice equals formatting once; clean input stays clean) seed
themselves from the repository's kernels, so the plain test suite
replays every seed in CI and `just fuzz` runs the mutation engine on
demand.
- **Oracle parity as its own CI step.** The push pipeline already ran
the live go-tool-asm comparison inside the suite; a dedicated step
now names that gate when it fails.
- **Man pages.** docs/man carries gasm(1) and one page per command,
written in roff: synopsis, description, every flag with its default,
exit status, worked examples and cross-references.
`just install-man` compresses them into ~/.local/share/man (MANDIR
overrides) and `just uninstall-man` removes them. A test builds the
binary and compares every command's `-h` output with its page, so the
pages cannot drift from the CLI.
## [0.33.0] — 2026-09-14 ### Changed
- **Canonical just recipes.** `just gates` is the definition of done
(build, fmt-check, vet, test, race). `install` now builds and copies
the binary into `~/.local/bin` (`BINDIR` overrides) instead of
downloading module dependencies, and `install-bin` is gone. The test
gate sweeps the logic packages (arch through verify; the hardware-bound
`debug` and the thin `cmd/gasm` sit outside it), so the coverage floor
is computed over the product code and the number is identical locally
and in CI. `fuzz` requires its target package.
- **The reported version comes from the build.** `gasm --version`
prints the version the toolchain recorded: the tag on a tagged
checkout, a pseudo-version naming the commit below one, `+dirty` on a
dirty tree and `(devel)` outside version control. Nothing is
injected with `-ldflags -X` any more.
- **CI realigned with the gate set.** The push pipeline runs the gates
minus race in one job, in the `gates` order, with a cached Go setup and
the module as the version source; a superseded run of the same branch
is cancelled instead of queueing; every `go test` runs under a
ten-minute bound that matches its job's; the race detector moved to a
hand-dispatched workflow and runs in the local gate before a tag is
cut, never on a push or a tag; the release builds without injection and
its smoke test requires the recorded tag and rejects `+dirty`.
- **The documents follow the standard set.** `docs/ARCHITECTURE.md` is
organised as Overview, Packages, Data flow, State and lifetime and
Dependencies, and carries a sequence diagram of the assembly path;
`docs/DEVELOPMENT.md` lists every recipe in one table and documents the
coverage floor, the CI and the release flow; `docs/CLI.md` gives the
synopsis, the commands, every flag with its default, the exit codes and
worked examples; `CONTRIBUTING.md` carries the Contributor terms and
states the commit trailer form, the one-logical-change rule and the
licence header rule. The repository's own assembly (the `verify`
trampolines and the test kernels) is in `gasm fmt` canonical form.
- **The README states the project's purpose and status.** It opens with
a warning that the tool is an experiment under active development,
version 0.x.x, free to change without warning, with 1.0.0 far off,
and already in active use on real assembly work. It describes both
goals (tooling for Plan 9 assembly, and Plan 9 assembly outside the
Go toolchain), argues the case for the syntax in a new Why Plan 9
assembly section, and carries a Direction section: extended
instruction support, full GOOBJ and ELF compilation, Linux and
FreeBSD, and the four architectures.
### Fixed
- **riscv64 JALR silently jumped to the wrong register.** The trampoline
form `JALR X0, 0(X5)` read the memory operand's base as the destination,
encoding a jump to X0 with no diagnostic; the destination is the first
operand. The leaf detection shared the confusion, so affected functions
also grew a bogus prologue. `JALR X0, 0(X1)` as written in the verify
trampoline was mis-encoded since its introduction.
- **DATA lines demanded their GLOBL first.** collectData processed the
declarations in file order, but the Plan 9 convention puts every DATA
line before its symbol's GLOBL; correctly ordered files (most of
GOROOT's) failed with "no matching GLOBL". Two passes: symbols are
registered before initialisers are applied.
- **The formatter lost idempotency on degenerate lines.** Illegal tokens
survived into the output, a label sharing its line with a
non-instruction split into a line the parser rejects, stray-operand
lines entered the alignment width computation, and rendered `/ *`,
`> >` sequences re-lexed as comments and shifts. The label, width and
spacing rules now agree between passes.
## [0.33.0] - 2026-09-14
### Added ### Added
@@ -67,7 +166,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
the toolchain picks, and the morestack block saves the link register the toolchain picks, and the morestack block saves the link register
with the toolchain's `OR` form on loong64. with the toolchain's `OR` form on loong64.
## [0.32.0] — 2026-08-31 ## [0.32.0] - 2026-08-31
### Added ### Added
@@ -218,7 +317,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
outputs and operand strictness with `go tool asm`. outputs and operand strictness with `go tool asm`.
- **asm help text.** Updated to list arm64 as a supported architecture. - **asm help text.** Updated to list arm64 as a supported architecture.
## [0.31.1] — 2026-08-20 ## [0.31.1] - 2026-08-20
### Fixed ### Fixed
@@ -226,9 +325,9 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
because the version variables in `justfile` and `cmd/gasm/main.go` were not because the version variables in `justfile` and `cmd/gasm/main.go` were not
bumped during the release commit. bumped during the release commit.
## [0.31.0] — 2026-08-20 ## [0.31.0] - 2026-08-20
The arm64 encoder (Phase 5 — complete) ships with ELF64 and GOOBJ emission, The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission,
verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a
real `go build`. The encoder covers the full integer instruction set, FP real `go build`. The encoder covers the full integer instruction set, FP
arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with
@@ -236,7 +335,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
### Added ### Added
- **arm64 encoder (Phase 5 — complete).** `gasm asm` can now assemble `_arm64.s` - **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s`
files: the AArch64 integer instruction set with the MOV pseudo-instruction and files: the AArch64 integer instruction set with the MOV pseudo-instruction and
its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with
logical bitmask encoding for values like `$1`), data-processing (shifted logical bitmask encoding for values like `$1`), data-processing (shifted
@@ -245,7 +344,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations),
jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth
verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5
(the other architectures — RISC-V, LoongArch, arm64) is now complete. (the other architectures; RISC-V, LoongArch, arm64) is now complete.
### Changed ### Changed
@@ -253,7 +352,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward
compatibility. compatibility.
## [0.30.0] — 2026-08-13 ## [0.30.0] - 2026-08-13
The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified
byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real
@@ -278,7 +377,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
- **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function - **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function
DWARF symbols the linker requires (the subprogram DIE and the `.debug_line` DWARF symbols the linker requires (the subprogram DIE and the `.debug_line`
program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are
in the architecture's MinLC units as the runtime expects — the amd64 link in the architecture's MinLC units as the runtime expects; the amd64 link
test now genuinely substitutes the gasm object, and the amd64/loong64 test now genuinely substitutes the gasm object, and the amd64/loong64
end-to-end GOOBJ link tests pass. end-to-end GOOBJ link tests pass.
- **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is - **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is
@@ -347,7 +446,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
`GOARCH=riscv64 go tool asm`. `GOARCH=riscv64 go tool asm`.
- **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used - **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used
hardware watchpoint slot 0, so a second `watch` call silently overwrote hardware watchpoint slot 0, so a second `watch` call silently overwrote
the first. Watchpoint slots are now tracked in the `Session` (DR0–DR3); the first. Watchpoint slots are now tracked in the `Session` (DR0-DR3);
`watch` picks the first free slot and reports an error if all four are in `watch` picks the first free slot and reports an error if all four are in
use, and `unwatch <slot>` clears one (no argument clears all). use, and `unwatch <slot>` clears one (no argument clears all).
@@ -359,7 +458,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
disassembly at PC, memory-write, watchpoints, and source-line mapping are disassembly at PC, memory-write, watchpoints, and source-line mapping are
all shipped. all shipped.
## [0.29.0] — 2026-08-07 ## [0.29.0] - 2026-08-07
RISC-V GOOBJ emission, YMM vector register display, named buffer allocation RISC-V GOOBJ emission, YMM vector register display, named buffer allocation
in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in
@@ -369,30 +468,30 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Added ### Added
- **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces - **RISC-V GOOBJ emission**; `gasm asm --format goobj` for RISC-V produces
linkable Go objects with funcdata, pc-value tables, and RISC-V relocation linkable Go objects with funcdata, pc-value tables, and RISC-V relocation
types (same format as amd64 GOOBJ, with the RISC-V architecture marker). types (same format as amd64 GOOBJ, with the RISC-V architecture marker).
- **`gasm diff`** — compare the machine code of two assembly files byte-for-byte; - **`gasm diff`**; compare the machine code of two assembly files byte-for-byte;
shows which functions differ and the first few differing bytes. shows which functions differ and the first few differing bytes.
- **`gasm profile`** — show the basic-block structure of each function: labels, - **`gasm profile`**; show the basic-block structure of each function: labels,
offsets, frame size, and NOSPLIT flag. offsets, frame size, and NOSPLIT flag.
- **LSP go-to-definition** — `textDocument/definition` navigates from a label - **LSP go-to-definition**; `textDocument/definition` navigates from a label
reference to its definition. reference to its definition.
- **did-you-mean** — when the RISC-V assembler encounters an undefined label, it - **did-you-mean**; when the RISC-V assembler encounters an undefined label, it
suggests the closest existing label using Levenshtein distance. suggests the closest existing label using Levenshtein distance.
- **YMM vector register display** — `regs` in the debugger now shows YMM - **YMM vector register display**; `regs` in the debugger now shows YMM
registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable). registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable).
- **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates - **Named buffer allocation**; `gasm debug --buf name:size:pattern` allocates
buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern; buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern;
buffer pointers are placed into the argument block at the matching positions. buffer pointers are placed into the argument block at the matching positions.
- **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused - **Crash input storage**; `FuzzResult.CrashInput` stores the input that caused
a crash or mismatch for reproducibility. a crash or mismatch for reproducibility.
- **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel - **ABI + fuzz combined**; `gasm verify --fuzz` now runs ABI checks (sentinel
registers, canary, stack bounds) alongside differential fuzz testing. registers, canary, stack bounds) alongside differential fuzz testing.
- **`gasm diff --map`** — compare functions whose names differ between files - **`gasm diff --map`**; compare functions whose names differ between files
(e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless (e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless
of suffix). Unmapped functions fall back to the original name match. of suffix). Unmapped functions fall back to the original name match.
- **`gasm verify --call`** — invoke a single function with user-supplied buffers - **`gasm verify --call`**; invoke a single function with user-supplied buffers
(`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns: (`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns:
`zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g. `zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g.
decoders) that crash on random input but should succeed on valid data. decoders) that crash on random input but should succeed on valid data.
@@ -402,23 +501,23 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Fixed ### Fixed
- **Signature parser** — grouped Go parameters like `dst, src []byte` are now - **Signature parser**; grouped Go parameters like `dst, src []byte` are now
parsed correctly (both get type `[]byte`). Previously the first name was parsed correctly (both get type `[]byte`). Previously the first name was
treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block
layout in both `verify --call` and the fuzzer. layout in both `verify --call` and the fuzzer.
- **Flaky JIT tests** — `runtime.KeepAlive` guards and package-level buffers - **Flaky JIT tests**; `runtime.KeepAlive` guards and package-level buffers
prevent GC from collecting heap objects whose addresses were passed to JIT prevent GC from collecting heap objects whose addresses were passed to JIT
code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`. code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`.
### Changed ### Changed
- **Removed external kernel test dependencies** — the verify test suite no - **Removed external kernel test dependencies**; the verify test suite no
longer references production kernels from the separate go-libraries project. longer references production kernels from the separate go-libraries project.
The remaining test suite uses only `testdata/verify/*.s` kernels, which are The remaining test suite uses only `testdata/verify/*.s` kernels, which are
part of this repository. Coverage is identical locally and in CI (80.3 %). part of this repository. Coverage is identical locally and in CI (80.3 %).
## [0.28.0] — 2026-08-03 ## [0.28.0] - 2026-08-03
RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV
pseudo-instruction, SB/global symbol references, ELF64 object emission, and pseudo-instruction, SB/global symbol references, ELF64 object emission, and
@@ -426,20 +525,20 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
### Added ### Added
- **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. - **RISC-V encoder**; RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
- **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping. - **MOV pseudo-instruction**; load, store, reg-to-reg, immediate, frame mapping.
- **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, - **RVC compression**; 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP,
C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND,
C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR). C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
- **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` - **SB/global symbols**; `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)`
encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations.
- **GLOBL/DATA** — data section layout in `AssembleFileRISCV`. - **GLOBL/DATA**; data section layout in `AssembleFileRISCV`.
- **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects - **ELF64 emission**; `gasm asm --format elf` produces EM_RISCV objects
(.text, .data, .symtab, .rela.text). (.text, .data, .symtab, .rela.text).
- **`gasm verify --ground-truth`** — byte-exact comparison against - **`gasm verify --ground-truth`**; byte-exact comparison against
`GOARCH=riscv64 go tool asm`. `GOARCH=riscv64 go tool asm`.
- **`gasm verify --profile`** — function layout listing for RISC-V. - **`gasm verify --profile`**; function layout listing for RISC-V.
- **CALL** — AUIPC + JALR pair encoding. - **CALL**; AUIPC + JALR pair encoding.
### Fixed ### Fixed
@@ -449,7 +548,7 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
(bit-interleaved format). (bit-interleaved format).
## [0.27.0] — 2026-08-01 ## [0.27.0] - 2026-08-01
Subprocess isolation for `--fuzz`: each function is fuzzed in its own child Subprocess isolation for `--fuzz`: each function is fuzzed in its own child
process, so a partial function (decoder) that faults on random garbage is process, so a partial function (decoder) that faults on random garbage is
@@ -460,7 +559,7 @@ the parent. CRASH is informational (exit 0); only MISMATCH is an error.
- `gasm verify --fuzz` no longer crashes the process on partial functions. - `gasm verify --fuzz` no longer crashes the process on partial functions.
## [0.26.0] — 2026-07-31 ## [0.26.0] - 2026-07-31
Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written
reference. It parses the `// func` signature from the assembly source, reference. It parses the `// func` signature from the assembly source,
@@ -471,7 +570,7 @@ area bit-for-bit.
### Added ### Added
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal - `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig`; universal
differential fuzz driven by the conventional `// func` comment. Each differential fuzz driven by the conventional `// func` comment. Each
version gets its own buffer set (deep copy) so functions that write to version gets its own buffer set (deep copy) so functions that write to
their arguments (histogram increments) don't corrupt the other's input. their arguments (histogram increments) don't corrupt the other's input.
@@ -486,16 +585,16 @@ area bit-for-bit.
over-copy paths read past the buffer on random garbage input. Subprocess over-copy paths read past the buffer on random garbage input. Subprocess
isolation (fork per function) is planned. Use `--ground-truth` for decoders. isolation (fork per function) is planned. Use `--ground-truth` for decoders.
## [0.25.0] — 2026-07-30 ## [0.25.0] - 2026-07-30
Universal ground-truth verification: `gasm verify --ground-truth` assembles Universal ground-truth verification: `gasm verify --ground-truth` assembles
any `.s` file with both gasm and `go tool asm`, then compares the machine any `.s` file with both gasm and `go tool asm`, then compares the machine
code byte-for-byte per function (relocation sites masked). No hand-written code byte-for-byte per function (relocation sites masked). No hand-written
reference needed — the Go toolchain IS the oracle. reference needed; the Go toolchain IS the oracle.
### Added ### Added
- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ - `verify`: `GroundTruth`; shells out to `go tool asm`, parses the GOOBJ
output (minimal reader: block offsets, nonpkg symbol table, data index) output (minimal reader: block offsets, nonpkg symbol table, data index)
and returns per-function code bytes. and returns per-function code bytes.
- `gasm verify --ground-truth`: compares gasm's output against the Go - `gasm verify --ground-truth`: compares gasm's output against the Go
@@ -504,11 +603,11 @@ reference needed — the Go toolchain IS the oracle.
linker fills) are masked before comparison. linker fills) are masked before comparison.
- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical. - Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
## [0.24.0] — 2026-07-29 ## [0.24.0] - 2026-07-29
The full analyze family and stereo PCM decode are now differentially tested. The full analyze family and stereo PCM decode are now differentially tested.
15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the
two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex two remaining (autocorrAVX2: FMA reassociation, lpcResidualAVX2: complex
multi-arg) are deferred. multi-arg) are deferred.
### Added ### Added
@@ -518,7 +617,7 @@ multi-arg) are deferred.
- `verify`: `decodeStereo16AVX2` differential test (500 random interleaved - `verify`: `decodeStereo16AVX2` differential test (500 random interleaved
stereo PCM buffers, both channels compared sample-by-sample). stereo PCM buffers, both channels compared sample-by-sample).
## [0.23.0] — 2026-07-28 ## [0.23.0] - 2026-07-28
The analyze family and 24-bit PCM decode join the differential suite. The analyze family and 24-bit PCM decode join the differential suite.
@@ -530,7 +629,7 @@ The analyze family and 24-bit PCM decode join the differential suite.
- `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM - `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM
buffers, sign-extension compared sample-by-sample). buffers, sign-extension compared sample-by-sample).
## [0.22.0] — 2026-07-27 ## [0.22.0] - 2026-07-27
The remaining go-flac encoder kernels join the differential suite. The remaining go-flac encoder kernels join the differential suite.
@@ -544,39 +643,39 @@ The remaining go-flac encoder kernels join the differential suite.
loop). loop).
## [0.21.0] — 2026-07-26 ## [0.21.0] - 2026-07-26
Differential testing extended to all four production kernels and the CLI Differential testing extended to all four production kernels and the CLI
exposes the full dynamic-analysis toolkit. exposes the full dynamic-analysis toolkit.
### Added ### Added
- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500 - `verify`: go-flac AVX2 differential tests; `decodeMono16AVX2` (500
random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and
all four decorrelation kernels (200 iterations each: left-side, side-right, all four decorrelation kernels (200 iterations each: left-side, side-right,
mid-side, interleave) compared bit-for-bit against the portable Go mid-side, interleave) compared bit-for-bit against the portable Go
references. references.
- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000 - `verify`: go-lz4 AVX-512 differential tests; `decodeBlockAVX512` (3 000
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes) fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0-1024 bytes)
against the same portable oracle as the AVX2 suite. against the same portable oracle as the AVX2 suite.
- `gasm verify --abi`: runs each NOSPLIT function with sentinel registers - `gasm verify --abi`: runs each NOSPLIT function with sentinel registers
and a red-zone canary, reporting violations. and a red-zone canary, reporting violations.
- `gasm verify --profile`: lists the static basic-block count per function. - `gasm verify --profile`: lists the static basic-block count per function.
## [0.20.0] — 2026-07-25 ## [0.20.0] - 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block Coverage profiling: the third pillar of Phase 3. Static basic-block
enumeration from the assembler's label map, combined with multi-input path enumeration from the assembler's label map, combined with multi-input path
diversity measurement — how many observationally distinct execution paths a diversity measurement; how many observationally distinct execution paths a
test corpus exercises. test corpus exercises.
### Added ### Added
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks - `verify`: `Kernel.Blocks` / `Kernel.BlockCount`; enumerate basic blocks
from the assembler's local-label map (every jump target is a block from the assembler's local-label map (every jump target is a block
boundary; the function entry is always a block). `decodeBlockAVX2` has boundary; the function entry is always a block). `decodeBlockAVX2` has
27 blocks. 27 blocks.
- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of - `verify`: `Kernel.ProfilePaths`; run the function with a corpus of
argument blocks and collect distinct output fingerprints (the result argument blocks and collect distinct output fingerprints (the result
words); reports path diversity as a lower bound on code coverage. words); reports path diversity as a lower bound on code coverage.
@@ -588,7 +687,7 @@ rt_sigaction handlers fragile in a Go process. The static + path-diversity
approach delivers the project's goal (proving the SIMD path and tail handling approach delivers the project's goal (proving the SIMD path and tail handling
execute) without fighting the runtime. execute) without fighting the runtime.
## [0.19.0] — 2026-07-24 ## [0.19.0] - 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now
has an ABI-checking variant that sets sentinels in the callee-saved registers has an ABI-checking variant that sets sentinels in the callee-saved registers
@@ -598,41 +697,41 @@ detects any illegal write below the stack pointer.
### Added ### Added
- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call - `verify`: `CallChecked` / `Kernel.CallFuncChecked`; ABI-checking JIT call
with sentinel registers and red-zone canary; returns an `ABIReport` with sentinel registers and red-zone canary; returns an `ABIReport`
(BPClobbered, R14Clobbered, RedZoneHit). (BPClobbered, R14Clobbered, RedZoneHit).
- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no - `verify`: the raw `leaveJITCheckedRaw` trampoline; a TEXT symbol with no
ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT
function's RET lands directly in the check code and sees the registers function's RET lands directly in the check code and sees the registers
exactly as the function left them. exactly as the function left them.
- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels - Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels
confirmed ABI-clean (BP preserved, R14 preserved, red zone intact). confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
## [0.18.0] — 2026-07-23 ## [0.18.0] - 2026-07-23
Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared fuzzed against a portable Go reference; 5 000 valid LZ4 blocks compared
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
codes. This is the automated form of the project's bit-identical contract. codes. This is the automated form of the project's bit-identical contract.
### Added ### Added
- `verify`: differential fuzz tests — a random LZ4 block generator produces - `verify`: differential fuzz tests; a random LZ4 block generator produces
valid blocks (literals, overlapping matches, extension bytes) and the valid blocks (literals, overlapping matches, extension bytes) and the
JIT-assembled kernel's output is compared byte-for-byte against a portable JIT-assembled kernel's output is compared byte-for-byte against a portable
Go decoder; a hostile-input suite confirms error-code agreement on random Go decoder; a hostile-input suite confirms error-code agreement on random
garbage (no crashes, same classification). garbage (no crashes, same classification).
## [0.17.0] — 2026-07-22 ## [0.17.0] - 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go Plan 9 amd64 kernels into executable memory and calls them directly; pure Go
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
toolchain. toolchain.
### Added ### Added
- `verify` package: JIT infrastructure — `Map` copies machine code into a - `verify` package: JIT infrastructure; `Map` copies machine code into a
W^X memory mapping, `Call` invokes it through an ABI0 trampoline that W^X memory mapping, `Call` invokes it through an ABI0 trampoline that
switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST` switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST`
parse, assemble and map a `.s` file in one step; `Kernel.CallFunc` parse, assemble and map a `.s` file in one step; `Kernel.CallFunc`
@@ -641,34 +740,34 @@ toolchain.
available functions; with `-smoke`, calls each NOSPLIT function with available functions; with `-smoke`, calls each NOSPLIT function with
zeroed arguments to confirm the trampoline works end-to-end. zeroed arguments to confirm the trampoline works end-to-end.
- Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2` - Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2`
kernels (699 and 146 bytes) assemble, map and execute correctly — kernels (699 and 146 bytes) assemble, map and execute correctly;
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes known-answer LZ4 blocks decode bit-for-bit, wide copies of 0-1024 bytes
match, malformed input returns the correct error codes. match, malformed input returns the correct error codes.
## [0.16.0] — 2026-07-21 ## [0.16.0] - 2026-07-21
The scalar conversions between vector and general-purpose registers — the The scalar conversions between vector and general-purpose registers; the
last of the amd64 EVEX instruction set. last of the amd64 EVEX instruction set.
### Added ### Added
- `asm`: the GPR-interchanging conversions, byte for byte against the Go - `asm`: the GPR-interchanging conversions, byte for byte against the Go
assembler (28 ground-truth cases including memory sources and extended assembler (28 ground-truth cases including memory sources and extended
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} GPRs): vector to GPR; the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and (EVEX only); GPR to vector; VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
vector source sits in vvvv (three Plan 9 operands). vector source sits in vvvv (three Plan 9 operands).
## [0.15.0] — 2026-07-20 ## [0.15.0] - 2026-07-20
The last of the EVEX conversions and narrowing/extending moves — the EVEX The last of the EVEX conversions and narrowing/extending moves; the EVEX
instruction set is now complete save for the GPR-interchanging forms. instruction set is now complete save for the GPR-interchanging forms.
### Added ### Added
- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y - `asm`: the unsigned and truncating conversions; VCVTPD2PS (and the X/Y
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
@@ -681,20 +780,20 @@ instruction set is now complete save for the GPR-interchanging forms.
D2M/Q2M), whose K register is a genuine operand rather than a mask and D2M/Q2M), whose K register is a genuine operand rather than a mask and
which therefore take no masking suffixes. which therefore take no masking suffixes.
## [0.14.0] — 2026-07-19 ## [0.14.0] - 2026-07-19
The floating-point helper and conversion tail of the AVX-512 set, plus The floating-point helper and conversion tail of the AVX-512 set, plus
gather and scatter with VSIB addressing — every encoding verified byte for gather and scatter with VSIB addressing; every encoding verified byte for
byte against the Go assembler. byte against the Go assembler.
### Added ### Added
- `asm`: the floating-point helpers — reciprocals and reciprocal square - `asm`: the floating-point helpers; reciprocals and reciprocal square
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask VFPCLASSSD/SS; a new immediate form whose reg field carries the opmask
destination). destination).
- `asm`: **gather and scatter with VSIB addressing.** The gathers take - `asm`: **gather and scatter with VSIB addressing.** The gathers take
both Go spellings: the VEX form with a vector mask register (OP mask, both Go spellings: the VEX form with a vector mask register (OP mask,
@@ -703,14 +802,14 @@ byte against the Go assembler.
data register (a ZMM index with an YMM destination encodes L'L = 10, as data register (a ZMM index with an YMM destination encodes L'L = 10, as
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
only (OP src, K, vsib). All eight gather and eight scatter widths. only (OP src, K, vsib). All eight gather and eight scatter widths.
- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets - `asm`: the remaining conversions; VCVTQQ2PS (the 512-bit source sets
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate). VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
## [0.13.0] — 2026-07-18 ## [0.13.0] - 2026-07-18
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every the opmask instructions and the EVEX rounding/SAE/broadcast suffixes; every
encoding verified byte for byte against the Go assembler. encoding verified byte for byte against the Go assembler.
### Added ### Added
@@ -718,7 +817,7 @@ encoding verified byte for byte against the Go assembler.
- `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth - `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS;
a new NDS-plus-immediate form with the K register in the reg field), the a new NDS-plus-immediate form with the K register in the reg field), the
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
@@ -732,15 +831,15 @@ encoding verified byte for byte against the Go assembler.
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
VPMOVQB). VPMOVQB).
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding - `asm`: the EVEX mnemonic suffixes the Go assembler accepts; the rounding
modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the
rounding control in L'L), suppress-all-exceptions `.SAE`, and memory rounding control in L'L), suppress-all-exceptions `.SAE`, and memory
broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled
by the element size) — each combinable with the `.Z` zeroing suffix, by the element size); each combinable with the `.Z` zeroing suffix,
validated against the Go assembler's bytes, and rejected on instructions validated against the Go assembler's bytes, and rejected on instructions
that do not support them. that do not support them.
## [0.12.0] — 2026-07-17 ## [0.12.0] - 2026-07-17
GOOBJ emission: gasm-assembled functions drop into a `go build` without the GOOBJ emission: gasm-assembled functions drop into a `go build` without the
Go assembler. Go assembler.
@@ -748,7 +847,7 @@ Go assembler.
### Added ### Added
- `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>` - `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>`
writes the Go toolchain's own object format — the one `cmd/link` consumes writes the Go toolchain's own object format; the one `cmd/link` consumes
directly: the functions as non-package symbols qualified with the package directly: the functions as non-package symbols qualified with the package
path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data, path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data,
one serialized `FuncInfo` per function (argument/frame sizes, the asm one serialized `FuncInfo` per function (argument/frame sizes, the asm
@@ -757,8 +856,8 @@ Go assembler.
real stack deltas: the assembler now tracks every stack-adjustment real stack deltas: the assembler now tracks every stack-adjustment
boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each
`RET`'s epilogue, so frame-pointer functions unwind correctly. The `RET`'s epilogue, so frame-pointer functions unwind correctly. The
object preamble — the version-and-experiment header the linker compares object preamble; the version-and-experiment header the linker compares
verbatim — is captured from the installed `go tool asm`, so the output is verbatim; is captured from the installed `go tool asm`, so the output is
always consistent with the toolchain that links it. always consistent with the toolchain that links it.
- `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL` - `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL`
entries in the GOOBJ output, with the instruction's displacement field entries in the GOOBJ output, with the instruction's displacement field
@@ -771,7 +870,7 @@ Go assembler.
pattern, instead of being rejected as non-integer. pattern, instead of being rejected as non-integer.
## [0.11.0] — 2026-07-16 ## [0.11.0] - 2026-07-16
Linkable object output: external symbols and relocatable ELF / Mach-O Linkable object output: external symbols and relocatable ELF / Mach-O
objects. objects.
@@ -790,7 +889,7 @@ objects.
external symbol; the Mach-O output is verified structurally with external symbol; the Mach-O output is verified structurally with
`debug/macho`. `debug/macho`.
- `asm`: **external symbol references.** A reference to a symbol no - `asm`: **external symbol references.** A reference to a symbol no
`GLOBL` in the file defines no longer aborts assembly — it is recorded `GLOBL` in the file defines no longer aborts assembly; it is recorded
as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and
becomes an undefined global symbol in the object output. The raw image becomes an undefined global symbol in the object output. The raw image
s (`--format raw`, the default) still reports them: only an object s (`--format raw`, the default) still reports them: only an object
@@ -802,7 +901,7 @@ objects.
writes; without `--format` the behaviour is unchanged (the concatenated writes; without `--format` the behaviour is unchanged (the concatenated
image). image).
## [0.10.0] — 2026-07-15 ## [0.10.0] - 2026-07-15
The EVEX floating-point and conversion set: the packed-double arithmetic, The EVEX floating-point and conversion set: the packed-double arithmetic,
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
@@ -810,23 +909,23 @@ verified byte for byte against the Go assembler.
### Added ### Added
- `asm`: the rest of the common EVEX/VEX floating-point set — packed double - `asm`: the rest of the common EVEX/VEX floating-point set; packed double
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
family in both VEX and EVEX — the EVEX scalar forms exist for masked and family in both VEX and EVEX; the EVEX scalar forms exist for masked and
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX). zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and - `asm`: the width-changing conversions; VCVTDQ2PS and VCVTPS2PD (VEX and
EVEX; the destination sets the length for PS→PD), the EVEX form of EVEX; the destination sets the length for PS→PD), the EVEX form of
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
source — a new operand form, since the destination is always XMM while source; a new operand form, since the destination is always XMM while
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
memory source). memory source).
- `asm`: masking and zeroing on every new form — the scalar SD/SS - `asm`: masking and zeroing on every new form; the scalar SD/SS
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
explicit K1–K7 operand and the `.Z` suffix the way Go writes them. explicit K1-K7 operand and the `.Z` suffix the way Go writes them.
### Changed ### Changed
@@ -837,35 +936,35 @@ verified byte for byte against the Go assembler.
shares the convention). shares the convention).
## [0.9.0] — 2026-07-14 ## [0.9.0] - 2026-07-14
AVX-512 masking and a wider EVEX integer set. AVX-512 masking and a wider EVEX integer set.
### Added ### Added
- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7` - `asm`: **EVEX masking** the way Go writes it; an explicit `K1`-`K7`
operand placed among the operands (merging mask), and a `.Z` mnemonic operand placed among the operands (merging mask), and a `.Z` mnemonic
suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS, suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS,
reg/rm, immediate-shift, align, extract, convert and move forms, including reg/rm, immediate-shift, align, extract, convert and move forms, including
masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is
rejected as an explicit mask, and `.Z` without a mask is an error, matching rejected as an explicit mask, and `.Z` without a mask is an error, matching
the Go assembler. the Go assembler.
- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, - `asm`: the common AVX-512 F/BW integer set; VPADDB/W, VPSUBB/W, VPANDD/Q,
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] Register indices 16-31 encode correctly (the mod=11 quirk carries rm[4]
in X̄). All verified byte for byte against the Go assembler. in X̄). All verified byte for byte against the Go assembler.
- `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by - `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by
`unknown-instruction` and exempted from `operand-count`. `unknown-instruction` and exempted from `operand-count`.
### Fixed ### Fixed
- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4] - `asm`: EVEX register-register operands with indices 16-31 encoded rm[4]
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
prefix bytes for X16+/Y16+ r/m operands. prefix bytes for X16+/Y16+ r/m operands.
## [0.8.0] — 2026-07-13 ## [0.8.0] - 2026-07-13
Standard CLI ergonomics. Standard CLI ergonomics.
@@ -881,20 +980,20 @@ Standard CLI ergonomics.
- The version is primarily available as the standard `gasm --version` / `-V` - The version is primarily available as the standard `gasm --version` / `-V`
flag; the `gasm version` spelling remains as an alias. flag; the `gasm version` spelling remains as an alias.
## [0.7.0] — 2026-07-12 ## [0.7.0] - 2026-07-12
The formatter behaves like `go fmt` and canonicalises block separation. The formatter behaves like `go fmt` and canonicalises block separation.
### Added ### Added
- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory - `gasm fmt` now works like `go fmt`: with no arguments; or with a directory
argument — it reformats every `.s` file below it in place and lists the argument; it reformats every `.s` file below it in place and lists the
changed files, skipping `.` and `_` directories (`.git`, `_refs`, …). changed files, skipping `.` and `_` directories (`.git`, `_refs`, …).
Explicit file arguments keep the `-w` / standard-output behaviour. Explicit file arguments keep the `-w` / standard-output behaviour.
### Changed ### Changed
- `s`: canonical blank-line layout — a new block (a label, `TEXT` or - `s`: canonical blank-line layout; a new block (a label, `TEXT` or
`GLOBL`) is preceded by exactly one blank line, neither more nor less. `GLOBL`) is preceded by exactly one blank line, neither more nor less.
Comments leading a block stay with it (the blank line goes before them), Comments leading a block stay with it (the blank line goes before them),
stacked labels share their block, the function's first label keeps hugging stacked labels share their block, the function's first label keeps hugging
@@ -903,7 +1002,7 @@ The formatter behaves like `go fmt` and canonicalises block separation.
kernels were reformatted with this release and remain byte-identical when kernels were reformatted with this release and remain byte-identical when
assembled. assembled.
## [0.6.0] — 2026-07-11 ## [0.6.0] - 2026-07-11
Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and
the encoder learns the legacy SSE moves. the encoder learns the legacy SSE moves.
@@ -912,13 +1011,13 @@ the encoder learns the legacy SSE moves.
- `lint`: **`register-clobber` is now calibrated to the Go ABI** - `lint`: **`register-clobber` is now calibrated to the Go ABI**
(`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based (`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based
ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15` ABI0 has no System V style callee-saved registers; amd64 `BX`, `R12`-`R15`
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
scratch, and hand-written kernels may clobber them freely. The rule now scratch, and hand-written kernels may clobber them freely. The rule now
audits only the registers Go fixes across calls: the frame pointer and the audits only the registers Go fixes across calls: the frame pointer and the
goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64 goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64
`X27`, loong64 `R22`), and the goroutine pointer is reported only when the `X27`, loong64 `R22`), and the goroutine pointer is reported only when the
function can reach the runtime (is not `NOSPLIT` or makes a call) — the function can reach the runtime (is not `NOSPLIT` or makes a call); the
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
may use it, exactly as the runtime's own assembly does. Both go-flac may use it, exactly as the runtime's own assembly does. Both go-flac
kernels now lint with zero diagnostics. kernels now lint with zero diagnostics.
@@ -935,21 +1034,21 @@ the encoder learns the legacy SSE moves.
### Added ### Added
- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names - `asm`: the legacy (non-VEX) SSE moves; `MOVOU`/`MOVO` (the Plan 9 names
for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar
`MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for `MOVSD`/`MOVSS`; and `VMOVDQU64` in the EVEX set. All verified byte for
byte against the Go assembler. byte against the Go assembler.
## [0.5.0] — 2026-07-10 ## [0.5.0] - 2026-07-10
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage. the Go toolchain, completing the production-kernel coverage.
### Added ### Added
- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the - `asm`: **EVEX (AVX-512) encoding**; the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and 5-bit register fields (Z0-Z31, X/Y 16-31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as V'̄ shared between vvvv and the SIB index), opmask registers (K0-K7) as
operands and as mask destinations, and the compressed disp8×N displacement operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's (the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac opcode tables prescribe). Covers every AVX-512 instruction the go-flac
@@ -959,20 +1058,20 @@ the Go toolchain, completing the production-kernel coverage.
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither. ; the kernels use neither.
- `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols - `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols
too; a reference is external only when no `GLOBL` in the file defines it. too; a reference is external only when no `GLOBL` in the file defines it.
### Fixed ### Fixed
- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; - `asm`: registers X16-Y31 force the EVEX encoding of dual-form mnemonics;
previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them. represent indices above 15 and silently truncated them.
- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of - `asm`: the VEX encoder now rejects vector register indices 16-31 instead of
encoding a truncated (wrong) register. encoding a truncated (wrong) register.
## [0.4.0] — 2026-07-09 ## [0.4.0] - 2026-07-09
The standalone assembler reaches the whole go-flac AVX2 kernel: static The standalone assembler reaches the whole go-flac AVX2 kernel: static
symbols assemble, and all 17 kernel functions now match the Go toolchain's symbols assemble, and all 17 kernel functions now match the Go toolchain's
@@ -980,10 +1079,10 @@ machine code byte for byte.
### Added ### Added
- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an - `asm`: **file-level assembly**; `AssembleFile` turns a parsed file into an
`Image`: the function bodies in source order followed by a data section `Image`: the function bodies in source order followed by a data section
built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned). built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned).
- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as - `asm`: **static-symbol (`SB`) operands**; `mask<>(SB)` references encode as
RIP-relative loads with a patched disp32, resolved against the image layout RIP-relative loads with a patched disp32, resolved against the image layout
so the output is self-consistent and position-independent. External so the output is self-consistent and position-independent. External
(non-file-local) symbols are rejected with a clear error: they need (non-file-local) symbols are rejected with a clear error: they need
@@ -992,7 +1091,7 @@ machine code byte for byte.
and writes the whole image (code + data) with `-o`. and writes the whole image (code + data) with `-o`.
## [0.3.0] — 2026-07-08 ## [0.3.0] - 2026-07-08
The assembler reaches byte-identical parity with the Go toolchain on the The assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that production go-flac AVX2 kernels: every one of the 15 kernel functions that
@@ -1002,18 +1101,18 @@ support).
### Added ### Added
- `asm`: the scalar instruction families the kernels use — `CMOVcc` and - `asm`: the scalar instruction families the kernels use; `CMOVcc` and
`SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT` `SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT`
(legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`, (legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`,
`MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the `MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the
legacy SSE encoding, as the Go assembler emits it), the traditional legacy SSE encoding, as the Go assembler emits it), the traditional
three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts
(`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the (`VPSRLQ X0, Y8, Y8`; the count in an XMM register or memory takes the
ordinary NDS form). ordinary NDS form).
- `asm`: **jump relaxation** — jumps start in the short (rel8) form and - `asm`: **jump relaxation**; jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating the expand to rel32 when the settled displacement does not fit, iterating the
layout to a fixed point (CALL is always rel32). layout to a fixed point (CALL is always rel32).
- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only - `asm`: **jump-to-jump folding**; a conditional jump to a label whose only
instruction is an unconditional jump is redirected to the ultimate instruction is an unconditional jump is redirected to the ultimate
target, replicating the Go toolchain's linker, which chases such chains target, replicating the Go toolchain's linker, which chases such chains
before it encodes branches. before it encodes branches.
@@ -1025,12 +1124,12 @@ support).
- `asm`: `CMP` with a register or memory operand computed **second − first** - `asm`: `CMP` with a register or memory operand computed **second − first**
instead of first − second, silently inverting every condition that followed instead of first − second, silently inverting every condition that followed
(`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records (`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records
first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` first − second; `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
with the first operand in reg — and is byte-identical to the Go assembler. with the first operand in reg; and is byte-identical to the Go assembler.
- `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg = - `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg =
source), the Go assembler's choice; the output is byte-identical. source), the Go assembler's choice; the output is byte-identical.
## [0.2.0] — 2026-07-07 ## [0.2.0] - 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms. and the moves, on top of the Phase 1 VEX forms.
@@ -1043,10 +1142,10 @@ and the moves, on top of the Phase 1 VEX forms.
- the immediate shuffle (`VPSHUFD`, `VPERMQ`), - the immediate shuffle (`VPSHUFD`, `VPERMQ`),
- the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, - the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`,
`VINSERTI128`), `VINSERTI128`),
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies - the lane extract (`VEXTRACTI128`, `VEXTRACTF128`; the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field), the ModRM.reg field, the XMM/memory destination the r/m field),
- the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, - the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`,
`VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector `VMOVSD`; each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler), move uses the store-form layout, matching the Go assembler),
- the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form, - the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form,
- the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`, - the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`,
@@ -1055,21 +1154,21 @@ and the moves, on top of the Phase 1 VEX forms.
the encoder now covers every integer, shuffle and FP instruction the the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use. go-flac AVX2 kernels use.
- `asm`: `CMP` accepts the immediate in the second operand position - `asm`: `CMP` accepts the immediate in the second operand position
(`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it (`CMPL CX, $31`), the spelling the Go assembler accepts, encoding it
identically to the immediate-first form. identically to the immediate-first form.
### Fixed ### Fixed
- `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as - `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as
the hardware requires — the previous value (`0000`) made the two-operand the hardware requires; the previous value (`0000`) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite. comparison (added this release) is now part of the test suite.
## [0.1.0] — 2026-07-06 ## [0.1.0] - 2026-07-06
Initial release — the Phase 1 foundation. Initial release; the Phase 1 foundation.
### Added ### Added
@@ -1082,11 +1181,11 @@ Initial release — the Phase 1 foundation.
- `arch`: register files and **complete** instruction tables for amd64, - `arch`: register files and **complete** instruction tables for amd64,
arm64, riscv64 and loong64, with the middle-dot symbol separator and static arm64, riscv64 and loong64, with the middle-dot symbol separator and static
(`<>`) symbols. Instruction names are generated from the Go toolchain's own (`<>`) symbols. Instruction names are generated from the Go toolchain's own
assembler source (`just gen`) — the `anames` opcode lists plus the common assembler source (`just gen`); the `anames` opcode lists plus the common
opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the
`.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump `.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump
spellings) — so every mnemonic the real assembler accepts is recognised. spellings); so every mnemonic the real assembler accepts is recognised.
- `lint`: conservative rules — `unknown-instruction`, `operand-count`, - `lint`: conservative rules; `unknown-instruction`, `operand-count`,
`undefined-label`, `duplicate-label`, `missing-ret`, `undefined-label`, `duplicate-label`, `missing-ret`,
`missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro
invocations are recognised (in-file `#define` names and underscore invocations are recognised (in-file `#define` names and underscore
@@ -1108,9 +1207,9 @@ Initial release — the Phase 1 foundation.
- `lsp`: a Language Server Protocol server over stdio providing completion, - `lsp`: a Language Server Protocol server over stdio providing completion,
hover documentation, document symbols, publish-diagnostics and semantic-token hover documentation, document symbols, publish-diagnostics and semantic-token
highlighting. highlighting.
- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ - `asm`: a standalone amd64 (x86-64) assembler; an instruction encoder (REX/
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — AVX2 SIMD across the three operand forms NDS, reg/rm and immediate-shift,
covering the bulk of the integer SIMD set) validated by round-trip decoding covering the bulk of the integer SIMD set) validated by round-trip decoding
against `golang.org/x/arch`, and an assembler that drives the parser's AST against `golang.org/x/arch`, and an assembler that drives the parser's AST
into the encoder with local-label resolution and `FP`/`SP` frame mapping into the encoder with local-label resolution and `FP`/`SP` frame mapping
+98 -78
View File
@@ -1,107 +1,127 @@
# Contributing to gasm-devkit # Contributing
Thanks for contributing to gasm-devkit. Contributions to **gasm-devkit** are governed by the Contributor terms
below; submitting one means you accept them.
## Contributor terms
1. This project belongs to its owner alone. The owner decides what is
accepted, in what form and when; the decision is final and needs no
justification.
2. By submitting a contribution you assign to Petr Balvín
<opensource@petrbalvin.org> all present and future copyright and
related rights in it, worldwide, for the full term of the rights,
with the right to relicense and sublicense without restriction,
including under proprietary terms.
3. Where that assignment is not effective, it counts as a perpetual,
irrevocable, royalty-free licence with the same scope.
4. To the fullest extent permitted by law, you waive any right of
attribution and integrity in the contribution. The project names no
contributors and keeps no credits list.
5. By submitting you represent that the work is yours and that you
hold the rights to assign it as above.
## Development setup ## Development setup
Requirements: Go 1.27 or later, the [just](https://github.com/casey/just) Requirements: Go 1.27.1, the exact version the `go` directive in `go.mod`
command runner, and a Linux host on amd64, arm64, riscv64 or loong64. declares, and [just](https://github.com/casey/just) for the recipes.
```sh ```sh
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
cd gasm-devkit cd gasm-devkit
just install # download module dependencies just build
just build # go vet + gofmt check just gates
just test # full suite, race detector, 80 % coverage gate
``` ```
## Workflow ## Workflow
1. Branch from `development`; never commit directly to `main` (`main` is 1. Branch from `development`. Never commit directly to `main`, which is release-only.
release-only: merge from `development`, then tag). 2. Commit in [Conventional Commits](https://www.conventionalcommits.org/) form:
2. Commit with [Conventional Commits](https://www.conventionalcommits.org/): `type(scope): description`, subject line only, imperative mood, lowercase after the
`type(scope): description`: subject line only, imperative mood, colon, no trailing full stop. Allowed types: `feat`, `fix`, `docs`, `style`,
lowercase after the colon, no trailing dot. Allowed types: `feat`, `refactor`, `perf`, `test`, `chore`, `ci`, `build`, `revert`.
`fix`, `docs`, `style`, `refactor`, `perf`, `test`, `chore`, `ci`, 3. One logical change per commit. A refactor, a behaviour change and a formatting pass
`build`, `revert`. The only line after the subject is the trailer: are three commits, never one.
`Assisted-by: <model-name>`. No `Co-Authored-By`, no `Signed-off-by`, 4. Record every user-visible change in `CHANGELOG.md` under `## [development]`.
no other trailers. 5. Add or update tests. Coverage stays at 80 percent or more; it is a hard gate.
3. Record every user-visible change in `CHANGELOG.md` under 6. Update the documentation when the public API, the configuration or the behaviour
`## [development]` (categories: Added, Changed, Fixed, Removed, changes.
Security). 7. Open a pull request against `development`.
4. Add or update tests; coverage must stay **at or above 80 %** (hard
gate, enforced by CI).
5. Update the documentation when behaviour, flags or the public surface
change.
6. Open a pull request against `development`.
Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`; Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`. The release
CI builds and publishes the binaries for all four architectures. workflow builds the assets and publishes the release and its notes.
## Code style ## Code style
`gofmt` and `go vet` via `just fmt` / `just build`; both must pass with `gofmt` and `go vet` run through `just fmt` and `just vet`, with zero diff and zero
zero output; `go fix -diff ./...` must report nothing on touched packages. warnings tolerated. `just gates` is the definition of done in one command, and the recipe
file names what it contains. Errors are checked explicitly, wrapped as
`fmt.Errorf("context: %w", err)`, and nothing panics outside `main`. The `golang`
skill holds the rules the project follows; the recipe file holds the commands.
- Standard library only in production code; `golang.org/x/arch` is used - `golang.org/x/arch` is the one module dependency, and it is linked into the binary:
in tests only (round-trip decoding) and is never linked into the `gasm` `gasm dis` and the debugger's listings decode through it. Everything else is the
binary. standard library.
- No cgo, no C, no external toolchains at runtime. - No cgo, no C, no external toolchain at runtime.
- Explicit `if err != nil`; errors wrapped with - The parser, lexer and formatter are hand-written; the `arch` instruction tables are
`fmt.Errorf("context: %w", err)`; no panics outside `main`. generated only by `_gen/gen.go` (`just gen`) and never edited by hand.
- The parser, lexer and formatter are hand-written; the `arch` instruction - Assembly committed to the repository goes through `gasm fmt` and `gasm lint`, so a
tables are generated only via `_gen/gen.go` (`just gen`), never edited. `.s` file that `gasm fmt -l .` lists is unfinished.
## Running a single test New source files open with the project's two-line licence header, whose SPDX
identifier matches `LICENSE`. Configuration files, workflows and dotfiles do not carry
it.
```sh ## AI contribution policy
go test -run TestVexGroundTruth ./asm/
go test -run TestGroundTruthBasic ./verify/
go test -run TestGOObjectLinkAndRun ./asm/
go test -run TestFuzzWideCopy ./verify/
```
The interactive debugger (`gasm debug`) requires a compiled binary on AI tools are welcome as productivity aids and are a normal part of modern software
`$PATH`; `go run` does not work for the traced child process. Install development. What matters is that the contribution stays understandable, reviewable and
first with `just install-bin`. genuinely useful.
## CI (Gitea Actions) - **Disclose the assistance.** If AI helped draft any part of a commit, issue, pull
request or review, say so.
- **Commit messages carry exactly one trailer**, as a git trailer on the line after a
blank line that closes the subject:
Workflows live in `.gitea/workflows/` and run on self-hosted runners: ```
Assisted-by: MODEL
```
Name the model that did the work, spelled the way its maker spells it, for example
`GLM 5.3`, `DeepSeek V4.1 Flash` or `Qwen 3.8 Flash`. No `Co-Authored-By`, no `Signed-off-by`,
no other trailers, and no prose: the trailer is the disclosure.
- **Issues and pull requests** attribute the assistance in a comment, for example
`_Assisted-by: GLM 5.3_`. It does not belong in the pull request description.
- **Take responsibility.** You are accountable for the accuracy, completeness and
intent of everything you submit, whether or not AI produced it.
- **Review before marking ready.** Read the diff carefully, run it locally, and add the
tests it needs. Do not mark a pull request ready until you can defend every change in
it.
- **Quality over quantity.** Contributions that look like un-reviewed output, or whose
author cannot engage substantively during review, may be closed.
- **Preferred models.** Prefer open-weight models with transparent training data and
minimal output filtering.
AI assists. It does not replace judgement.
## Continuous integration
Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | What it does | | Workflow | Trigger | What it does |
|----------|---------|--------------| |---|---|---|
| Test | push / PR to `development` | gofmt check, `go vet`, `go test -race`, 80 % coverage gate | | Test | push or pull request to `development` | build, format check, vet, modernisation, the test suite with the coverage floor |
| Release | tag `v*` | cross-compiles binaries for linux/{amd64,arm64,riscv64,loong64} and publishes the Gitea release | | Release | a `v*` tag | the same gates as Test, then the matrix build, the proven version and the release itself; the race detector runs locally in `just gates` before the tag is cut |
The Definition of Done (`just build` + `just test` + `just fmt`) must The local equivalent is `just gates`, which is the same set plus the race detector. The
still pass locally before pushing. race detector also has its own workflow, dispatched by hand; it never runs on a push or a
tag, where it would double the time and the memory a shared runner cannot spare.
## AI Contribution Policy
AI tools are welcome as productivity aids. What matters is that
contributions remain understandable, reviewable, and genuinely useful.
- **Disclose AI use.** If you used AI to draft or generate any part of a
commit, issue, pull request, or code review, say so clearly.
- **Commit messages:** end every commit with exactly one trailer:
`Assisted-by: <model-name>` (e.g. `Assisted-by: GLM 5.3`).
- **Pull requests and issues:** attribute AI assistance in one trailing
line, e.g. `_Assisted-by: GLM 5.3_`. Do not paste it into the PR
description as a section.
- **Take responsibility.** You remain accountable for the accuracy,
completeness, and intent of everything you submit.
- **Review before marking ready.** Read AI-generated diffs carefully, run
them locally, and add or update tests where appropriate.
- **Preferred models.** Prefer open-weight models with transparent
training data: **GLM**, **DeepSeek**, and **MiMo**.
## Reporting bugs ## Reporting bugs
Open an issue at Open an issue at `https://sourcedock.dev/petrbalvin/gasm-devkit/issues` with the
[sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit/issues) version, the operating system and architecture, the exact command, the full output,
with the version (`gasm --version`), OS and architecture, the exact and the expected against the actual behaviour.
command, the full output, and the expected versus actual behaviour.
**Security issues:** email **opensource@petrbalvin.org** instead of opening **Security issues do not go in the issue tracker.** Report them as
a public issue. [SECURITY.md](SECURITY.md) describes, to **opensource@petrbalvin.org**.
+111 -23
View File
@@ -1,13 +1,63 @@
# gasm-devkit # Plan 9 assembly tooling, inside and outside Go
Developer tooling for **GAsm**, Go's built-in Plan 9 assembler. > **Warning: this is an experiment.** gasm-devkit is under active
> development and is not stable. The version is 0.x.x: commands, flags,
> output formats and behaviour can change without warning at any time.
> A 1.0.0 release is light years away. Nothing in this document is a
> stability promise. For all of that, this is not a paper project: gasm
> is already in active use and is tested on real assembly work.
Go ships an assembler but no tooling for it: there is no syntax highlighting, **GAsm** is Go's Plan 9 assembler, and Go ships it without tooling:
no autocomplete, no linter, no static analyser, no formatter, no standalone there is no formatter, no linter, no static analyser, no standalone
assembler and no debugger for `.s` files. Developers write assembly blind, assembler and no debugger for `.s` files. Developers write assembly
validate it by benchmark, and debug it by print statement. gasm-devkit is the blind, validate it by benchmark, and debug it by print statement.
missing toolkit: a single, self-contained binary, `gasm`, that brings proper gasm-devkit is the missing toolkit: a single, self-contained binary,
developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64. `gasm`, that serves both purposes.
- **Help develop Plan 9 assembly.** Formatting, linting, disassembly,
dynamic verification, a source-level debugger and a language server,
for `.s` files in Go programs.
- **Use Plan 9 assembly outside the Go toolchain.** `gasm asm` encodes
on its own, with no Go installation in the loop, and writes raw
images, linkable ELF objects with DWARF5 debug sections, or the Go
toolchain's own GOOBJ format, which `go build` consumes in place of
the toolchain's output.
## Why Plan 9 assembly
Plan 9 assembly is the quiet triumph of the field. One syntax across
every architecture Go builds for: the same source-first operand order,
the same four pseudo-registers, the same frame convention, whether the
target is x86, ARM, RISC-V or LoongArch. Learn it once and you can
read a kernel on any of them.
Compare the alternatives. Intel syntax and AT&T syntax disagree on the
one question every instruction answers, which operand is the source
and which is the destination, so half the world writes it one way,
half the other, and every assembly programmer carries both in their
head forever. GNU as settles the argument with directives that switch
dialects mid-file (`.intel_syntax noprefix`), a percent sign on every
register and a dollar on every immediate: punctuation that carries
nothing the operand order did not already say. And the x86 family
fragments again underneath: NASM is not MASM is not GAS, each with its
own directive zoo and macro language, so every project picks a dialect
and every reader learns a different one by accident.
Plan 9 assembly has none of it. Registers are bare names. Memory is
one notation, `offset(base)`, extended by an index and a scale when
the instruction needs it. Arguments arrive named and offset-checked:
`x+0(FP)` is the argument x, on every architecture, and `go vet`
polices the offsets against the Go prototype.
```text
AT&T (GNU as): movq %rax, -16(%rbp)
Plan 9 (Go): MOVQ AX, total-16(SP)
```
The same lines, but only one of them tells you what the number is for.
The syntax is uppercase, regular and boring, which is the highest
compliment a language for machine code can earn. gasm-devkit exists
to give that syntax the tooling it deserves.
## Features ## Features
@@ -47,12 +97,11 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
assembly files byte-for-byte, `gasm profile` shows basic-block structure, assembly files byte-for-byte, `gasm profile` shows basic-block structure,
`gasm audit-instructions` diffs the encoder against the installed toolchain, `gasm audit-instructions` diffs the encoder against the installed toolchain,
and `gasm scaffold` generates a differential test skeleton for a kernel. and `gasm scaffold` generates a differential test skeleton for a kernel.
- **Complete instruction coverage.** The instruction tables are generated
from the Go toolchain's own assembler source, so the toolkit recognises
every mnemonic the real assembler accepts; `just gen` refreshes them.
### Architecture support ### Architecture support
Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised | | Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------| |--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases | | AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
@@ -63,27 +112,64 @@ developer tooling to Plan 9 assembly on amd64, arm64, riscv64 and loong64.
"Common opcodes" are the instructions shared by every architecture (`RET`, "Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally `JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
carries the traditional conditional-jump spellings (`JZ`, `JNZ`, `JA`, `JC`, carries the traditional conditional-jump spellings (`JZ`, `JNZ`, `JA`, `JC`,
...) that the assembler accepts as aliases. Regenerating the tables is one ...) that the assembler accepts as aliases. The tables are generated from
command (`just gen`) and requires only a Go installation; the committed output the Go toolchain's own assembler source (`just gen` refreshes them), so
has no runtime dependency on the toolchain. every mnemonic the real assembler accepts is recognised; what the encoder
can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 108 of 627 files (17.2 %)
assembling for every target architecture today, with the top failure
reasons per architecture; the number moves with every release.
## Direction
The plan, in the order it is being worked:
- **Extended instruction support.** Two layers. First, encoding
coverage for every mnemonic the Go toolchain itself accepts, closed in
order of how often real code needs each instruction;
`gasm audit-instructions` measures the gap. Second, the larger work:
an extended instruction set the toolchain does not know at all. The
toolchain-derived tables stay generated and untouched; only the
extended instructions are hand-maintained, with their own spellings
and encoders, verified by execution on real hardware because the
toolchain offers no ground truth to compare against. The gaps exist
on every architecture, amd64 included.
- **Full GOOBJ and ELF compilation.** The destination is a complete,
standalone compilation path: linkable ELF objects for consumers outside
Go, and GOOBJ objects that `go build` links directly. Through GOOBJ, a
Go program will be able to use machine instructions that the Go
toolchain itself does not support; through ELF, Plan 9 assembly becomes
usable outside Go entirely.
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
architectures and is where the binary builds. FreeBSD follows: the
JIT's executable-memory mapping and the ptrace debugger layer are the
two pieces of porting work. Other unix systems may follow those two.
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
No others are planned.
## Install ## Install
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
linux/loong64 are on the linux/loong64 are on the
[releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases). [releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases).
From source (Go 1.27 or later): From source (Go 1.27.1):
```sh ```sh
go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest
``` ```
Or from a repository checkout, with the development version stamped: Or from a repository checkout:
```sh ```sh
just install-bin just install
``` ```
The installed binary reports the version the toolchain recorded: the tag
on a tagged checkout, a pseudo-version naming the commit below one.
## Quick start ## Quick start
```sh ```sh
@@ -142,9 +228,9 @@ infers the target architecture from the file-name suffix
## Development ## Development
```sh ```sh
just install # download module dependencies just build # compile, zero errors and zero warnings
just build # go vet + gofmt check, zero errors and zero warnings just test # the suite, no cache, the 80 % coverage floor
just test # full suite, race detector, 80 % coverage gate just gates # build, fmt-check, vet, test, race: the definition of done
just fmt # gofmt the tree just fmt # gofmt the tree
just gen # regenerate the instruction tables from the Go toolchain just gen # regenerate the instruction tables from the Go toolchain
``` ```
@@ -155,14 +241,16 @@ recipe.
## Documentation ## Documentation
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/CLI.md](docs/CLI.md): full command reference - [docs/CLI.md](docs/CLI.md): full command reference
- man pages: `just install-man` installs gasm(1) and one page per command
into ~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes - [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [docs/DECISIONS.md](docs/DECISIONS.md): deferred design decisions
- [CHANGELOG.md](CHANGELOG.md): release history - [CHANGELOG.md](CHANGELOG.md): release history
## Licence ## Licence
BSD-3-Clause — see [LICENSE](LICENSE). BSD-3-Clause; see [LICENSE](LICENSE).
Copyright © 2026 [Petr Balvín](https://petrbalvin.org) Copyright © 2026 [Petr Balvín](https://petrbalvin.org)
+40
View File
@@ -0,0 +1,40 @@
# Security policy
## Supported versions
Security fixes go to the newest release and to the `development` branch. Older
releases do not receive them.
| Version | Supported |
|---|---|
| 0.33.0 | yes |
| older releases | no |
## Reporting a vulnerability
**Do not open a public issue for a security problem.** A public report tells everyone
about the flaw before there is a fix. Report it privately to
**opensource@petrbalvin.org**.
Include:
- the version or commit you tested, and the platform
- what the problem is, and what an attacker gains from it
- the smallest reproducer you have, ideally a test or a single command
- a suggested fix, if you have one
## What to expect
- A human reads the report, and you get an acknowledgement.
- You are kept informed while the fix is being made, and told when it ships.
- The fix is released before the details are published, and the timing is agreed with
you.
- The reporter is credited in the release notes unless they ask otherwise.
## Out of scope
- Findings that require the attacker to already run code as the user, or to have local
access.
- Missing hardening with no demonstrated impact.
- Flaws in a third-party dependency: report them to that project, and to this one only
when this project's use of it makes them reachable.
+36
View File
@@ -236,6 +236,11 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
return encodeARM64BranchCond(mnem, enc.op, ops, pc, offsets, resolve) return encodeARM64BranchCond(mnem, enc.op, ops, pc, offsets, resolve)
} }
// Unconditional register branches (BR, BLR).
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FUncondBranch {
return encodeARM64RegBranch(mnem, enc.op, ops)
}
// ADD/SUB immediate. // ADD/SUB immediate.
if mnem == "ADD" || mnem == "ADDW" || mnem == "SUB" || mnem == "SUBW" || if mnem == "ADD" || mnem == "ADDW" || mnem == "SUB" || mnem == "SUBW" ||
mnem == "CMP" || mnem == "CMPW" || mnem == "CMN" || mnem == "CMNW" { mnem == "CMP" || mnem == "CMPW" || mnem == "CMN" || mnem == "CMNW" {
@@ -337,6 +342,24 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
} }
op := ops[0] op := ops[0]
// Register-indirect: JMP (R0) is BR R0, CALL (R0) is BLR R0. The
// toolchain's spelling carries no offset and no index; anything else
// is reported rather than silently dropped.
if op.Addr.Sym == nil && op.Addr.Base != "" {
if op.Addr.Offset != 0 || op.Addr.Index != "" {
return nil, fmt.Errorf("%s: invalid indirect branch operand %q", mnem, op.Raw)
}
rn := arm64RegNum(op.Addr.Base)
if rn < 0 {
return nil, fmt.Errorf("%s: unknown branch register %q", mnem, op.Addr.Base)
}
opc := uint32(0) // BR
if link {
opc = 1 // BLR
}
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a // Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
// relocation (R_CALLARM64 either way). // relocation (R_CALLARM64 either way).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" { if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
@@ -373,6 +396,19 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
return a64wordLE(a64Branch(bop, int32(rel))), nil return a64wordLE(a64Branch(bop, int32(rel))), nil
} }
// encodeARM64RegBranch encodes BR/BLR through a register operand:
// BR Xn = 0xd61f0000 | Rn<<5, BLR Xn = 0xd63f0000 | Rn<<5.
func encodeARM64RegBranch(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 1 {
return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(ops))
}
rn := arm64RegNum(operandRegName(ops[0]))
if rn < 0 {
return nil, fmt.Errorf("%s expects a register operand", mnem)
}
return a64wordLE(uint32(baseOp) | 31<<16 | uint32(rn)<<5), nil
}
// encodeARM64BranchCond encodes a conditional branch (B.cond) to a label. // encodeARM64BranchCond encodes a conditional branch (B.cond) to a label.
func encodeARM64BranchCond(mnem string, baseOp uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) { func encodeARM64BranchCond(mnem string, baseOp uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
if len(ops) != 1 { if len(ops) != 1 {
+39
View File
@@ -572,3 +572,42 @@ func leWords(b []byte) []uint32 {
} }
return w return w
} }
// TestArm64IndirectBranch pins the indirect branch forms in a leaf function:
// JMP (Rn) lowers to BR Rn, matching the toolchain's spelling, and the raw
// BR/BLR mnemonics encode directly (a gasm superset the toolchain's front
// end does not accept). CALL (Rn) shares the BLR path and its non-leaf
// prologue parity is covered by the ground-truth kernel.
func TestArm64IndirectBranch(t *testing.T) {
src := `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (R0)
BR R5
BLR R6
RET
`
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
want := []uint32{
0xd61f0000, // BR R0
0xd61f00a0, // BR R5
0xd63f00c0, // BLR R6
0xd65f03c0, // RET (BR LR)
}
got := leWords(img.Code)
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
+66 -1
View File
@@ -337,7 +337,18 @@ func computeFrame(t *ast.Text) frameInfo {
if t.Frame != nil && t.Frame.Imm.HasVal { if t.Frame != nil && t.Frame.Imm.HasVal {
fi.size = int(t.Frame.Imm.Val) fi.size = int(t.Frame.Imm.Val)
} }
if fi.size > 0 { if fi.size == 0 && hasCall(t) {
// The toolchain gives a frameless function containing a CALL an
// 8-byte frame for the pushed base pointer: the prologue saves BP
// with no stack adjustment, every RET pops it back, FP references
// pass one extra slot, and the virtual SP is the hardware SP.
fi.size = 8
fi.useFP = true
fi.fpAdjust = int64(fi.size) + 16 // return address + saved BP + args base
fi.spAdjust = 0
fi.prologue = []byte{0x55, 0x48, 0x89, 0xE5} // PUSHQ BP; MOVQ SP, BP
fi.epilogue = []byte{0x5D} // POPQ BP
} else if fi.size > 0 {
fi.useFP = true fi.useFP = true
fi.fpAdjust = int64(fi.size) + 16 // frame + saved BP + return address fi.fpAdjust = int64(fi.size) + 16 // frame + saved BP + return address
fi.spAdjust = int64(fi.size) fi.spAdjust = int64(fi.size)
@@ -518,6 +529,13 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) { if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
return 5, nil // opcode + rel32, always the long form return 5, nil // opcode + rel32, always the long form
} }
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
code, err := encodeIndirectJump(s, mnem)
if err != nil {
return 0, err
}
return len(code), nil
}
return jumpSize(mnem, long), nil return jumpSize(mnem, long), nil
} }
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link) code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
@@ -585,6 +603,15 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
} }
return append(prefix, code...), ps, nil return append(prefix, code...), ps, nil
} }
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
}
return append(prefix, code...), nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve) code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
} else { } else {
code, ps, err = encodeNormal(s, fi, link) code, ps, err = encodeNormal(s, fi, link)
@@ -706,6 +733,44 @@ func labelName(op *ast.Operand) (string, bool) {
return "", false return "", false
} }
// indirectJumpTarget reports whether the JMP/CALL operand addresses a
// register or a memory location rather than a label or a static symbol.
// A bare identifier is a register when the register table knows the name and
// a label otherwise, which is exactly how the parser cannot distinguish them.
func indirectJumpTarget(s *ast.Instr) bool {
if len(s.Operands) != 1 || s.Operands[0].Kind != ast.OpAddr {
return false
}
a := s.Operands[0].Addr
if a.Base != "" || a.Index != "" {
return true
}
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if _, ok := ParseReg(a.Sym.Name); ok {
return true
}
}
return false
}
// encodeIndirectJump assembles a JMP/CALL through a register or memory
// operand, which carries no relocation and no label to resolve.
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encodeIndirectBranch(mnem, ops); err != nil {
return nil, err
}
return e.out, nil
}
// spReg is the hardware stack pointer used to realise FP/SP pseudo-operands. // spReg is the hardware stack pointer used to realise FP/SP pseudo-operands.
var spReg = Reg{idx: 4, size: 8} var spReg = Reg{idx: 4, size: 8}
+2 -2
View File
@@ -202,8 +202,8 @@ TEXT ·withframe(SB), NOSPLIT, $16-16
} }
// TestAssembleVexKernel assembles the horizontal-sum reduction the go-flac // TestAssembleVexKernel assembles the horizontal-sum reduction the go-flac
// kernels end with — exercising the VEX moves, shuffle and extract forms // kernels end with; exercising the VEX moves, shuffle and extract forms
// through the full parser → encoder path — and checks the output is // through the full parser → encoder path; and checks the output is
// byte-identical to the Go assembler's. // byte-identical to the Go assembler's.
func TestAssembleVexKernel(t *testing.T) { func TestAssembleVexKernel(t *testing.T) {
fn := firstText(t, ` fn := firstText(t, `
+1 -1
View File
@@ -52,7 +52,7 @@ func elfTestImage(t *testing.T) *Image {
} }
// TestAssembleFileExternals checks that a reference to a symbol no GLOBL // TestAssembleFileExternals checks that a reference to a symbol no GLOBL
// defines is recorded as an external relocation instead of failing — the // defines is recorded as an external relocation instead of failing; the
// raw image leaves the displacement zero, the object emitters carry it. // raw image leaves the displacement zero, the object emitters carry it.
func TestAssembleFileExternals(t *testing.T) { func TestAssembleFileExternals(t *testing.T) {
img := elfTestImage(t) img := elfTestImage(t)
+14 -4
View File
@@ -40,10 +40,20 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeRet() return e.encodeRet()
case upper == "NOP": case upper == "NOP":
return e.emit(&instr{opcode: []byte{0x90}, modrm: -1, sib: -1}) return e.emit(&instr{opcode: []byte{0x90}, modrm: -1, sib: -1})
case upper == "CALL": case upper == "CALL" || upper == "JMP":
return e.encodeJmpRel(ops, []byte{0xE8}) // Through a register or memory: FF /2 (CALL) or FF /4 (JMP).
case upper == "JMP": // Anything else is a rel32 against a label resolved by the assembler.
return e.encodeJmpRel(ops, []byte{0xE9}) if len(ops) == 1 {
switch ops[0].(type) {
case Reg, Mem:
return e.encodeIndirectBranch(upper, ops)
}
}
opcode := []byte{0xE8}
if upper == "JMP" {
opcode = []byte{0xE9}
}
return e.encodeJmpRel(ops, opcode)
} }
if cc, ok := condCode(upper); ok { if cc, ok := condCode(upper); ok {
return e.encodeJcc(cc, ops) return e.encodeJcc(cc, ops)
+35 -2
View File
@@ -181,6 +181,39 @@ func TestControl(t *testing.T) {
checkOp(t, x86asm.JBE, "JLS", Imm(0)) checkOp(t, x86asm.JBE, "JLS", Imm(0))
} }
// TestIndirectControlFlow pins the indirect JMP/CALL forms: FF /4 for JMP and
// FF /2 for CALL through a register or memory. A REX appears only for the
// extended registers, never REX.W: the branch operand size is fixed at 64
// bits in long mode.
func TestIndirectControlFlow(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"JMP AX", "JMP", []Operand{AX}, "ffe0"},
{"CALL AX", "CALL", []Operand{AX}, "ffd0"},
{"JMP (BX)", "JMP", []Operand{Ptr(BX, 0, 8)}, "ff23"},
{"CALL (BX)", "CALL", []Operand{Ptr(BX, 0, 8)}, "ff13"},
{"JMP 8(BX)", "JMP", []Operand{Ptr(BX, 8, 8)}, "ff6308"},
{"CALL -16(BX)", "CALL", []Operand{Ptr(BX, -16, 8)}, "ff53f0"},
{"JMP R8", "JMP", []Operand{Reg{idx: 8, size: 2}}, "41ffe0"},
{"CALL R9", "CALL", []Operand{Reg{idx: 9, size: 2}}, "41ffd1"},
{"JMP R15", "JMP", []Operand{Reg{idx: 15, size: 2}}, "41ffe7"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestSSEMoveGroundTruth checks the legacy (non-VEX) SSE moves byte for byte // TestSSEMoveGroundTruth checks the legacy (non-VEX) SSE moves byte for byte
// against the Go assembler. wantOp is the decoder's name, which differs from // against the Go assembler. wantOp is the decoder's name, which differs from
// the Plan 9 spelling for the octa moves (MOVOU = MOVDQU, MOVO = MOVDQA). // the Plan 9 spelling for the octa moves (MOVOU = MOVDQU, MOVO = MOVDQA).
@@ -230,7 +263,7 @@ func TestSSEMoveGroundTruth(t *testing.T) {
// TestGoFlacScalarTail encodes the scalar tail of an analyze kernel to confirm // TestGoFlacScalarTail encodes the scalar tail of an analyze kernel to confirm
// the encoder handles a realistic instruction sequence. // the encoder handles a realistic instruction sequence.
func TestGoFlacScalarTail(t *testing.T) { func TestGoFlacScalarTail(t *testing.T) {
// MOVQ swin_base+0(FP), SI — modelled as MOVQ disp(reg), reg. // MOVQ swin_base+0(FP), SI; modelled as MOVQ disp(reg), reg.
checkSyntax(t, "mov rsi, qword ptr [rax+0x10]", "MOVQ", Ptr(AX, 0x10, 8), SI) checkSyntax(t, "mov rsi, qword ptr [rax+0x10]", "MOVQ", Ptr(AX, 0x10, 8), SI)
checkSyntax(t, "lea r9, ptr [rsi+4*rbx]", "LEAQ", Idx(SI, BX, 4, 0, 8), Reg{idx: 9, size: 8}) checkSyntax(t, "lea r9, ptr [rsi+4*rbx]", "LEAQ", Idx(SI, BX, 4, 0, 8), Reg{idx: 9, size: 8})
checkSyntax(t, "and r10, -0x8", "ANDQ", Imm(-8), Reg{idx: 10, size: 8}) checkSyntax(t, "and r10, -0x8", "ANDQ", Imm(-8), Reg{idx: 10, size: 8})
@@ -421,7 +454,7 @@ func TestSSEShuffleGroundTruth(t *testing.T) {
} }
// TestMOVQXMMGroundTruth pins the SSE2 packed-quadword move encodings: // TestMOVQXMMGroundTruth pins the SSE2 packed-quadword move encodings:
// loads and register moves on F3 0F 7E, stores on 66 0F D6 — the forms // loads and register moves on F3 0F 7E, stores on 66 0F D6; the forms
// the GPR-move fallback silently corrupted. // the GPR-move fallback silently corrupted.
func TestMOVQXMMGroundTruth(t *testing.T) { func TestMOVQXMMGroundTruth(t *testing.T) {
cases := []struct { cases := []struct {
+13 -13
View File
@@ -16,7 +16,7 @@ import (
// kernels use: NDS arithmetic, immediate and variable shifts, shuffles with // kernels use: NDS arithmetic, immediate and variable shifts, shuffles with
// an immediate, lane extracts, narrowing stores, broadcasts from a GPR or // an immediate, lane extracts, narrowing stores, broadcasts from a GPR or
// memory, mask destinations, mask moves, disp8×N compression and the 5-bit // memory, mask destinations, mask moves, disp8×N compression and the 5-bit
// register fields (X/Y 16–31, Z 0–31). // register fields (X/Y 16-31, Z 0-31).
func TestEvexGroundTruth(t *testing.T) { func TestEvexGroundTruth(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
@@ -58,7 +58,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VMOVDQU32 16(SI)(R15*4),Z4", "VMOVDQU32", []Operand{Idx(SI, vreg(t, "R15"), 4, 16, 64), vreg(t, "Z4")}, "62b17e486fa4be10000000"}, {"VMOVDQU32 16(SI)(R15*4),Z4", "VMOVDQU32", []Operand{Idx(SI, vreg(t, "R15"), 4, 16, 64), vreg(t, "Z4")}, "62b17e486fa4be10000000"},
{"VMOVDQU32 Z0,4(SI)(AX*1)", "VMOVDQU32", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f17e487f840604000000"}, {"VMOVDQU32 Z0,4(SI)(AX*1)", "VMOVDQU32", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f17e487f840604000000"},
{"VMOVDQU32 Z3,(DI)(R15*4)", "VMOVDQU32", []Operand{vreg(t, "Z3"), Idx(DI, vreg(t, "R15"), 4, 0, 64)}, "62b17e487f1cbf"}, {"VMOVDQU32 Z3,(DI)(R15*4)", "VMOVDQU32", []Operand{vreg(t, "Z3"), Idx(DI, vreg(t, "R15"), 4, 0, 64)}, "62b17e487f1cbf"},
// VMOVDQU64 — the W1 qword variant. // VMOVDQU64; the W1 qword variant.
{"VMOVDQU64 (SI)(R15*4),Z3", "VMOVDQU64", []Operand{Idx(SI, vreg(t, "R15"), 4, 0, 64), vreg(t, "Z3")}, "62b1fe486f1cbe"}, {"VMOVDQU64 (SI)(R15*4),Z3", "VMOVDQU64", []Operand{Idx(SI, vreg(t, "R15"), 4, 0, 64), vreg(t, "Z3")}, "62b1fe486f1cbe"},
{"VMOVDQU64 Z0,4(SI)(AX*1)", "VMOVDQU64", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f1fe487f840604000000"}, {"VMOVDQU64 Z0,4(SI)(AX*1)", "VMOVDQU64", []Operand{vreg(t, "Z0"), Idx(SI, AX, 1, 4, 64)}, "62f1fe487f840604000000"},
{"VMOVDQU64 Z1,Z2", "VMOVDQU64", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1fe487fca"}, {"VMOVDQU64 Z1,Z2", "VMOVDQU64", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1fe487fca"},
@@ -77,7 +77,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VPSHUFB Z1,Z2,Z3", "VPSHUFB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d4800d9"}, {"VPSHUFB Z1,Z2,Z3", "VPSHUFB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d4800d9"},
{"VMOVDQU8 Z1,Z2", "VMOVDQU8", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f17f487fca"}, {"VMOVDQU8 Z1,Z2", "VMOVDQU8", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f17f487fca"},
{"VMOVDQU16 Z1,Z2", "VMOVDQU16", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff487fca"}, {"VMOVDQU16 Z1,Z2", "VMOVDQU16", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff487fca"},
// Indices 16–31: rm[4] rides in X̄ for register operands. // Indices 16-31: rm[4] rides in X̄ for register operands.
{"VPSHUFD $1,X16,X17", "VPSHUFD", []Operand{Imm(1), vreg(t, "X16"), vreg(t, "X17")}, "62a17d0870c801"}, {"VPSHUFD $1,X16,X17", "VPSHUFD", []Operand{Imm(1), vreg(t, "X16"), vreg(t, "X17")}, "62a17d0870c801"},
{"VMOVUPD (DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 0, 64), vreg(t, "Z14")}, "6271fd481037"}, {"VMOVUPD (DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 0, 64), vreg(t, "Z14")}, "6271fd481037"},
{"VMOVUPD 64(DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 64, 64), vreg(t, "Z14")}, "6271fd48107701"}, {"VMOVUPD 64(DI),Z14", "VMOVUPD", []Operand{Ptr(DI, 64, 64), vreg(t, "Z14")}, "6271fd48107701"},
@@ -96,7 +96,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VPBROADCASTD 4(SI),Z10", "VPBROADCASTD", []Operand{Ptr(SI, 4, 4), vreg(t, "Z10")}, "62727d48585601"}, {"VPBROADCASTD 4(SI),Z10", "VPBROADCASTD", []Operand{Ptr(SI, 4, 4), vreg(t, "Z10")}, "62727d48585601"},
{"VPBROADCASTQ R8,X31", "VPBROADCASTQ", []Operand{vreg(t, "R8"), vreg(t, "X31")}, "6242fd087cf8"}, {"VPBROADCASTQ R8,X31", "VPBROADCASTQ", []Operand{vreg(t, "R8"), vreg(t, "X31")}, "6242fd087cf8"},
{"VPBROADCASTQ AX,Z9", "VPBROADCASTQ", []Operand{AX, vreg(t, "Z9")}, "6272fd487cc8"}, {"VPBROADCASTQ AX,Z9", "VPBROADCASTQ", []Operand{AX, vreg(t, "Z9")}, "6272fd487cc8"},
// Register indices 16–31 exist only in EVEX encodings. // Register indices 16-31 exist only in EVEX encodings.
{"VPBROADCASTD AX,Y30", "VPBROADCASTD", []Operand{AX, vreg(t, "Y30")}, "62627d287cf0"}, {"VPBROADCASTD AX,Y30", "VPBROADCASTD", []Operand{AX, vreg(t, "Y30")}, "62627d287cf0"},
// Packed double arithmetic / unpack (EVEX forms carry W=1). // Packed double arithmetic / unpack (EVEX forms carry W=1).
{"VSUBPD Z1,Z2,Z3", "VSUBPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed485cd9"}, {"VSUBPD Z1,Z2,Z3", "VSUBPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed485cd9"},
@@ -107,7 +107,7 @@ func TestEvexGroundTruth(t *testing.T) {
{"VUNPCKHPD Z1,Z2,Z3", "VUNPCKHPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed4815d9"}, {"VUNPCKHPD Z1,Z2,Z3", "VUNPCKHPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed4815d9"},
{"VSUBPD 64(AX),Z1,Z2", "VSUBPD", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5485c5001"}, {"VSUBPD 64(AX),Z1,Z2", "VSUBPD", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5485c5001"},
{"VSUBPD Z17,Z18,Z19", "VSUBPD", []Operand{vreg(t, "Z17"), vreg(t, "Z18"), vreg(t, "Z19")}, "62a1ed405cd9"}, {"VSUBPD Z17,Z18,Z19", "VSUBPD", []Operand{vreg(t, "Z17"), vreg(t, "Z18"), vreg(t, "Z19")}, "62a1ed405cd9"},
// VMOVDDUP — duplicate the low double; disp8×N = 64 at 512 bits, and // VMOVDDUP; duplicate the low double; disp8×N = 64 at 512 bits, and
// X16/X17 force EVEX (the mod=11 rm[4] extension rides in X̄). // X16/X17 force EVEX (the mod=11 rm[4] extension rides in X̄).
{"VMOVDDUP Z1,Z2", "VMOVDDUP", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff4812d1"}, {"VMOVDDUP Z1,Z2", "VMOVDDUP", []Operand{vreg(t, "Z1"), vreg(t, "Z2")}, "62f1ff4812d1"},
{"VMOVDDUP 64(AX),Z1", "VMOVDDUP", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1")}, "62f1ff48124801"}, {"VMOVDDUP 64(AX),Z1", "VMOVDDUP", []Operand{Ptr(AX, 64, 64), vreg(t, "Z1")}, "62f1ff48124801"},
@@ -150,7 +150,7 @@ func TestEvexGroundTruth(t *testing.T) {
} }
} }
// TestEvexMasking checks the AVX-512 mask operand (K1–K7, placed freely among // TestEvexMasking checks the AVX-512 mask operand (K1-K7, placed freely among
// the operands) and the .Z zeroing suffix, byte for byte against the Go // the operands) and the .Z zeroing suffix, byte for byte against the Go
// assembler. // assembler.
func TestEvexMasking(t *testing.T) { func TestEvexMasking(t *testing.T) {
@@ -241,11 +241,11 @@ func TestEvexMasking(t *testing.T) {
} }
} }
// TestEvexExtendedGroundTruth covers the wider EVEX/AVX-512 set — ternary // TestEvexExtendedGroundTruth covers the wider EVEX/AVX-512 set; ternary
// logic, lane shuffles/inserts/extracts, compares with a K destination, // logic, lane shuffles/inserts/extracts, compares with a K destination,
// permutes, the wider integer families, expand/compress, broadcasts, // permutes, the wider integer families, expand/compress, broadcasts,
// rotates and word shifts, the opmask instructions, the EVEX suffixes // rotates and word shifts, the opmask instructions, the EVEX suffixes
// (rounding/SAE/broadcast) and the aligned/scalar moves — byte for byte // (rounding/SAE/broadcast) and the aligned/scalar moves; byte for byte
// against the Go assembler. // against the Go assembler.
func TestEvexExtendedGroundTruth(t *testing.T) { func TestEvexExtendedGroundTruth(t *testing.T) {
mem64 := func(base Reg) Operand { return Ptr(base, 0, 64) } mem64 := func(base Reg) Operand { return Ptr(base, 0, 64) }
@@ -275,7 +275,7 @@ func TestEvexExtendedGroundTruth(t *testing.T) {
{"VMULPD.RZ_SAE.Z", "VMULPD.RZ_SAE.Z", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f1edf959d9"}, {"VMULPD.RZ_SAE.Z", "VMULPD.RZ_SAE.Z", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f1edf959d9"},
{"VMAXPD.SAE", "VMAXPD.SAE", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed585fd9"}, {"VMAXPD.SAE", "VMAXPD.SAE", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f1ed585fd9"},
{"VADDPD.BCST", "VADDPD.BCST", []Operand{mem64(AX), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5585810"}, {"VADDPD.BCST", "VADDPD.BCST", []Operand{mem64(AX), vreg(t, "Z1"), vreg(t, "Z2")}, "62f1f5585810"},
// Packed single arithmetic (same opcodes, no mandatory prefix) — // Packed single arithmetic (same opcodes, no mandatory prefix);
// ZMM, YMM and XMM widths, rounding and broadcast. // ZMM, YMM and XMM widths, rounding and broadcast.
{"VADDPS", "VADDPS", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f16c4858d9"}, {"VADDPS", "VADDPS", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f16c4858d9"},
{"VMULPS", "VMULPS", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ec59d9"}, {"VMULPS", "VMULPS", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ec59d9"},
@@ -399,8 +399,8 @@ func TestEvexExtendedGroundTruth(t *testing.T) {
} }
// TestEvexHelperGroundTruth covers the floating-point helper and conversion // TestEvexHelperGroundTruth covers the floating-point helper and conversion
// tail of the EVEX set — reciprocals, rsqrt, getexp/getmant, scalef, // tail of the EVEX set; reciprocals, rsqrt, getexp/getmant, scalef,
// rndscale, reduce, fixupimm, range, fpclass, the remaining conversions — // rndscale, reduce, fixupimm, range, fpclass, the remaining conversions;
// plus gather/scatter with VSIB addressing, byte for byte against the Go // plus gather/scatter with VSIB addressing, byte for byte against the Go
// assembler. // assembler.
func TestEvexHelperGroundTruth(t *testing.T) { func TestEvexHelperGroundTruth(t *testing.T) {
@@ -502,9 +502,9 @@ func TestEvexHelperGroundTruth(t *testing.T) {
} }
// TestEvexGprGroundTruth covers the scalar conversions between vector and // TestEvexGprGroundTruth covers the scalar conversions between vector and
// general-purpose registers — the signed and truncated VCVT{,T}S{D,S}2SI // general-purpose registers; the signed and truncated VCVT{,T}S{D,S}2SI
// forms (VEX and EVEX), the unsigned EVEX-only forms, and the GPR-to-vector // forms (VEX and EVEX), the unsigned EVEX-only forms, and the GPR-to-vector
// VCVTSI2*/VCVTUSI2* forms with the preserved vector source in vvvv — byte // VCVTSI2*/VCVTUSI2* forms with the preserved vector source in vvvv; byte
// for byte against the Go assembler, including memory sources and extended // for byte against the Go assembler, including memory sources and extended
// GPRs. // GPRs.
func TestEvexGprGroundTruth(t *testing.T) { func TestEvexGprGroundTruth(t *testing.T) {
+4 -4
View File
@@ -158,8 +158,8 @@ DATA mask<>+8(SB)/8, $0x800f0e0d0c0b0a09
t.Errorf("funcinfo bytes %x", fi) t.Errorf("funcinfo bytes %x", fi)
} }
// The pc-value tables of addq (non-package indices 0–3, so global // The pc-value tables of addq (non-package indices 0-3, so global
// indices 7–10): pcsp a flat zero over the whole function, pcinline a // indices 7-10): pcsp a flat zero over the whole function, pcinline a
// flat -1, both with the pc delta in MinLC (1) units. // flat -1, both with the pc delta in MinLC (1) units.
pcsp := data[le.Uint32(didx[4*7:]):] pcsp := data[le.Uint32(didx[4*7:]):]
if got := pcsp[:3]; !bytes.Equal(got, []byte{0x02, 19, 0x00}) { if got := pcsp[:3]; !bytes.Equal(got, []byte{0x02, 19, 0x00}) {
@@ -294,7 +294,7 @@ TEXT ·framed(SB), NOSPLIT, $8-0
} }
for i := range wantPCs { for i := range wantPCs {
if pcs[i] != wantPCs[i] || vals[i] != wantVals[i] { if pcs[i] != wantPCs[i] || vals[i] != wantVals[i] {
t.Errorf("pcsp[%d] = (%d,%d), want (%d,%d) — all: %v %v", i, pcs[i], vals[i], wantPCs[i], wantVals[i], pcs, vals) t.Errorf("pcsp[%d] = (%d,%d), want (%d,%d); all: %v %v", i, pcs[i], vals[i], wantPCs[i], wantVals[i], pcs, vals)
} }
} }
// The last two steps unwind the epilogue to zero. // The last two steps unwind the epilogue to zero.
@@ -333,7 +333,7 @@ TEXT ·useext(SB), NOSPLIT, $0-8
// TestGOObjectLinkAndRun is the end-to-end check: assemble the test // TestGOObjectLinkAndRun is the end-to-end check: assemble the test
// functions to a GOOBJ, swap it into a go build in place of the toolchain's // functions to a GOOBJ, swap it into a go build in place of the toolchain's
// assembly object, link, and run — the output must match the baseline // assembly object, link, and run; the output must match the baseline
// binary the Go assembler produced. Skipped when no Go toolchain is // binary the Go assembler produced. Skipped when no Go toolchain is
// available. // available.
func TestGOObjectLinkAndRun(t *testing.T) { func TestGOObjectLinkAndRun(t *testing.T) {
+18
View File
@@ -575,6 +575,24 @@ func (e *enc) encodeJmpRel(ops []Operand, opcode []byte) error {
return e.emit(&instr{opcode: opcode, modrm: -1, sib: -1, imm: le32(int64(imm))}) return e.emit(&instr{opcode: opcode, modrm: -1, sib: -1, imm: le32(int64(imm))})
} }
// encodeIndirectBranch encodes JMP/CALL through a register or memory operand:
// FF /4 for JMP, FF /2 for CALL. The operand size is fixed at 64 bits in
// 64-bit mode, so no REX.W is emitted; a REX appears only for R8-R15 bases.
func (e *enc) encodeIndirectBranch(mnem string, ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 1 operand, got %d", mnem, len(ops))
}
digit := 4 // JMP r/m64
if mnem == "CALL" {
digit = 2 // CALL r/m64
}
i := &instr{opcode: []byte{0xFF}, modrm: -1, sib: -1}
if err := setRMDigit(i, digit, ops[0], 8); err != nil {
return err
}
return e.emit(i)
}
// condCode maps a Plan 9 conditional-jump mnemonic to its x86 condition code. // condCode maps a Plan 9 conditional-jump mnemonic to its x86 condition code.
func condCode(upper string) (int, bool) { func condCode(upper string) (int, bool) {
if len(upper) < 2 || upper[0] != 'J' || upper == "JMP" { if len(upper) < 2 || upper[0] != 'J' || upper == "JMP" {
+3 -3
View File
@@ -19,8 +19,8 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
// TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel — // TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel;
// all functions plus the file-local mask24 constant — and checks that every // all functions plus the file-local mask24 constant; and checks that every
// static-symbol load resolves to the right bytes in the image. // static-symbol load resolves to the right bytes in the image.
func TestAssembleGoFlacAVX2Kernel(t *testing.T) { func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx2_amd64.s" path := "../../go-libraries/go-flac/avx2_amd64.s"
@@ -81,7 +81,7 @@ func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
} }
// TestAssembleGoFlacAVX512Kernel assembles the whole production AVX-512 // TestAssembleGoFlacAVX512Kernel assembles the whole production AVX-512
// kernel — all functions plus the file-global idx16 constant — and checks // kernel, all functions plus the file-global idx16 constant, and checks
// that the static-symbol load resolves to the right bytes in the image. // that the static-symbol load resolves to the right bytes in the image.
func TestAssembleGoFlacAVX512Kernel(t *testing.T) { func TestAssembleGoFlacAVX512Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx512_amd64.s" path := "../../go-libraries/go-flac/avx512_amd64.s"
+2 -2
View File
@@ -94,9 +94,9 @@ DATA ·table<>+0(SB)/8, $0x1122334455667788
} }
// The debug_line program: LNE_set_address (the R_ADDR relocation // The debug_line program: LNE_set_address (the R_ADDR relocation
// carries the function address), then one row per line change — the // carries the function address), then one row per line change; the
// TEXT is on line 4 (a leading blank line precedes the include), the // TEXT is on line 4 (a leading blank line precedes the include), the
// instructions on lines 5–9 — an advance to the 20-byte end and an // instructions on lines 5-9; an advance to the 20-byte end and an
// end-of-sequence. // end-of-sequence.
linesOff := le.Uint32(dataIdx[4*2:]) linesOff := le.Uint32(dataIdx[4*2:])
lines := dataBlk[linesOff : linesOff+21] lines := dataBlk[linesOff : linesOff+21]
+77 -70
View File
@@ -440,83 +440,90 @@ type dataSym struct {
} }
// collectData gathers the file's static symbols (GLOBL) and their initial // collectData gathers the file's static symbols (GLOBL) and their initial
// contents (DATA) into byte buffers, in declaration order. // contents (DATA) into byte buffers. Two passes: the Plan 9 convention puts
// every DATA line before its symbol's GLOBL, so the symbols are registered
// before the initialisers are applied.
func collectData(f *ast.File) ([]dataSym, error) { func collectData(f *ast.File) ([]dataSym, error) {
index := map[string]int{} index := map[string]int{}
var syms []dataSym var syms []dataSym
for _, d := range f.Decls { for _, d := range f.Decls {
switch dd := d.(type) { gd, ok := d.(*ast.Globl)
case *ast.Globl: if !ok {
if dd.Name == nil || dd.Name.Pseudo != "SB" { continue
continue }
} if gd.Name == nil || gd.Name.Pseudo != "SB" {
name := dd.Name.Name continue
if _, dup := index[name]; dup { }
return nil, fmt.Errorf("duplicate GLOBL %q", name) name := gd.Name.Name
} if _, dup := index[name]; dup {
size := 0 return nil, fmt.Errorf("duplicate GLOBL %q", name)
if dd.Size != nil && dd.Size.Imm.HasVal { }
size = int(dd.Size.Imm.Val) size := 0
} if gd.Size != nil && gd.Size.Imm.HasVal {
index[name] = len(syms) size = int(gd.Size.Imm.Val)
ds := dataSym{ }
name: name, index[name] = len(syms)
pkg: dd.Name.Pkg, ds := dataSym{
buf: make([]byte, size), name: name,
size: size, pkg: gd.Name.Pkg,
static: dd.Name.Static, buf: make([]byte, size),
} size: size,
for _, f := range dd.Flags { static: gd.Name.Static,
switch f { }
case "RODATA": for _, f := range gd.Flags {
ds.rodata = true switch f {
case "DUPOK": case "RODATA":
ds.dupok = true ds.rodata = true
default: case "DUPOK":
// Legacy numeric flag constants (runtime/textflag.h): ds.dupok = true
// DUPOK is 2, RODATA is 8; combinations arrive as one default:
// number (e.g. 10 = RODATA|DUPOK). // Legacy numeric flag constants (runtime/textflag.h):
if n, err := strconv.Atoi(f); err == nil { // DUPOK is 2, RODATA is 8; combinations arrive as one
if n&2 != 0 { // number (e.g. 10 = RODATA|DUPOK).
ds.dupok = true if n, err := strconv.Atoi(f); err == nil {
} if n&2 != 0 {
if n&8 != 0 { ds.dupok = true
ds.rodata = true }
} if n&8 != 0 {
ds.rodata = true
} }
} }
} }
syms = append(syms, ds) }
syms = append(syms, ds)
case *ast.Data: }
if dd.Name == nil || dd.Name.Pseudo != "SB" { for _, d := range f.Decls {
continue dd, ok := d.(*ast.Data)
} if !ok {
i, ok := index[dd.Name.Name] continue
if !ok { }
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name) if dd.Name == nil || dd.Name.Pseudo != "SB" {
} continue
if dd.Value == nil || !dd.Value.Imm.HasVal { }
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name) i, ok := index[dd.Name.Name]
} if !ok {
w := dd.Width return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
switch w { }
case 1, 2, 4, 8: if dd.Value == nil || !dd.Value.Imm.HasVal {
default: return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w) }
} w := dd.Width
off := dd.Name.Offset switch w {
buf := syms[i].buf case 1, 2, 4, 8:
if off < 0 || off+int64(w) > int64(len(buf)) { default:
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf)) return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
} }
v := dd.Value.Imm.Val off := dd.Name.Offset
if dd.Value.Imm.Neg { buf := syms[i].buf
v = -v if off < 0 || off+int64(w) > int64(len(buf)) {
} return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
for j := range w { }
buf[off+int64(j)] = byte(v >> (8 * j)) v := dd.Value.Imm.Val
} if dd.Value.Imm.Neg {
v = -v
}
for j := range w {
buf[off+int64(j)] = byte(v >> (8 * j))
} }
} }
return syms, nil return syms, nil
+2 -2
View File
@@ -10,8 +10,8 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
// TestAssembleFileStaticData checks the whole-image layout — code, padding // TestAssembleFileStaticData checks the whole-image layout; code, padding
// and the data section — and that the RIP-relative displacements of static // and the data section; and that the RIP-relative displacements of static
// symbol loads resolve to the right bytes. // symbol loads resolve to the right bytes.
func TestAssembleFileStaticData(t *testing.T) { func TestAssembleFileStaticData(t *testing.T) {
f, errs := parser.Parse("d_amd64.s", ` f, errs := parser.Parse("d_amd64.s", `
+44
View File
@@ -6,6 +6,7 @@ package asm
import ( import (
"fmt" "fmt"
"math/bits" "math/bits"
"strconv"
"strings" "strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast" "sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -258,6 +259,9 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
// 16-bit branches (BEQ/BNE/BLT/BGE/BLTU/BGEU) and JIRL. // 16-bit branches (BEQ/BNE/BLT/BGE/BLTU/BGEU) and JIRL.
if op, ok := l64branchTable[mnem]; ok { if op, ok := l64branchTable[mnem]; ok {
if mnem == "JIRL" {
return encodeLOONG64Jirl(op, ops)
}
return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve) return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve)
} }
// Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ, // Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ,
@@ -554,6 +558,46 @@ func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[stri
return l64wordLE(l64bbl(opc, v)), nil return l64wordLE(l64bbl(opc, v)), nil
} }
// encodeLOONG64Jirl encodes the raw JIRL spelling, JIRL rd, rj, offset, the
// form the verify trampolines use. The (rj) indirect form without an offset
// is handled by encodeLOONG64Branch.
func encodeLOONG64Jirl(op uint32, ops []*ast.Operand) ([]byte, error) {
if len(ops) != 3 {
return nil, fmt.Errorf("JIRL expects 3 operands, got %d", len(ops))
}
rd := l64Reg(ops[0])
rj := l64Reg(ops[1])
if rd < 0 || rj < 0 {
return nil, fmt.Errorf("invalid register operand")
}
off, ok := l64offsetOperand(ops[2])
if !ok {
return nil, fmt.Errorf("JIRL expects an immediate offset, got %q", ops[2].Raw)
}
if (int64(off)<<16)>>16 != int64(off) {
return nil, fmt.Errorf("JIRL offset %d out of the 16-bit range", off)
}
return l64wordLE(l64irr16(op, int(off), rj, rd)), nil
}
// l64offsetOperand reads a bare numeric branch offset: an immediate ($n) or a
// plain number, which parses as an empty address carrying the digits in Raw.
func l64offsetOperand(op *ast.Operand) (int32, bool) {
if op.Imm.HasVal {
v := op.Imm.Val
if op.Imm.Neg {
v = -v
}
return int32(v), true
}
if op.Kind == ast.OpAddr && op.Addr.Sym == nil && op.Addr.Base == "" && op.Addr.Index == "" {
if v, err := strconv.ParseInt(op.Raw, 0, 64); err == nil {
return int32(v), true
}
}
return 0, false
}
// encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU): // encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU):
// INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain // INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain
// turns into the 21-bit BEQZ/BNEZ form when the register is the only operand. // turns into the 21-bit BEQZ/BNEZ form when the register is the only operand.
+37
View File
@@ -291,3 +291,40 @@ done:
t.Errorf("code = % x\nwant % x", code, want) t.Errorf("code = % x\nwant % x", code, want)
} }
} }
// TestLOONG64IndirectBranch pins the indirect branch encodings: JMP (Rj) and
// JAL (Rj) lower to jirl, and the raw JIRL spelling encodes the written
// offset (the Go loong64 assembler deletes raw JIRL instructions entirely,
// so this form is a gasm-only superset with faithful semantics).
func TestLOONG64IndirectBranch(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (R4)
JIRL R0, R4, 8
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x4C000080, // jirl r0, r4, 0
0x4C002080, // jirl r0, r4, 8
0x4C000020, // jirl r0, r1, 0 (RET)
)
// JAL (R5) links, so the toolchain gives the function its autosize-8
// prologue and epilogue around the call and the closing RET.
fn = firstTextLOONG64(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JAL (R5)
RET
`)
code = assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x29FFE061, // addi.d r1, r2, -8 (prologue)
0x02FFE063, // addi.d r3, r3, -8
0x29C00061, // st.d r1, r2, 0 (prologue saves RA)
0x4C0000A1, // jirl r1, r5, 0
0x28C00061, // ld.d r1, r2, 0 (epilogue restores RA)
0x02C02063, // addi.d r3, r3, 8
0x4C000020, // jirl r0, r1, 0 (RET)
)
}
+1 -1
View File
@@ -280,7 +280,7 @@ done:
// TestLOONG64_pcsp checks the stack-adjustment table of a framed function: // TestLOONG64_pcsp checks the stack-adjustment table of a framed function:
// the prologue raises the SP delta by autosize (in effect from the third // the prologue raises the SP delta by autosize (in effect from the third
// instruction) and the RET's epilogue restores it to zero, with the pc deltas // instruction) and the RET's epilogue restores it to zero, with the pc deltas
// in MinLC (4) units — byte-identical to `go tool asm`. // in MinLC (4) units; byte-identical to `go tool asm`.
func TestLOONG64_pcsp(t *testing.T) { func TestLOONG64_pcsp(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
+181 -12
View File
@@ -135,6 +135,43 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
return out, offsets, relocs, lines, spadj, nil return out, offsets, relocs, lines, spadj, nil
} }
// riscvImmAlias maps the R-type ALU mnemonics onto their I-type immediate
// forms: the toolchain accepts ADD $imm, rj, rd and emits addi. Applied
// whenever the first operand is an immediate.
var riscvImmAlias = map[string]string{
"ADD": "ADDI",
"ADDW": "ADDIW",
"AND": "ANDI",
"OR": "ORI",
"XOR": "XORI",
"SLL": "SLLI",
"SRL": "SRLI",
"SRA": "SRAI",
"SLLW": "SLLIW",
"SRLW": "SRLIW",
"SRAW": "SRAIW",
}
// riscvNormaliseImmAlias rewrites the mnemonic to its immediate form when the
// first operand is an immediate: the toolchain accepts ADD $imm, rj, rd and
// emits addi, and SUB $imm becomes addi with the negated immediate. The
// second result reports that negation; the operand itself is left untouched
// because several passes normalise the same instruction.
func riscvNormaliseImmAlias(mnem string, ops []*ast.Operand) (string, bool) {
if len(ops) >= 2 && isImmOperand(ops[0]) {
switch strings.ToUpper(mnem) {
case "SUB":
return "ADDI", true
case "SUBW":
return "ADDIW", true
}
if alias, ok := riscvImmAlias[strings.ToUpper(mnem)]; ok {
return alias, false
}
}
return mnem, false
}
// riscvInstrSize returns the encoded size in bytes of a RISC-V instruction. // riscvInstrSize returns the encoded size in bytes of a RISC-V instruction.
// Most instructions are 4 bytes; MOV with a large immediate and I-type // Most instructions are 4 bytes; MOV with a large immediate and I-type
// arithmetic with a large immediate expand to several (possibly compressed) // arithmetic with a large immediate expand to several (possibly compressed)
@@ -142,10 +179,12 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int { func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
mnem := instr.Mnemonic.Text mnem := instr.Mnemonic.Text
ops := instr.Operands ops := instr.Operands
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
if mnem == "RET" { if mnem == "RET" {
return len(riscvReturn(fi)) return len(riscvReturn(fi))
} }
if mnem == "MOV" && len(ops) == 2 { if strings.HasPrefix(mnem, "MOV") && len(ops) == 2 {
// MOV $sym(SB), rd → 8 bytes (AUIPC + ADDI). // MOV $sym(SB), rd → 8 bytes (AUIPC + ADDI).
if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil && ops[0].Imm.Sym.Pseudo == "SB" { if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil && ops[0].Imm.Sym.Pseudo == "SB" {
return 8 return 8
@@ -173,7 +212,11 @@ func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
} }
// I-type arithmetic with a large immediate expands to several instructions. // I-type arithmetic with a large immediate expands to several instructions.
if (mnem == "ADDI" || mnem == "ANDI" || mnem == "ORI" || mnem == "XORI") && len(ops) >= 1 && isImmOperand(ops[0]) { if (mnem == "ADDI" || mnem == "ANDI" || mnem == "ORI" || mnem == "XORI") && len(ops) >= 1 && isImmOperand(ops[0]) {
return riscvItypeImmediateSize(mnem, immFromOperand(ops[0])) imm := immFromOperand(ops[0])
if immNeg {
imm = -imm
}
return riscvItypeImmediateSize(mnem, imm)
} }
return 4 return 4
} }
@@ -192,6 +235,8 @@ func isBranchLike(mnem string) bool {
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) { func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) {
mnem := instr.Mnemonic.Text mnem := instr.Mnemonic.Text
ops := instr.Operands ops := instr.Operands
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
var word uint32 var word uint32
// Handle pseudo-instructions and special cases first. // Handle pseudo-instructions and special cases first.
@@ -208,6 +253,18 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
} }
op := ops[0] op := ops[0]
if op.Addr.Sym == nil || op.Addr.Sym.Pseudo != "SB" { if op.Addr.Sym == nil || op.Addr.Sym.Pseudo != "SB" {
// CALL (X5): an indirect call, the toolchain's JALR X1, 0(X5).
if op.Addr.Sym == nil && op.Addr.Base != "" {
if op.Addr.Offset != 0 || op.Addr.Index != "" {
return nil, fmt.Errorf("CALL: invalid indirect operand %q", op.Raw)
}
rs1 := riscvRegNum(op.Addr.Base)
if rs1 < 0 {
return nil, fmt.Errorf("CALL: unknown branch register %q", op.Addr.Base)
}
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
return nil, fmt.Errorf("CALL: local branch target is not supported (use CALL sym(SB))") return nil, fmt.Errorf("CALL: local branch target is not supported (use CALL sym(SB))")
} }
if relocs != nil { if relocs != nil {
@@ -229,6 +286,18 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
target = labelFromOperand(ops[0]) target = labelFromOperand(ops[0])
// JMP (X5): an indirect branch, the toolchain's JALR X0, 0(X5).
if ops[0].Addr.Sym == nil && ops[0].Addr.Base != "" {
if ops[0].Addr.Offset != 0 || ops[0].Addr.Index != "" {
return nil, fmt.Errorf("JMP: invalid indirect operand %q", ops[0].Raw)
}
rs1 := riscvRegNum(ops[0].Addr.Base)
if rs1 < 0 {
return nil, fmt.Errorf("JMP: unknown branch register %q", ops[0].Addr.Base)
}
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
} }
targetOff, ok := offsets[target] targetOff, ok := offsets[target]
if !ok { if !ok {
@@ -255,14 +324,50 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
// MOV is a pseudo-instruction that the Go assembler uses for loads, // MOV is a pseudo-instruction that the Go assembler uses for loads,
// stores, register moves and immediate loads. // stores, register moves and immediate loads. The width suffixes
case "MOV": // (MOVB/MOVH/MOVW and unsigned forms) select the access width, and
// MOVD/MOVF address the FP registers.
case "MOV", "MOVB", "MOVBU", "MOVH", "MOVHU", "MOVW", "MOVWU", "MOVF", "MOVD":
return encodeRISCVMov(instr, fi, relocs) return encodeRISCVMov(instr, fi, relocs)
// JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1). // JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1).
case "JALR": case "JALR":
return encodeRISCVJALR(instr, fi) return encodeRISCVJALR(instr, fi)
// Branch-zero pseudos: BEQZ/BNEZ compare against X0, and BLTZ/BGEZ/
// BLEZ/BGTZ reorder the register operands of BLT/BGE accordingly.
case "BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ":
if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
}
rs := regFromOperand(ops[0])
if rs < 0 {
return nil, fmt.Errorf("%s: invalid register", mnem)
}
target := labelFromOperand(ops[1])
targetOff, ok := offsets[target]
if !ok {
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
}
var enc riscvEnc
rs1, rs2 := rs, 0
switch mnem {
case "BEQZ":
enc = riscvEnc{0x63, 0x0, 0x00} // beq rs, x0
case "BNEZ":
enc = riscvEnc{0x63, 0x1, 0x00} // bne rs, x0
case "BLTZ":
enc = riscvEnc{0x63, 0x4, 0x00} // blt rs, x0
case "BGEZ":
enc = riscvEnc{0x63, 0x5, 0x00} // bge rs, x0
case "BLEZ":
enc, rs1, rs2 = riscvEnc{0x63, 0x5, 0x00}, 0, rs // bge x0, rs
case "BGTZ":
enc, rs1, rs2 = riscvEnc{0x63, 0x4, 0x00}, 0, rs // blt x0, rs
}
word = riscvBType(enc, rs1, rs2, int32(targetOff-pc))
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
// System instructions with no operands. // System instructions with no operands.
case "FENCE", "ECALL", "EBREAK": case "FENCE", "ECALL", "EBREAK":
enc, ok := riscvInstrTable[mnem] enc, ok := riscvInstrTable[mnem]
@@ -457,6 +562,9 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// two-operand form INSTR $imm, rd uses rd as the source. // two-operand form INSTR $imm, rd uses rd as the source.
case len(ops) == 3 && isITypeInstr(mnem): case len(ops) == 3 && isITypeInstr(mnem):
imm := immFromOperand(ops[0]) // immediate imm := immFromOperand(ops[0]) // immediate
if immNeg {
imm = -imm // SUB $imm arrived through the ADDI alias
}
rs1 := regFromOperand(ops[1]) // source register rs1 := regFromOperand(ops[1]) // source register
rd := regFromOperand(ops[2]) // destination rd := regFromOperand(ops[2]) // destination
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
@@ -466,6 +574,9 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
case len(ops) == 2 && isITypeInstr(mnem): case len(ops) == 2 && isITypeInstr(mnem):
imm := immFromOperand(ops[0]) imm := immFromOperand(ops[0])
if immNeg {
imm = -imm
}
rd := regFromOperand(ops[1]) rd := regFromOperand(ops[1])
if rd < 0 { if rd < 0 {
return nil, fmt.Errorf("invalid register in %s", mnem) return nil, fmt.Errorf("invalid register in %s", mnem)
@@ -602,7 +713,7 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV load: invalid operand") return nil, fmt.Errorf("MOV load: invalid operand")
} }
return riscvFrameMemOp(riscvEnc{0x03, 0x3, 0x00}, false, rd, rs1, off), nil return riscvFrameMemOp(riscvMovEnc(strings.ToUpper(instr.Mnemonic.Text), false), false, rd, rs1, off), nil
} }
// Register → memory (store). // Register → memory (store).
@@ -619,21 +730,70 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rs2 < 0 || rs1 < 0 { if rs2 < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV store: invalid operand") return nil, fmt.Errorf("MOV store: invalid operand")
} }
return riscvFrameMemOp(riscvEnc{0x23, 0x3, 0x00}, true, rs2, rs1, off), nil return riscvFrameMemOp(riscvMovEnc(strings.ToUpper(instr.Mnemonic.Text), true), true, rs2, rs1, off), nil
} }
// Register → register (ADDI $0, src, dst). // Register → register: MOVD/MOVF are FP moves (fsgnj with rs2 = rs1),
// everything else is ADDI $0, src, dst.
{ {
rs1 := regFromOperand(src) rs1 := regFromOperand(src)
rd := regFromOperand(dst) rd := regFromOperand(dst)
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("MOV: invalid register operand") return nil, fmt.Errorf("MOV: invalid register operand")
} }
mnem := strings.ToUpper(instr.Mnemonic.Text)
if mnem == "MOVD" || mnem == "MOVF" {
op := uint32(0x20000053) // FSGNJ.S
if mnem == "MOVD" {
op = 0x22000053 // FSGNJ.D
}
return wordLE(op | uint32(rs1)<<15 | uint32(rs1)<<20 | uint32(rd)<<7), nil
}
word := riscvIType(riscvEnc{0x13, 0x0, 0x00}, rd, rs1, 0) word := riscvIType(riscvEnc{0x13, 0x0, 0x00}, rd, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
} }
} }
// riscvMovEnc returns the load (store=false) or store (store=true) opcode for
// a MOV-family mnemonic: the suffix selects the access width, MOVD and MOVF
// select the FP load/store opcodes, and bare MOV is the 64-bit integer form.
func riscvMovEnc(mnem string, store bool) riscvEnc {
if store {
switch mnem {
case "MOVB":
return riscvEnc{0x23, 0x0, 0x00} // SB
case "MOVH":
return riscvEnc{0x23, 0x1, 0x00} // SH
case "MOVW":
return riscvEnc{0x23, 0x2, 0x00} // SW
case "MOVF":
return riscvEnc{0x27, 0x2, 0x00} // FSW
case "MOVD":
return riscvEnc{0x27, 0x3, 0x00} // FSD
}
return riscvEnc{0x23, 0x3, 0x00} // SD
}
switch mnem {
case "MOVB":
return riscvEnc{0x03, 0x0, 0x00} // LB
case "MOVBU":
return riscvEnc{0x03, 0x4, 0x00} // LBU
case "MOVH":
return riscvEnc{0x03, 0x1, 0x00} // LH
case "MOVHU":
return riscvEnc{0x03, 0x5, 0x00} // LHU
case "MOVW":
return riscvEnc{0x03, 0x2, 0x00} // LW
case "MOVWU":
return riscvEnc{0x03, 0x6, 0x00} // LWU
case "MOVF":
return riscvEnc{0x07, 0x2, 0x00} // FLW
case "MOVD":
return riscvEnc{0x07, 0x3, 0x00} // FLD
}
return riscvEnc{0x03, 0x3, 0x00} // LD
}
// riscvFrameMemOp encodes a register-relative load (store=false, I-type // riscvFrameMemOp encodes a register-relative load (store=false, I-type
// width 0x03) or store (store=true, S-type width 0x23) of the 64-bit width // width 0x03) or store (store=true, S-type width 0x23) of the 64-bit width
// at off(rs1). Offsets beyond the signed 12-bit range materialise the // at off(rs1). Offsets beyond the signed 12-bit range materialise the
@@ -886,25 +1046,34 @@ func word16(w uint16) []byte {
} }
// encodeRISCVJALR encodes the JALR indirect jump/call instruction. // encodeRISCVJALR encodes the JALR indirect jump/call instruction.
// Plan 9: JALR rs1, rd (2 regs) or JALR offset(rs1) (memory → rd=X1). // Plan 9: JALR rs1, rd (2 regs), JALR rd, offset(rs1) (the trampoline
// form), or JALR offset(rs1) (memory → rd=X1).
func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) { func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
ops := instr.Operands ops := instr.Operands
// JALR rd, offset(rs1): the memory operand's base is the jump-target
// register, not the destination.
if len(ops) == 2 && isMemOperand(ops[1]) {
rd := regFromOperand(ops[0])
rs1, imm := memFromOperandWithFrame(ops[1], fi)
if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid register operand")
}
return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, imm)), nil
}
if len(ops) == 2 { if len(ops) == 2 {
rs1 := regFromOperand(ops[0]) rs1 := regFromOperand(ops[0])
rd := regFromOperand(ops[1]) rd := regFromOperand(ops[1])
if rd < 0 || rs1 < 0 { if rd < 0 || rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid register operand") return nil, fmt.Errorf("JALR: invalid register operand")
} }
word := riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, 0) return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, rd, rs1, 0)), nil
return wordLE(word), nil
} }
if len(ops) == 1 { if len(ops) == 1 {
rs1, imm := memFromOperandWithFrame(ops[0], fi) rs1, imm := memFromOperandWithFrame(ops[0], fi)
if rs1 < 0 { if rs1 < 0 {
return nil, fmt.Errorf("JALR: invalid memory operand") return nil, fmt.Errorf("JALR: invalid memory operand")
} }
word := riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, imm) return wordLE(riscvIType(riscvEnc{0x67, 0x0, 0x00}, 1, rs1, imm)), nil
return wordLE(word), nil
} }
return nil, fmt.Errorf("JALR expects 1 or 2 operands, got %d", len(ops)) return nil, fmt.Errorf("JALR expects 1 or 2 operands, got %d", len(ops))
} }
+2
View File
@@ -328,6 +328,8 @@ var riscvCvtTable = map[string]riscvCvtEnc{
"FCVTSWU": {0x68, 0x1, 0x53}, // uint32 → float32 "FCVTSWU": {0x68, 0x1, 0x53}, // uint32 → float32
"FCVTSL": {0x68, 0x2, 0x53}, // int64 → float32 "FCVTSL": {0x68, 0x2, 0x53}, // int64 → float32
"FCVTSLU": {0x68, 0x3, 0x53}, // uint64 → float32 "FCVTSLU": {0x68, 0x3, 0x53}, // uint64 → float32
"FCLASSS": {0x70, 0x0, 0x53}, // classify float32 → GPR mask
"FCLASSD": {0x70, 0x0, 0x53}, // classify float64 → GPR mask
"FCVTDW": {0x69, 0x0, 0x53}, // int32 → float64 "FCVTDW": {0x69, 0x0, 0x53}, // int32 → float64
"FCVTDWU": {0x69, 0x1, 0x53}, // uint32 → float64 "FCVTDWU": {0x69, 0x1, 0x53}, // uint32 → float64
"FCVTDL": {0x69, 0x2, 0x53}, // int64 → float64 "FCVTDL": {0x69, 0x2, 0x53}, // int64 → float64
+22 -1
View File
@@ -289,7 +289,7 @@ TEXT ·cmp(SB), NOSPLIT, $0
} }
func TestRISCV_forwardBranch(t *testing.T) { func TestRISCV_forwardBranch(t *testing.T) {
// Forward label reference — must not fail. // Forward label reference; must not fail.
fn := firstTextRISCV(t, `#include "textflag.h" fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·fwd(SB), NOSPLIT, $0 TEXT ·fwd(SB), NOSPLIT, $0
ADDI $1, X10, X10 ADDI $1, X10, X10
@@ -761,3 +761,24 @@ sub:
t.Error("expected error for CALL to local label, got nil") t.Error("expected error for CALL to local label, got nil")
} }
} }
// TestRISCVIndirectBranch pins the indirect branch encodings: JMP (X5) is the
// toolchain's JALR X0, 0(X5), and the trampoline form JALR rd, offset(rs1)
// takes its destination from the first operand (regression: the base
// register was once read as the destination, silently jumping to X0).
func TestRISCVIndirectBranch(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0-0
JMP (X5)
JALR X0, 0(X6)
JALR X28, 0(X9)
RET
`)
code := assembleRISCVHelper(t, fn)
wantWords(t, code,
0x00028067, // jalr x0, 5(x0), 0
0x00030067, // jalr x0, 6(x0), 0
0x00048e67, // jalr x28, 9(x0), 0
0x00008067, // jalr x0, 1(x0), 0 (RET)
)
}
+10 -3
View File
@@ -93,12 +93,19 @@ func riscvIsLeaf(t *ast.Text) bool {
return false return false
} }
case "JALR": case "JALR":
// JALR rs1, rd, a call when rd is X1; JALR offset(rs1) always // JALR rd, offset(rs1) links when the destination register (the
// links to X1. // first operand) is X1; JALR rs1, rd links when the second
// register is X1; JALR offset(rs1) always links to X1.
if len(in.Operands) == 1 { if len(in.Operands) == 1 {
return false return false
} }
if len(in.Operands) >= 2 && regFromOperand(in.Operands[1]) == 1 { if isMemOperand(in.Operands[1]) {
if regFromOperand(in.Operands[0]) == 1 {
return false
}
continue
}
if regFromOperand(in.Operands[1]) == 1 {
return false return false
} }
} }
+5 -5
View File
@@ -173,7 +173,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VPMULLD Y1,Y2,Y3", "VPMULLD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d40d9", ""}, {"VPMULLD Y1,Y2,Y3", "VPMULLD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d40d9", ""},
{"VPUNPCKLDQ Y4,Y3,Y5", "VPUNPCKLDQ", []Operand{vreg(t, "Y4"), vreg(t, "Y3"), vreg(t, "Y5")}, "c5e562ec", ""}, {"VPUNPCKLDQ Y4,Y3,Y5", "VPUNPCKLDQ", []Operand{vreg(t, "Y4"), vreg(t, "Y3"), vreg(t, "Y5")}, "c5e562ec", ""},
{"VPERMD Y1,Y2,Y3", "VPERMD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d36d9", ""}, {"VPERMD Y1,Y2,Y3", "VPERMD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c4e26d36d9", ""},
// Floating point (packed and scalar) and FMA — same NDS form, the pp // Floating point (packed and scalar) and FMA; same NDS form, the pp
// bits and map select the operation. // bits and map select the operation.
{"VADDPD Y9,Y8,Y8", "VADDPD", []Operand{vreg(t, "Y9"), vreg(t, "Y8"), vreg(t, "Y8")}, "c4413d58c1", ""}, {"VADDPD Y9,Y8,Y8", "VADDPD", []Operand{vreg(t, "Y9"), vreg(t, "Y8"), vreg(t, "Y8")}, "c4413d58c1", ""},
{"VADDPD X1,X2,X3", "VADDPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e958d9", ""}, {"VADDPD X1,X2,X3", "VADDPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e958d9", ""},
@@ -217,7 +217,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VEXTRACTI128 $1,Y8,X9", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d39c101", ""}, {"VEXTRACTI128 $1,Y8,X9", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d39c101", ""},
{"VEXTRACTI128 $1,Y8,(DI)", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), Ptr(DI, 0, 16)}, "c4637d390701", ""}, {"VEXTRACTI128 $1,Y8,(DI)", "VEXTRACTI128", []Operand{Imm(1), vreg(t, "Y8"), Ptr(DI, 0, 16)}, "c4637d390701", ""},
{"VEXTRACTF128 $1,Y8,X9", "VEXTRACTF128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d19c101", ""}, {"VEXTRACTF128 $1,Y8,X9", "VEXTRACTF128", []Operand{Imm(1), vreg(t, "Y8"), vreg(t, "X9")}, "c4437d19c101", ""},
// Moves — each direction picks its own opcode and VEX.W. // Moves; each direction picks its own opcode and VEX.W.
{"VMOVDQU (SI),Y1", "VMOVDQU", []Operand{Ptr(SI, 0, 32), vreg(t, "Y1")}, "c5fe6f0e", ""}, {"VMOVDQU (SI),Y1", "VMOVDQU", []Operand{Ptr(SI, 0, 32), vreg(t, "Y1")}, "c5fe6f0e", ""},
{"VMOVDQU Y3,(DI)", "VMOVDQU", []Operand{vreg(t, "Y3"), Ptr(DI, 0, 32)}, "c5fe7f1f", ""}, {"VMOVDQU Y3,(DI)", "VMOVDQU", []Operand{vreg(t, "Y3"), Ptr(DI, 0, 32)}, "c5fe7f1f", ""},
{"VMOVDQU X1,X2", "VMOVDQU", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fa7fca", ""}, {"VMOVDQU X1,X2", "VMOVDQU", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fa7fca", ""},
@@ -234,7 +234,7 @@ func TestVexGroundTruth(t *testing.T) {
{"VMOVD AX,X0", "VMOVD", []Operand{AX, vreg(t, "X0")}, "c5f96ec0", ""}, {"VMOVD AX,X0", "VMOVD", []Operand{AX, vreg(t, "X0")}, "c5f96ec0", ""},
{"VMOVSD (SI),X8", "VMOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X8")}, "c57b1006", ""}, {"VMOVSD (SI),X8", "VMOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X8")}, "c57b1006", ""},
{"VMOVSD X8,(SI)", "VMOVSD", []Operand{vreg(t, "X8"), Ptr(SI, 0, 8)}, "c57b1106", ""}, {"VMOVSD X8,(SI)", "VMOVSD", []Operand{vreg(t, "X8"), Ptr(SI, 0, 8)}, "c57b1106", ""},
// Packed double arithmetic and unpack — the NDS form, the opcode // Packed double arithmetic and unpack; the NDS form, the opcode
// selects the operation. // selects the operation.
{"VSUBPD Y1,Y2,Y3", "VSUBPD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ed5cd9", ""}, {"VSUBPD Y1,Y2,Y3", "VSUBPD", []Operand{vreg(t, "Y1"), vreg(t, "Y2"), vreg(t, "Y3")}, "c5ed5cd9", ""},
{"VDIVPD X1,X2,X3", "VDIVPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e95ed9", ""}, {"VDIVPD X1,X2,X3", "VDIVPD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5e95ed9", ""},
@@ -255,12 +255,12 @@ func TestVexGroundTruth(t *testing.T) {
{"VMINSS X6,X7,X8", "VMINSS", []Operand{vreg(t, "X6"), vreg(t, "X7"), vreg(t, "X8")}, "c5425dc6", ""}, {"VMINSS X6,X7,X8", "VMINSS", []Operand{vreg(t, "X6"), vreg(t, "X7"), vreg(t, "X8")}, "c5425dc6", ""},
{"VMAXSS X1,X2,X3", "VMAXSS", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5ea5fd9", ""}, {"VMAXSS X1,X2,X3", "VMAXSS", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")}, "c5ea5fd9", ""},
{"VADDSD 8(AX),X1,X2", "VADDSD", []Operand{Ptr(AX, 8, 8), vreg(t, "X1"), vreg(t, "X2")}, "c5f3585008", ""}, {"VADDSD 8(AX),X1,X2", "VADDSD", []Operand{Ptr(AX, 8, 8), vreg(t, "X1"), vreg(t, "X2")}, "c5f3585008", ""},
// VMOVDDUP — duplicate the low double (reg=dst, rm=src, F2 pp). // VMOVDDUP; duplicate the low double (reg=dst, rm=src, F2 pp).
{"VMOVDDUP X1,X2", "VMOVDDUP", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fb12d1", ""}, {"VMOVDDUP X1,X2", "VMOVDDUP", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5fb12d1", ""},
{"VMOVDDUP Y1,Y2", "VMOVDDUP", []Operand{vreg(t, "Y1"), vreg(t, "Y2")}, "c5ff12d1", ""}, {"VMOVDDUP Y1,Y2", "VMOVDDUP", []Operand{vreg(t, "Y1"), vreg(t, "Y2")}, "c5ff12d1", ""},
{"VMOVDDUP 8(AX),X1", "VMOVDDUP", []Operand{Ptr(AX, 8, 8), vreg(t, "X1")}, "c5fb124808", ""}, {"VMOVDDUP 8(AX),X1", "VMOVDDUP", []Operand{Ptr(AX, 8, 8), vreg(t, "X1")}, "c5fb124808", ""},
// Conversions: DQ→PS (no prefix), PS→PD (Go emits it without the F3 // Conversions: DQ→PS (no prefix), PS→PD (Go emits it without the F3
// prefix — see the table comment), DQ→PD. // prefix; see the table comment), DQ→PD.
{"VCVTDQ2PS X1,X2", "VCVTDQ2PS", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85bd1", ""}, {"VCVTDQ2PS X1,X2", "VCVTDQ2PS", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85bd1", ""},
{"VCVTDQ2PS Y3,Y4", "VCVTDQ2PS", []Operand{vreg(t, "Y3"), vreg(t, "Y4")}, "c5fc5be3", ""}, {"VCVTDQ2PS Y3,Y4", "VCVTDQ2PS", []Operand{vreg(t, "Y3"), vreg(t, "Y4")}, "c5fc5be3", ""},
{"VCVTPS2PD X1,X2", "VCVTPS2PD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85ad1", ""}, {"VCVTPS2PD X1,X2", "VCVTPS2PD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "c5f85ad1", ""},
+206 -1
View File
@@ -37,17 +37,29 @@ import (
// construction and are excluded from the diff; the other architectures list // construction and are excluded from the diff; the other architectures list
// their conditional branches outright. // their conditional branches outright.
func cmdAuditInstructions(args []string) error { func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [amd64|arm64|riscv64|loong64]", ` fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj), known-but-unencodable names (the backlog) and go- gasm asm --format goobj), known-but-unencodable names (the backlog) and go-
only names (feature gaps). The Go side is probed black-box with a battery only names (feature gaps). The Go side is probed black-box with a battery
of bare mnemonics, so the audit tracks whatever toolchain `+"`go env GOROOT`"+` of bare mnemonics, so the audit tracks whatever toolchain `+"`go env GOROOT`"+`
provides. provides.
With --corpus the audit changes shape: it assembles every .s file under the
given directory (default GOROOT/src) with the gasm encoder only, no
toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
`) `)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return err return err
} }
if *corpus {
return cmdAuditCorpus(fs.Args())
}
archName := "amd64" archName := "amd64"
switch n := len(fs.Args()); { switch n := len(fs.Args()); {
case n > 1: case n > 1:
@@ -305,3 +317,196 @@ func gasmAssembles(a arch.Arch, name, shape string) bool {
func sanitize(name string) string { func sanitize(name string) string {
return strings.NewReplacer(".", "_", "$", "_").Replace(name) return strings.NewReplacer(".", "_", "$", "_").Replace(name)
} }
// --- corpus audit -----------------------------------------------------------
// corpusTarget is one architecture row of the corpus report.
type corpusTarget struct {
a arch.Arch
name string
}
// corpusTally accumulates one architecture's attempts over the corpus.
type corpusTally struct {
attempted int
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
}
func (t *corpusTally) fail(path, reason string) {
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
}
// cmdAuditCorpus implements audit-instructions --corpus.
func cmdAuditCorpus(args []string) error {
if len(args) > 1 {
return fmt.Errorf("audit-instructions --corpus takes at most one directory argument")
}
root := ""
if len(args) == 1 {
root = args[0]
} else {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
return fmt.Errorf("locate GOROOT: %w", err)
}
root = filepath.Join(strings.TrimSpace(string(out)), "src")
}
stats, err := runCorpusAudit(root)
if err != nil {
return err
}
printCorpusStats(stats)
return nil
}
// corpusStats is the outcome of one corpus audit run.
type corpusStats struct {
root string
files int
generic int // files attempted for all four architectures
full int // files that assembled for every target architecture
targets []corpusTarget
tallies []*corpusTally
}
// runCorpusAudit assembles every .s file under root and returns the stats.
func runCorpusAudit(root string) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
return nil, err
}
targets := []corpusTarget{
{arch.AMD64, "amd64"},
{arch.ARM64, "arm64"},
{arch.RISCV, "riscv64"},
{arch.LOONG64, "loong64"},
}
tallies := make([]*corpusTally, len(targets))
for i := range tallies {
tallies[i] = &corpusTally{reasons: map[string]int{}, example: map[string]string{}}
}
// full is the north-star number: a file counts when every architecture
// its name allows assembles it.
full, generic := 0, 0
for _, path := range files {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
for i, tg := range targets {
if tg.a == a {
wanted = append(wanted, i)
}
}
} else {
generic++
for i := range targets {
wanted = append(wanted, i)
}
}
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
continue
}
t.assembled++
}
if ok && len(wanted) > 0 {
full++
}
}
return &corpusStats{
root: root,
files: len(files),
generic: generic,
full: full,
targets: targets,
tallies: tallies,
}, nil
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures)\n", s.root, s.files, s.generic)
fmt.Printf(" assemble for every target architecture: %d (%.1f%%)\n", s.full, 100*float64(s.full)/float64(max(s.files, 1)))
for i, tg := range s.targets {
t := s.tallies[i]
fmt.Printf(" %s: %d/%d attempted\n", tg.name, t.assembled, t.attempted)
for _, r := range topReasons(t) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
}
}
// corpusReason buckets an assembly or parse failure for the histogram.
func corpusReason(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
return "instruction not encodable"
case strings.Contains(msg, "undefined label"):
return "undefined label"
case strings.Contains(msg, "undefined symbol"), strings.Contains(msg, "external symbol"), strings.Contains(msg, "file-level assembly"):
return "undefined symbol or external"
case strings.Contains(msg, "operand"), strings.Contains(msg, "operand form"):
return "unsupported operand form"
default:
return "other: " + firstLine(msg)
}
}
// topReasons returns at most five reasons, most frequent first.
func topReasons(t *corpusTally) []string {
type kv struct {
k string
n int
}
var kvs []kv
for k, n := range t.reasons {
kvs = append(kvs, kv{k, n})
}
slices.SortFunc(kvs, func(a, b kv) int { return b.n - a.n })
if len(kvs) > 5 {
kvs = kvs[:5]
}
out := make([]string, len(kvs))
for i, kv := range kvs {
out[i] = kv.k
}
return out
}
// firstLine returns the first line of an error message, truncated.
func firstLine(msg string) string {
if i := strings.IndexByte(msg, '\n'); i >= 0 {
msg = msg[:i]
}
if len(msg) > 80 {
msg = msg[:80]
}
return msg
}
+54 -20
View File
@@ -18,6 +18,7 @@ import (
"os/exec" "os/exec"
"path/filepath" "path/filepath"
"runtime" "runtime"
"runtime/debug"
"slices" "slices"
"sort" "sort"
"strconv" "strconv"
@@ -36,9 +37,16 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/verify" "sourcedock.dev/petrbalvin/gasm-devkit/verify"
) )
// version is the release version, stamped at build time via // version reports the release the toolchain recorded for this build: the
// -ldflags "-X main.version=…" (defaulting to the current release). // tag on a tag, a pseudo-version below one, and (devel) outside version
var version = "0.33.0" // control. Nothing is injected; the recorded value cannot go stale.
func version() string {
bi, ok := debug.ReadBuildInfo()
if !ok || bi.Main.Version == "" {
return "(devel)"
}
return bi.Main.Version
}
func main() { func main() {
if len(os.Args) < 2 { if len(os.Args) < 2 {
@@ -88,13 +96,13 @@ func main() {
} }
} }
// cmdVersion prints the release version. // cmdVersion prints the recorded version.
func cmdVersion() int { func cmdVersion() int {
fmt.Printf("gasm %s\n", version) fmt.Printf("gasm %s\n", version())
return 0 return 0
} }
// ANSI color helpers for terminal output. // ANSI colour helpers for terminal output.
const ( const (
colorReset = "\033[0m" colorReset = "\033[0m"
colorBold = "\033[1m" colorBold = "\033[1m"
@@ -103,7 +111,7 @@ const (
colorGray = "\033[90m" colorGray = "\033[90m"
) )
// isTTY reports whether the writer is a terminal (for color output). // isTTY reports whether the writer is a terminal (for colour output).
func isTTY(w io.Writer) bool { func isTTY(w io.Writer) bool {
if f, ok := w.(*os.File); ok { if f, ok := w.(*os.File); ok {
stat, _ := f.Stat() stat, _ := f.Stat()
@@ -119,7 +127,7 @@ func usage(w io.Writer) {
bold, cyan, yellow, gray, reset = colorBold, colorCyan, colorYellow, colorGray, colorReset bold, cyan, yellow, gray, reset = colorBold, colorCyan, colorYellow, colorGray, colorReset
} }
fmt.Fprintf(w, "%sgasm %s%s: developer tooling for Go's Plan 9 assembler (GAsm)%s\n\n", bold, version, reset, reset) fmt.Fprintf(w, "%sgasm %s%s: developer tooling for Go's Plan 9 assembler (GAsm)%s\n\n", bold, version(), reset, reset)
fmt.Fprintf(w, "gasm bundles a lexer, parser, formatter, linter, standalone assembler and\n") fmt.Fprintf(w, "gasm bundles a lexer, parser, formatter, linter, standalone assembler and\n")
fmt.Fprintf(w, "language server for Plan 9 assembly into one self-contained binary.\n\n") fmt.Fprintf(w, "language server for Plan 9 assembly into one self-contained binary.\n\n")
@@ -442,7 +450,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
`) `)
fs.Parse(args) fs.Parse(args)
srv := lsp.New(os.Stdin, os.Stdout) srv := lsp.New(os.Stdin, os.Stdout)
srv.SetVersion(version) srv.SetVersion(version())
if err := srv.Run(); err != nil { if err := srv.Run(); err != nil {
fmt.Fprintln(os.Stderr, "gasm lsp:", err) fmt.Fprintln(os.Stderr, "gasm lsp:", err)
return 1 return 1
@@ -451,7 +459,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
} }
func cmdAsm(args []string) int { func cmdAsm(args []string) int {
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-o out] <file>", ` fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>", `
Assemble FILE without the Go toolchain: every TEXT function is encoded to Assemble FILE without the Go toolchain: every TEXT function is encoded to
machine code and printed as a hex dump. Supported architectures: amd64 machine code and printed as a hex dump. Supported architectures: amd64
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP, (including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
@@ -469,13 +477,22 @@ requires -p, the package path, and the installed Go toolchain).
out := fs.String("o", "", "write the output to this file") out := fs.String("o", "", "write the output to this file")
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)") format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)") pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
fs.Parse(args) fs.Parse(args)
if fs.NArg() != 1 { if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-o out] <file>") fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>")
return 2 return 2
} }
path := fs.Arg(0) path := fs.Arg(0)
targetArch := arch.FromFilename(path) targetArch := arch.FromFilename(path)
if *archName != "" {
a, err := auditArch(*archName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
return 2
}
targetArch = a
}
src, err := readSource(path) src, err := readSource(path)
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "gasm:", err) fmt.Fprintln(os.Stderr, "gasm:", err)
@@ -494,8 +511,10 @@ requires -p, the package path, and the installed Go toolchain).
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err) fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1 return 1
} }
if len(img.Funcs) == 0 { if len(img.Funcs) == 0 && len(img.Data) == 0 {
fmt.Fprintln(os.Stderr, "gasm asm: no assemblable TEXT functions found") // A file with neither code nor data assembles to nothing, which is
// almost always a wrong architecture rather than an intent.
fmt.Fprintln(os.Stderr, "gasm asm: no assemblable TEXT functions or GLOBL data found")
return 1 return 1
} }
for _, fn := range img.Funcs { for _, fn := range img.Funcs {
@@ -586,7 +605,7 @@ requires -p, the package path, and the installed Go toolchain).
// cmdDiff compares the machine code of two assembly files. // cmdDiff compares the machine code of two assembly files.
func cmdDiff(args []string) int { func cmdDiff(args []string) int {
set := newCommand("diff", "gasm diff <file1.s> <file2.s>", ` set := newCommand("diff", "gasm diff [-GOARCH arch] <file1.s> <file2.s>", `
Compare the machine code produced by assembling two files. Compare the machine code produced by assembling two files.
Shows which functions differ and the byte-level differences. Shows which functions differ and the byte-level differences.
Useful for verifying that two implementations produce identical code, Useful for verifying that two implementations produce identical code,
@@ -596,12 +615,22 @@ Use --map to compare functions whose names differ between the files,
e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix. e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
`) `)
mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names") mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names")
archName := set.String("GOARCH", "", "target architecture for both files: amd64, arm64, riscv64 or loong64")
set.Parse(args) set.Parse(args)
if set.NArg() != 2 { if set.NArg() != 2 {
fmt.Fprintln(os.Stderr, "usage: gasm diff <file1.s> <file2.s>") fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>")
return 2 return 2
} }
path1, path2 := set.Arg(0), set.Arg(1) path1, path2 := set.Arg(0), set.Arg(1)
forced := arch.Unknown
if *archName != "" {
a, err := auditArch(*archName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %v\n", err)
return 2
}
forced = a
}
// Parse the name mapping (file1 name → file2 name). // Parse the name mapping (file1 name → file2 name).
nameMap := make(map[string]string) nameMap := make(map[string]string)
@@ -617,12 +646,12 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
} }
// Assemble both files. // Assemble both files.
img1, err := assemblePath(path1) img1, err := assemblePath(path1, forced)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err) fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err)
return 1 return 1
} }
img2, err := assemblePath(path2) img2, err := assemblePath(path2, forced)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err) fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err)
return 1 return 1
@@ -697,8 +726,9 @@ func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
} }
} }
// assemblePath reads, parses and assembles a file (used by cmdDiff). // assemblePath reads, parses and assembles a file (used by cmdDiff). A
func assemblePath(path string) (*asm.Image, error) { // non-Unknown forced architecture overrides the file-name suffix.
func assemblePath(path string, forced arch.Arch) (*asm.Image, error) {
src, err := readSource(path) src, err := readSource(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -710,7 +740,11 @@ func assemblePath(path string) (*asm.Image, error) {
if len(errs) > 0 { if len(errs) > 0 {
return nil, fmt.Errorf("parse errors") return nil, fmt.Errorf("parse errors")
} }
return assembleFile(arch.FromFilename(path), f) target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
} }
// printByteDiff shows the first few byte differences between two code blocks. // printByteDiff shows the first few byte differences between two code blocks.
+59 -3
View File
@@ -219,8 +219,9 @@ func TestCmdVersion(t *testing.T) {
if code != 0 { if code != 0 {
t.Fatalf("code = %d", code) t.Fatalf("code = %d", code)
} }
if !strings.Contains(out, version) { got := version()
t.Errorf("version output %q does not mention %q", out, version) if !strings.Contains(out, got) {
t.Errorf("version output %q does not mention %q", out, got)
} }
} }
@@ -274,7 +275,7 @@ func TestVerifySmokeCrashIsolation(t *testing.T) {
} }
if exitErr, ok := err.(*exec.ExitError); ok { if exitErr, ok := err.(*exec.ExitError); ok {
if ws, ok := exitErr.Sys().(syscall.WaitStatus); ok && ws.Signaled() { if ws, ok := exitErr.Sys().(syscall.WaitStatus); ok && ws.Signaled() {
t.Fatalf("verify died from %v — the crash was not isolated:\n%s", ws.Signal(), out) t.Fatalf("verify died from %v; the crash was not isolated:\n%s", ws.Signal(), out)
} }
} }
if !strings.Contains(string(out), "CRASH") { if !strings.Contains(string(out), "CRASH") {
@@ -292,3 +293,58 @@ func TestSweepCheckLines(t *testing.T) {
t.Errorf("sweepCheckLines = %q, want %q", got, want) t.Errorf("sweepCheckLines = %q, want %q", got, want)
} }
} }
// TestRunCorpusAudit drives the corpus audit over a small fixture tree: one
// suffixed amd64 file, one suffixed arm64 file whose body is not arm64, one
// generic file, and one file that does not parse.
func TestRunCorpusAudit(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("good_amd64.s", "#include \"textflag.h\"\nTEXT ·add(SB), NOSPLIT, $0-0\n\tMOVQ AX, BX\n\tRET\n")
write("bad_arm64.s", "#include \"textflag.h\"\nTEXT ·f(SB), NOSPLIT, $0-0\n\tMOVQ AX, BX\n\tRET\n")
write("generic.s", "#include \"textflag.h\"\nTEXT ·g(SB), NOSPLIT, $0-0\n\tRET\n")
write("broken.s", "#include \"textflag.h\"\nTEXT ·b(SB), NOSPLIT, $0-0\n\tJMP nowhere\n\tRET\n")
stats, err := runCorpusAudit(dir)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
if stats.files != 4 {
t.Errorf("files = %d, want 4", stats.files)
}
if stats.generic != 2 {
t.Errorf("generic = %d, want 2 (generic.s and broken.s)", stats.generic)
}
// good_amd64 and generic.s assemble everywhere they are attempted.
if stats.full != 2 {
t.Errorf("full = %d, want 2", stats.full)
}
get := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
// amd64: good_amd64 + generic.s + broken.s; the broken file fails to parse.
if a := get("amd64"); a.attempted != 3 || a.assembled != 2 {
t.Errorf("amd64 = %d/%d, want 2/3", a.assembled, a.attempted)
}
// arm64: bad_arm64 (MOVQ is not arm64) + generic.s + broken.s.
if a := get("arm64"); a.attempted != 3 || a.assembled != 1 {
t.Errorf("arm64 = %d/%d, want 1/3", a.assembled, a.attempted)
}
if r := get("amd64").reasons["instruction not encodable"]; r != 0 {
t.Errorf("amd64 unexpected unencodable reason: %d", r)
}
if r := get("arm64").reasons["instruction not encodable"]; r != 1 {
t.Errorf("arm64 unencodable reasons = %d, want 1", r)
}
}
+157
View File
@@ -0,0 +1,157 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"os/exec"
"path/filepath"
"regexp"
"strings"
"testing"
)
// TestManPagesTrackTheCLI builds the binary once, then compares every
// command's live `-h` output with its docs/man/gasm-<command>.1 page: the
// flag sets must agree both ways, and the page's SYNOPSIS line must carry
// the command's usage line. A flag or a usage change that skips the man
// page fails here, so the pages cannot drift from the binary.
func TestManPagesTrackTheCLI(t *testing.T) {
if testing.Short() {
t.Skip("builds the gasm binary")
}
bin := filepath.Join(t.TempDir(), "gasm")
if out, err := exec.Command("go", "build", "-o", bin, ".").CombinedOutput(); err != nil {
t.Fatalf("build gasm: %v\n%s", err, out)
}
for _, cmd := range []string{
"tokens", "parse", "fmt", "lint", "asm", "dis", "verify",
"debug", "diff", "profile", "audit-instructions", "scaffold", "lsp",
} {
t.Run(cmd, func(t *testing.T) {
raw, err := os.ReadFile(filepath.Join("..", "..", "docs", "man", "gasm-"+cmd+".1"))
if err != nil {
t.Fatalf("read man page: %v", err)
}
page := string(raw)
out, _ := exec.Command(bin, cmd, "-h").CombinedOutput()
help := string(out)
binFlags := helpFlags(help)
pageFlags := roffFlags(page)
for f := range binFlags {
if !pageFlags[f] {
t.Errorf("flag -%s is in the binary's help but missing from the man page", f)
}
}
for f := range pageFlags {
if !binFlags[f] {
t.Errorf("flag -%s is in the man page but the binary does not accept it", f)
}
}
want := helpUsage(help)
got := roffSynopsis(page)
if want != "" && got != want {
t.Errorf("SYNOPSIS drift:\n page: %s\nbinary: %s", got, want)
}
})
}
}
// helpFlags extracts the flag names from a `gasm <cmd> -h` output.
func helpFlags(help string) map[string]bool {
m := map[string]bool{}
inFlags := false
for line := range strings.SplitSeq(help, "\n") {
if strings.TrimRight(line, " \t") == "Flags:" {
inFlags = true
continue
}
if !inFlags {
continue
}
if !strings.HasPrefix(line, " -") {
continue
}
token := strings.FieldsFunc(strings.TrimLeft(line, " "), func(r rune) bool {
return r == ' ' || r == '\t'
})
if len(token) == 0 {
continue
}
m[strings.TrimLeft(token[0], "-")] = true
}
return m
}
var roffEscape = regexp.MustCompile(`\\f[BIRP]`)
// roffFlags extracts the flag names from a man page's OPTIONS section.
func roffFlags(page string) map[string]bool {
m := map[string]bool{}
inOptions := false
for line := range strings.SplitSeq(page, "\n") {
if strings.HasPrefix(line, ".SH ") {
inOptions = strings.HasPrefix(line, ".SH OPTIONS")
continue
}
if !inOptions {
continue
}
// Flag entries are written as either `.B \-flag` or `\fB\-flag`.
var body string
switch {
case strings.HasPrefix(line, `.B \-`):
body = line[3:]
case strings.HasPrefix(line, `\fB\-`):
body = line[1:]
default:
continue
}
name := roffEscape.ReplaceAllString(body, "")
name = strings.ReplaceAll(name, `\-`, "-")
name = strings.TrimSpace(name)
if i := strings.IndexAny(name, " \t"); i >= 0 {
name = name[:i]
}
m[strings.TrimLeft(name, "-")] = true
}
return m
}
// helpUsage returns the command's usage line without the "Usage: " prefix.
func helpUsage(help string) string {
for line := range strings.SplitSeq(help, "\n") {
if strings.HasPrefix(line, "Usage: ") {
return normaliseUsage(line[len("Usage: "):])
}
}
return ""
}
// roffSynopsis returns the page's SYNOPSIS usage line, unescaped.
func roffSynopsis(page string) string {
inSyn := false
for line := range strings.SplitSeq(page, "\n") {
if strings.HasPrefix(line, ".SH ") {
inSyn = strings.HasPrefix(line, ".SH SYNOPSIS")
continue
}
if !inSyn || !strings.HasPrefix(line, ".B ") {
continue
}
return normaliseUsage(strings.ReplaceAll(line[3:], `\-`, "-"))
}
return ""
}
// normaliseUsage flattens whitespace and drops the roff font escapes so that
// the binary's usage line and the page's SYNOPSIS line compare equal.
func normaliseUsage(s string) string {
s = roffEscape.ReplaceAllString(s, "")
return strings.Join(strings.Fields(s), " ")
}
+100 -14
View File
@@ -4,7 +4,9 @@ How gasm-devkit is put together and why.
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit) Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
## Design goals ## Overview
Three design goals shape everything below.
1. **A real AST, not a grammar hack.** The linter, analyser, assembler and 1. **A real AST, not a grammar hack.** The linter, analyser, assembler and
language server all need to *reason* about assembly, not just colour it. language server all need to *reason* about assembly, not just colour it.
@@ -20,10 +22,10 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
through two vendor-neutral interfaces: a CLI and an LSP server. No editor through two vendor-neutral interfaces: a CLI and an LSP server. No editor
owns the toolkit; the toolkit is offered to editors on standard terms. owns the toolkit; the toolkit is offered to editors on standard terms.
## Pipeline The components, and how data moves between them:
```mermaid ```mermaid
graph TD flowchart TD
SRC["source .s"] --> LEX["lexer<br/>token stream"] SRC["source .s"] --> LEX["lexer<br/>token stream"]
LEX --> PAR["parser<br/>AST + diagnostics"] LEX --> PAR["parser<br/>AST + diagnostics"]
LEX --> FMT["format<br/>re-space tokens"] LEX --> FMT["format<br/>re-space tokens"]
@@ -42,9 +44,34 @@ graph TD
The lexer is the shared foundation: the parser builds the AST from it, the The lexer is the shared foundation: the parser builds the AST from it, the
formatter re-spaces its tokens directly, and the language server uses it for formatter re-spaces its tokens directly, and the language server uses it for
semantic highlighting. semantic highlighting. The phases follow a dependency chain: Phase 1 (static
analysis) builds only on the AST, Phase 2 (the standalone assembler) emits
object code, and Phases 3 (dynamic analysis) and 4 (the debugger) both consume
the execution substrate that the assembler provides.
## Components ## Packages
| Package | Responsibility |
|---|---|
| `token` | token kinds and positions |
| `lexer` | hand-written scanner; permissive, and it never panics |
| `ast` | the typed syntax tree: declarations, lines, operands |
| `parser` | line-oriented parser producing the AST and its diagnostics |
| `arch` | register and instruction tables for the four architectures |
| `lint` | static checks over the AST |
| `format` | canonical formatter over the token stream |
| `lsp` | the language server |
| `asm` | standalone assembler: encoders, image layout, object emitters |
| `verify` | JIT execution, ABI checks, differential fuzzing |
| `debug` | interactive ptrace debugger |
| `cmd/gasm` | the CLI |
| `_gen` | rebuilds the `arch` tables from the Go toolchain source |
The boundaries matter as much as the responsibilities: `ast` records syntax
only, and whether a name is a register or a label is left to `arch`, so the
parser stays architecture-agnostic. `asm` and `verify` are the only packages
that touch machine code and executable memory, and `cmd/gasm` owns no logic
beyond flags and output.
### `token` and `lexer` ### `token` and `lexer`
@@ -134,7 +161,8 @@ Two deeper analyses sit on top of the AST:
- **`unreachable-code`.** Code after a `RET` and before the next label is - **`unreachable-code`.** Code after a `RET` and before the next label is
dead. The check is suppressed for any function whose reachability cannot be dead. The check is suppressed for any function whose reachability cannot be
decided statically: those using PC-relative jumps (`JMP 2(PC)`), decided statically: those using PC-relative jumps (`JMP 2(PC)`),
register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`, or a `JMP`/`CALL`
through a register or memory operand), or living in a
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator: file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator:
code after it is occasionally intentional metadata. code after it is occasionally intentional metadata.
- **`register-clobber` (register liveness).** The linter builds the function's - **`register-clobber` (register liveness).** The linter builds the function's
@@ -206,7 +234,7 @@ The standalone assembler (Phase 2). Its core is an amd64 instruction encoder:
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set, a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
with the Plan 9 operand order (source first) mapped onto the x86 encoding. with the Plan 9 operand order (source first) mapped onto the x86 encoding.
Every encoding is validated by decoding it again with `golang.org/x/arch`, the Every encoding is validated by decoding it again with `golang.org/x/arch`, the
one module dependency, used in tests only and never linked into the binary. one module dependency, which also backs the `gasm dis` listings.
A **RISC-V encoder** (Phase 5, RV64IMAFDC + RVC compression) encodes the full A **RISC-V encoder** (Phase 5, RV64IMAFDC + RVC compression) encodes the full
integer, atomic, float/double, FMA and CSR instruction sets with the MOV integer, atomic, float/double, FMA and CSR instruction sets with the MOV
@@ -383,8 +411,8 @@ the Go ABI fixes across calls (amd64 `BP`/`R14`, arm64 `R29`/`R28`, riscv64
raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the
saved registers before Go code resumes. riscv64 is validated end to saved registers before Go code resumes. riscv64 is validated end to
end under qemu-user emulation; arm64 shares the same stack convention and end under qemu-user emulation; arm64 shares the same stack convention and
fix; loong64 stays ground-truth-only until hardware validation (see fix; loong64 stays ground-truth-only until hardware validation.
docs/DECISIONS.md). `gasm verify` runs the JIT checks when the host `gasm verify` runs the JIT checks when the host
matches the kernel's architecture and the toolchain comparisons matches the kernel's architecture and the toolchain comparisons
elsewhere. elsewhere.
@@ -436,14 +464,72 @@ watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot
block it), and `--cover` runs to completion with a breakpoint on every block it), and `--cover` runs to completion with a breakpoint on every
label and reports which blocks executed. label and reports which blocks executed.
## Extension points ### Extending the toolkit
- **New architecture:** add an entry to the generator in `_gen`, run - **New architecture:** add an entry to the generator in `_gen`, run
`just gen`, and add a `buildXXX()` register file plus a case in `ForArch`. `just gen`, and add a `buildXXX()` register file plus a case in `ForArch`.
- **New lint rule:** add a function in `lint` and a rule-code constant. - **New lint rule:** add a function in `lint` and a rule-code constant.
- **New LSP feature:** add a method case in `dispatch` and a handler. - **New LSP feature:** add a method case in `dispatch` and a handler.
The phases follow a dependency chain. Phase 1 (static analysis) builds only on ## Data flow
the AST; Phase 2 (the standalone assembler) emits object code; Phases 3
(dynamic analysis) and 4 (the debugger) both consume the execution substrate The main operation, assembling one file:
that the assembler provides.
```mermaid
sequenceDiagram
participant User
participant CLI as gasm CLI
participant Parser as parser
participant Asm as asm
participant Go as go toolchain
User->>CLI: gasm asm --format goobj -p pkg -o k.o k_amd64.s
CLI->>Parser: Parse(path, src)
Parser-->>CLI: AST, diagnostics
CLI->>Asm: AssembleFile(AST)
Asm->>Asm: encode operands, settle label offsets, lay out data
Asm-->>CLI: Image, code and data and relocations
CLI->>Asm: GOObject(pkg, path)
Asm->>Go: go list -json -export, externals only
Go-->>Asm: package and symbol indices
Asm-->>CLI: Go object bytes
CLI-->>User: wrote N bytes to k.o
```
Errors are produced where the parse or the encoding fails and become values at
the CLI boundary: the parser returns a diagnostic list and never aborts a file,
`AssembleFile` returns an error, and `cmd/gasm` prints what it has to stderr
and returns a non-zero exit code. The formatter and the linter take the same
AST by a different route: `gasm fmt` re-spaces the token stream and `gasm lint`
walks the parsed file, so neither depends on an encoding.
## State and lifetime
- The analysis packages (`lexer`, `parser`, `format`, `lint`, `arch`) hold only
read-only lookup tables and no mutable state: every call allocates its own
tokens and AST, and any number of goroutines may read the `arch` tables.
- A `verify.Kernel` owns one executable mapping, which `Close` releases. The
JIT trampolines keep the Go stack pointer and the checked-call sentinels in
package globals, so a call is a process-wide, one-at-a-time operation. The
`gasm verify` sweeps therefore run each function in a child process, which
contains a crash and keeps the globals unshared.
- `lsp.Server` is long-lived: it runs a single read and dispatch loop over the
stream and touches its document store only from that loop, so one server
serves one connection.
- A `debug.Session` owns a traced child process and pins its goroutine to the
forking OS thread, because ptrace requests must stay on that thread.
## Dependencies
- **`golang.org/x/arch`** (v0.30.0) is the one module dependency: it is the
disassembler backend (`gasm dis` and the debugger's listings) and the source
of the register metadata the encoder consults (`asm/reg.go`, `asm/vex.go`).
The tests additionally decode through it to validate the encodings.
- **The Go toolchain**, as an oracle and never as a library: `go tool asm`
supplies the object preamble and the ground truth for `gasm verify
--ground-truth`, `go list -json -export` locates the archives of the packages
a GOOBJ object references, and `_gen` parses
`$GOROOT/src/cmd/internal/obj/<arch>/anames.go` to rebuild the tables.
- **Linux process interfaces** for the dynamic work: `mmap` and `mprotect` for
the JIT mapping, ptrace with `/proc/pid/mem` for the debugger. That is why
`verify` runs a JIT check only when the host architecture matches the
kernel's, and why `debug` is Linux-only.
+419 -162
View File
@@ -1,215 +1,472 @@
# CLI Reference # Command line
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit) The reference below is taken from the program's own `--help`. If the two disagree, the
program is right and this file is a defect.
`gasm` is a single binary with subcommands. Run `gasm --help` for an The same reference is installed as man pages: `just install-man` puts gasm(1) and one
overview, or `gasm <command> -h` for a command's usage and flags. page per command into ~/.local/share/man (`MANDIR` overrides), and a test compares each
page against the binary so the two cannot drift apart.
## Global Flags ## Synopsis
| Flag | Description | ```sh
|------|-------------| gasm [global flags] <command> [command flags] [arguments]
| `-h`, `--help` | Show help | ```
| `-V`, `--version` | Print the version |
## `gasm tokens <file>` ## Commands
Print the lexical token stream of FILE: position, token kind, and text, | Command | Purpose |
one token per line. FILE may be `-` to read standard input. |---|---|
| `tokens` | print the lexical token stream |
| `parse` | parse a file and report syntax errors |
| `fmt` | canonicalise the formatting of `.s` files |
| `lint` | run the static checks |
| `asm` | assemble `.s` files to machine code |
| `dis` | disassemble machine code or an assembled file |
| `verify` | JIT-assemble and run the dynamic checks |
| `debug` | interactive source-level debugger |
| `diff` | compare the machine code of two `.s` files |
| `profile` | show the basic-block structure of the functions |
| `audit-instructions` | diff the encoder against the toolchain's name table |
| `scaffold` | generate a differential test skeleton for a kernel |
| `lsp` | run the language server over stdio |
| `version` | print the version |
## `gasm parse <file>` ## tokens
Parse FILE and report syntax errors on stderr. On success, prints how ```text
many declarations and TEXT functions the file contains. Usage: gasm tokens <file>
```
## `gasm fmt [-w|-l|-d] [path...]` Print the lexical token stream of FILE: position, token kind and text, one
token per line. FILE may be `-` to read standard input.
Canonicalise the formatting of Plan 9 assembly sources: indentation, ```sh
operand spacing, per-function mnemonic alignment, and blank-line layout. gasm tokens hello_amd64.s
```
| Flag | Description | ```text
|------|-------------| 1:1 # "#"
| `-w` | Write result to the source file (default: print to stdout) | 1:2 IDENT "include"
| `-l` | List files whose formatting differs, one per line; write nothing | 1:10 STRING "\"textflag.h\""
| `-d` | Print a unified diff of the canonical formatting instead | ```
With no arguments, or with a directory argument, every `.s` file below ## parse
it is reformatted in place and the names of changed files are listed
(`go fmt` style). `.` and `_` directories are skipped.
## `gasm lint <file...>` ```text
Usage: gasm parse <file>
```
Run static checks and print diagnostics as Parse FILE and report syntax errors on stderr. On success, print how many
`file:line:col: severity: message [code]`. Exit status is non-zero when declarations and TEXT functions the file contains.
an error-severity diagnostic is found.
| Flag | Description | ```sh
|------|-------------| gasm parse hello_amd64.s
| `-disable` | Comma-separated rule codes to disable | ```
```text
hello_amd64.s: OK, 2 declarations, 1 functions
```
## fmt
```text
Usage: gasm fmt [-w|-l|-d] [path...]
```
| Flag | Default | Effect |
|---|---|---|
| `-w` | off | write the result back to the source file |
| `-l` | off | list the files whose formatting differs; write nothing |
| `-d` | off | print a unified diff of the canonical formatting instead |
`-l` and `-d` are mutually exclusive. With no arguments, or with a directory
argument, every `.s` file below it is reformatted in place and the names of the
changed files are listed, the way `go fmt` does; `.` and `_` directories are
skipped. Explicit file arguments print to stdout unless `-w` is given.
```sh
gasm fmt -l kernel_amd64.s
```
Empty output means every file is formatted, which is the shape a CI check
wants; `-d` shows what would change:
```sh
gasm fmt -d ugly_amd64.s
```
```text
--- ugly_amd64.s
+++ ugly_amd64.s
@@ -2,8 +2,8 @@
// func add(a, b int) int
TEXT ·add(SB), NOSPLIT, $0-24
- MOVQ a+0(FP), AX
- ADDQ b+8(FP), AX
+ MOVQ a+0(FP), AX
+ ADDQ b+8(FP), AX
```
## lint
```text
Usage: gasm lint <file...>
```
| Flag | Default | Effect |
|---|---|---|
| `-disable` | empty | comma-separated rule codes to disable |
Diagnostics are printed as `file:line:col: severity: message [code]`. The exit
status is non-zero when an error-severity diagnostic is found; warnings (the
register-clobber audit, for example) do not affect it.
Rules: `unknown-instruction`, `operand-count`, `undefined-label`, Rules: `unknown-instruction`, `operand-count`, `undefined-label`,
`duplicate-label`, `missing-ret`, `missing-textflag-include`, `duplicate-label`, `missing-ret`, `missing-textflag-include`,
`abi-argsize`, `unreachable-code`, `register-clobber`, `abi-argsize`, `unreachable-code`, `register-clobber`,
`funcdata-pcdata`, `unused-label`, `invalid-textflag`, `funcdata-pcdata`, `unused-label`, `invalid-textflag`,
`stack-imbalance`, `register-width-mismatch`, `abi0-register-args`, `stack-imbalance`, `register-width-mismatch`, `abi0-register-args`,
`nonportable-register-name` and `unencodable-instruction`. `nonportable-register-name`, `unencodable-instruction` and
`reserved-register-write`.
## `gasm asm [--format raw|elf|goobj] [-p pkg] [-o out] <file>` ```sh
gasm lint kernel_amd64.s
```
Assemble FILE to machine code (amd64, arm64, riscv64, loong64). ## asm
| Flag | Description | ```text
|------|-------------| Usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>
| `--format` | Output format: `raw` (default), `elf`, `goobj` | ```
| `-p` | Package path (required for `--format goobj`) |
| `-o` | Write output to file (default: hex dump to stdout) |
## `gasm dis [-a arch] <file>` | Flag | Default | Effect |
|---|---|---|
| `-format` | `raw` | output format: `raw` (concatenated image), `elf` or `goobj` (Go object) |
| `-p` | empty | package path for `--format goobj`, qualifying the exported symbols |
| `-GOARCH` | empty | target architecture: `amd64`, `arm64`, `riscv64` or `loong64`; overrides the file-name suffix |
| `-o` | empty | write the output to this file instead of a hex dump on stdout |
Disassemble machine code to instruction text (via `golang.org/x/arch`). Supported architectures: amd64 (VEX/AVX2 and EVEX/AVX-512 included), arm64,
riscv64 (RV64IMAFDC and RVC) and loong64, taken from the file's `_arch.s`
suffix or from `-GOARCH`, which is how files whose names carry no
recognisable suffix (most of GOROOT's, for example `cpu_x86.s`) are
assembled. `raw` concatenates the functions and the data section into one
self-consistent image; `elf` emits a relocatable object that links with the
system toolchain; `goobj` emits the Go toolchain's own object format, which
`cmd/link` consumes directly.
With a `.s` file, the file is assembled first and the listing follows the ```sh
real layout: one block per `TEXT` function, local labels printed at their gasm asm hello_amd64.s
offsets. The architecture comes from the file name suffix, or from `-a`. ```
With any other file, or `-` for standard input, the bytes are
disassembled linearly and `-a` selects the architecture (amd64, arm64,
riscv64 or loong64).
| Flag | Description | ```text
|------|-------------| add: 16 bytes
| `-a` | Architecture for raw input without a `_arch.s` name | 0000: 48 8b 44 24 08 48 03 44 24 10 48 89 44 24 18 c3
```
## `gasm verify [flags] <file.s>` ## dis
Assemble FILE, map it into executable memory, and run dynamic checks. ```text
Usage: gasm dis [-a arch] <file>
```
| Flag | Description | | Flag | Default | Effect |
|------|-------------| |---|---|---|
| `--ground-truth` | Compare machine code byte-for-byte against `go tool asm` | | `-a` | empty | architecture for raw input without a `_arch.s` name |
| `--fuzz` | Differential fuzz: JIT both gasm and go-tool-asm, compare outputs |
| `-n` | Fuzz iterations per function (default: 1000) |
| `--abi` | Run ABI-checking calls (sentinel registers + red zone) |
| `--abi-n` | Number of ABI check iterations with varied inputs (default: 100) |
| `--profile` | List basic-block structure per function |
| `--smoke` | Call each NOSPLIT function with zeroed args |
| `--call <func>` | Invoke a single function with `--buf` instead of the sweeps |
| `--buf <spec>` | Buffer spec for `--call`: `name:size:pattern[,name:size:pattern]` |
| `--args <spec>` | Scalar args for `--call`: `name=value[,name=value]` (decimal or `0x` hex) |
| `--repeat <n>` | Number of times to repeat a `--call` invocation (default: 1) |
| `--save-corpus <dir>` | With `--fuzz`: write each failing input to DIR as replayable JSON |
| `--replay <dir>` | Re-run saved corpus entries (JSON in DIR), one child process per entry |
The `--fuzz` mode runs each function in a subprocess; a partial function With a `.s` file the file is assembled first and the listing follows the real
(e.g. a decoder that faults on malformed input) is reported as layout: one block per `TEXT` function, local labels printed at their offsets.
`CRASH` without killing the parent. Use `--call` with `--buf` to invoke With any other file, or `-` for standard input, the bytes are disassembled
partial functions with valid data instead. linearly and `-a` selects the architecture (amd64, arm64, riscv64 or loong64).
The `--call` mode parses the `// func` signature, allocates the requested ```sh
buffers (`zero`, `ones`, `seq`, or a hex blob), builds the ABI0 argument gasm dis hello_amd64.s
block with buffer pointers/lengths/capacities at the matching parameter ```
offsets, and prints the arg block before and after the call, showing
return values and any output written to the buffers. Scalar parameters
are supplied with `--args` (decimal, or `0x` hex) at their ABI0 offsets.
The `--save-corpus` mode records the logical arguments (buffer contents and ```text
scalars, not raw pointers) of every failing fuzz input as JSON. `--replay` add: 16 bytes
rebuilds a live argument block from each entry and calls it in its own child 0000: 48 8b 44 24 08 mov rax, qword ptr [rsp+0x8]
process, reporting `OK`, `CRASH (reproduced)` or `FAIL` per entry and 0005: 48 03 44 24 10 add rax, qword ptr [rsp+0x10]
exiting non-zero when any entry fails. 000a: 48 89 44 24 18 mov qword ptr [rsp+0x18], rax
000f: c3 ret
```
## `gasm debug [--func <name>] [--buf spec] [--script file] <file.s>` ## verify
Interactive debugger for JIT-assembled functions (amd64, arm64, riscv64, ```text
loong64). Requires a compiled binary on `$PATH` (not `go run`). Usage: gasm verify [-smoke] [-abi] [-fuzz] [-ground-truth] [-profile] [-call] <file.s>
```
| Flag | Description | | Flag | Default | Effect |
|------|-------------| |---|---|---|
| `--func` | Function to debug (required) | | `--ground-truth` | off | compare the machine code byte-for-byte against `go tool asm` |
| `--buf` | Buffer spec: `name:size:pattern[,name:size:pattern]` | | `--fuzz` | off | differential fuzz against the `go tool asm` build |
| `--args <file>` | File containing the ABI0 argument block | | `-n` | 1000 | fuzz iterations per function |
| `--script <file>` | Run REPL commands from a file (one per line) and exit; `-` reads stdin | | `--abi` | off | ABI-checking calls: sentinel registers and a red-zone canary |
| `--timeout <dur>` | Kill the debuggee after this duration (e.g. `30s`); for headless `--script` runs | | `--abi-n` | 100 | ABI check iterations with varied inputs |
| `--cover` | Run to completion with a breakpoint on every instruction; report which executed, how often, and which labels were reached | | `--profile` | off | list the basic-block structure per function |
| `--smoke` | off | call each NOSPLIT function with zeroed arguments |
| `--call` | empty | invoke a single function with `--buf` instead of the sweeps |
| `--buf` | empty | buffer spec for `--call`: `name:size:pattern[,name:size:pattern]` |
| `--args` | empty | scalar args for `--call`: `name=value[,name=value]` (decimal or `0x` hex) |
| `--repeat` | 1 | number of times to repeat a `--call` invocation |
| `--save-corpus` | empty | with `--fuzz`: write each failing input to this directory as replayable JSON |
| `--replay` | empty | re-run saved corpus entries, one child process per entry |
The JIT checks run when the host matches the file's architecture; the
toolchain comparison works everywhere. `--fuzz`, `--smoke` and `--abi` run each
function in its own child process, so a partial function that faults on random
input is reported as `CRASH` instead of ending the sweep; `--call` with `--buf`
invokes such a function with valid data. loong64 stays on the ground-truth path
until hardware validation.
```sh
gasm verify --ground-truth hello_amd64.s
```
```text
hello_amd64.s: 1 functions JIT-loaded
add: MATCH (16 bytes)
ground truth: 1/1 functions byte-identical
add: 16 bytes, args=24, frame=0 NOSPLIT
```
```sh
gasm verify --call add --args a=2,b=3 hello_amd64.s
```
```text
add: 16 bytes, args=24
signature: func add(a int, b int) int
scalars:
a = 2
b = 3
args before: 02 00 00 00 00 00 00 00 03 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 (24 bytes)
args after: 02 00 00 00 00 00 00 00 03 00 00 00 00 00 00 00 05 00 00 00 00 00 00 00 (24 bytes)
call 1: OK
```
## debug
```text
Usage: gasm debug <file.s> --func <name>
```
| Flag | Default | Effect |
|---|---|---|
| `-func` | empty | the function to debug, required |
| `-buf` | empty | buffer spec: `name:size:pattern[,name:size:pattern]` (zero, ones, seq or hex) |
| `-args` | empty | file containing the ABI0 argument block |
| `-script` | empty | run REPL commands from a file, one per line, and exit; `-` reads stdin |
| `-timeout` | 0 | kill the debuggee after this duration, for headless `-script` runs |
| `-cover` | off | run to completion with a breakpoint on every instruction and report which executed |
The debugger spawns the debuggee from the `gasm` binary on `$PATH`, so install
it first with `just install`; `go run` does not work for the traced child.
Requires Linux (ptrace) and all four architectures are supported.
REPL commands: REPL commands:
| Command | Description | | Command | Effect |
|---------|-------------| |---|---|
| `break <label\|addr> [if <reg> <op> <val>]` | Set a breakpoint, optionally conditional | | `break <label\|addr> [if <reg> <op> <val>]` | set a breakpoint, optionally conditional |
| `delete <label\|addr>` | Remove a breakpoint | | `delete <label\|addr>` | remove a breakpoint |
| `info break` | List all breakpoints | | `info break` | list the breakpoints |
| `step [n]`, `s` | Single-step n instructions | | `step [n]`, `s` | single-step n instructions |
| `next`, `n` | Step over CALL | | `next`, `n` | step over a CALL |
| `finish`, `fin` | Run until the function returns | | `finish`, `fin` | run until the function returns |
| `continue`, `c` | Run until breakpoint, watchpoint or exit | | `continue`, `c` | run until a breakpoint, watchpoint or exit |
| `disas [n]`, `u` | Disassemble n instructions at PC | | `disas [n]`, `u` | disassemble n instructions at the PC |
| `regs` | Print general-purpose + vector/FP registers | | `regs` | print the general-purpose and vector/FP registers |
| `where` | Show source line and nearest label at PC | | `where` | show the source line and the nearest label at the PC |
| `stack` | Show stack near RSP (return address + ABI0 args) | | `stack` | show the stack near RSP, the return address and the ABI0 args |
| `bt`, `backtrace` | Backtrace (current frame + return address) | | `bt`, `backtrace` | backtrace: the current frame and the return address |
| `x [addr] [len]` | Hex-dump memory | | `x [addr] [len]` | hex-dump memory |
| `w <addr> <val...>` | Write bytes to memory | | `w <addr> <val...>` | write bytes to memory |
| `set <reg> <value>` | Set a register | | `set <reg> <value>` | set a register |
| `watch <addr> [r\|w] [size]` | Set a hardware watchpoint (write by default) | | `watch <addr> [r\|w] [size]` | set a hardware watchpoint, write by default |
| `unwatch [<slot>]` | Clear one or all watchpoints | | `unwatch [<slot>]` | clear one watchpoint or all of them |
| `labels`, `l` | List function labels and offsets | | `labels`, `l` | list the function's labels and offsets |
| `help`, `h`, `?` | Show command help | | `help`, `h`, `?` | show the command help |
| `quit`, `q` | Kill the debuggee and exit | | `quit`, `q` | kill the debuggee and exit |
## `gasm diff [--map old=new,...] <file1.s> <file2.s>` ```sh
gasm debug --func add --cover hello_amd64.s
```
Compare the machine code produced by assembling two files. Shows which ## diff
functions differ and the first few differing bytes. Useful for verifying
that two implementations produce identical code, or for tracking encoding
changes between Go assembler versions.
| Flag | Description | ```text
|------|-------------| Usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>
| `--map` | Comma-separated `old=new` pairs to match functions with different names | ```
Without `--map`, functions are paired by exact name. With `--map`, a | Flag | Default | Effect |
function named `old` in the first file is compared against the function |---|---|---|
named `new` in the second file (e.g. `--map wideCopyAVX2=wideCopyAVX512` | `-GOARCH` | empty | target architecture for both files, overriding the file-name suffixes |
pairs AVX2 and AVX-512 variants regardless of suffix). | `-map` | empty | comma-separated `old=new` pairs to match functions with different names |
## `gasm profile <file.s>` Functions are paired by exact name unless `--map` says otherwise, so
`--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless of suffix.
The exit status is non-zero when anything differs.
Show the basic-block structure of functions in an assembly file. Lists ```sh
each function's labels, their offsets, and the block boundaries. This is gasm diff hello_amd64.s hello_amd64.s
the static structure; for runtime execution counts, use `gasm verify ```
--fuzz` which exercises the code paths.
## `gasm audit-instructions [amd64|arm64|riscv64|loong64]` ```text
add: identical (16 bytes)
all functions identical
```
Compare the gasm encoder for the given architecture (default amd64) ## profile
against the installed `go tool asm` and print the diff: superset
encodings (gasm-only spellings, shippable via `gasm asm --format goobj`),
known-but-unencodable names (the encoder backlog) and go-only names
(feature gaps). The Go side is probed black-box with a battery of operand
shapes per mnemonic, so the audit tracks whatever toolchain
`go env GOROOT` provides. On non-amd64 architectures the backlog is an
over-approximation: a name counts as encodable only when a probe shape
assembles cleanly, so a name whose real forms the battery misses lands
in the backlog.
## `gasm scaffold differential <file.s>` ```text
Usage: gasm profile <file.s>
```
Print a differential test skeleton for every `// func` signature in Show the basic-block structure of each function: its labels, their offsets and
FILE. The generated test seeds random states, drives the kernel and a the block boundaries. This is the static structure; for runtime execution
portable reference (`<name>Portable`), and compares outputs counts use `gasm debug --cover`, and for input coverage `gasm verify --fuzz`.
byte-for-byte. Write the reference bodies, place the file in the
kernel's package, and run it in CI.
## `gasm lsp` ```sh
gasm profile hello_amd64.s
```
Run the language server over standard input/output (JSON-RPC 2.0 with ```text
Content-Length framing). Point an LSP-capable editor at the binary and add: 16 bytes, args=24, frame=0 NOSPLIT
associate it with `.s` files. The target architecture is inferred from basic blocks: 1
the file-name suffix (`_amd64.s`, `_arm64.s`, `_riscv64.s`, ```
`_loong64.s`).
Provides: completion, hover, document symbols, push and pull ## audit-instructions
diagnostics, semantic tokens, go-to-definition, find references, rename,
document formatting, inlay hints, code actions, signature help, document ```text
highlights, workspace symbol search, #include document links, and Usage: gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]
folding ranges for function bodies. ```
Compare the gasm encoder for the given architecture (default amd64) against the
installed `go tool asm` and print the diff: superset encodings (gasm-only
spellings, shippable via `gasm asm --format goobj`), known-but-unencodable
names (the encoder backlog) and go-only names (feature gaps). The Go side is
probed black-box with a battery of operand shapes per mnemonic, so the audit
tracks whatever toolchain `go env GOROOT` provides. On non-amd64
architectures the backlog is an over-approximation: a name counts as encodable
only when a probe shape assembles cleanly, so a name whose real forms the
battery misses lands in the backlog.
```sh
gasm audit-instructions amd64
```
```text
gasm table (amd64, families excluded): 1542 mnemonics
gasm encodable: 580 go tool asm recognized: 1542
shared: 580
```
With `--corpus` the audit changes shape: it assembles every `.s` file under
DIR (default `GOROOT/src`) with the gasm encoder only, no toolchain probing.
A file whose name carries a recognisable `_arch` suffix is attempted for that
architecture; a file without one is attempted for all four, exactly as a
`GOARCH` build would compile it. The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
over GOROOT takes under a second.
```sh
gasm audit-instructions --corpus
gasm audit-instructions --corpus "$(go env GOROOT)/src/crypto"
```
```text
corpus /usr/local/go/src: 627 files (365 generic, attempted for all architectures)
assemble for every target architecture: 108 (17.2%)
amd64: 77/464 attempted
148 instruction not encodable
e.g. /usr/local/go/src/cmd/asm/internal/asm/testdata/386enc.s
...
```
## scaffold
```text
Usage: gasm scaffold differential <file.s>
```
Print a differential test skeleton for every `// func` signature in FILE. The
generated test seeds random states, drives the kernel and a portable reference
(`<name>Portable`), and compares the outputs byte-for-byte. Write the reference
bodies, place the file in the kernel's package, and run it in CI.
```sh
gasm scaffold differential kernel_amd64.s > kernel_differential_test.go
```
## lsp
```text
Usage: gasm lsp
```
Run the language server over standard input/output, JSON-RPC 2.0 with
`Content-Length` framing. Point an LSP-capable editor at the binary and
associate it with `.s` files; the target architecture is inferred from the
file-name suffix (`_amd64.s`, `_arm64.s`, `_riscv64.s`, `_loong64.s`).
Provides: completion, hover, document symbols, push and pull diagnostics,
semantic tokens, go-to-definition, find references, rename, document
formatting, inlay hints, code actions, signature help, document highlights,
workspace symbol search, #include document links, and folding ranges for
function bodies. Definition, references and rename work across every open
document.
## version
```text
Usage: gasm version
```
Print the version the toolchain recorded for the build, the same string as
`gasm --version`: the tag on a tagged checkout, a pseudo-version naming the
commit below one, with `+dirty` appended on a dirty tree and `(devel)` outside
version control.
## Global flags
| Flag | Default | Effect |
|---|---|---|
| `-h`, `--help` | off | print the usage |
| `-V`, `--version` | off | print the version |
## Exit codes
| Code | Meaning |
|---|---|
| `0` | success |
| `1` | a failure the program detected: a parse or assembly error, an error-severity lint diagnostic, a mismatch in `verify`, a file that cannot be read |
| `2` | the arguments were wrong: a missing or extra argument, an unknown command or format, an invalid `--map` pair |
## Examples
Assemble a kernel, check it, and run it:
```sh
gasm lint kernel_amd64.s
gasm fmt -l kernel_amd64.s
gasm asm -o kernel.bin kernel_amd64.s
gasm verify --ground-truth kernel_amd64.s
```
Link the kernel into a Go program through the toolchain's own object format:
```sh
gasm asm --format goobj -p example.com/kernel -o kernel.o kernel_amd64.s
```
Find which labels a failing kernel reaches, headlessly:
```sh
gasm debug --func decodeBlockAVX2 --cover --script cmds.txt --timeout 30s kernel_amd64.s
```
-111
View File
@@ -1,111 +0,0 @@
# Deferred decisions
Design decisions deliberately postponed, with enough context to pick them up
again without re-deriving the analysis. Each entry records what is deferred,
why, the options on the table, and the trigger that should reopen it.
---
## GOOBJ external (cross-package) symbol references
**Status:** resolved (v0.29.0+, 2026-08-07).
**Approach taken.** Instead of parsing the compiler's iexport data (which
would have required either `golang.org/x/tools` or an in-house parser), the
resolver reads the **GOOBJ data directly** from the target package's `.a`
archive. The `.a` file contains a `_go_.o` member whose GOOBJ s is the
same one gasm writes; the parser reuses the same layout (`blkSymdef`,
`blkNonpkgdef`, the string table), so no new dependency was needed.
**How it works.**
1. `go list -json -export <pkg>` finds the target package's `.a` file.
2. `extractGOOBJ` reads the ar archive, finds the `_go_.o` member, skips
the `"go object …\n!\n"` preamble and parses the GOOBJ header.
3. `goobjFile.symbols()` walks `blkSymdef` and `blkNonpkgdef` in definition
order (the same order the linker uses) to build the symbol-to-index
mapping.
4. `resolveExternalSymbols` wires the resolved `{PkgIdx, SymIdx}` into the
GOOBJ emission.
The resolver is invoked automatically when `img.Externals` is non-empty; it
runs `go list` as a subprocess (consistent with `toolchainObjectPreamble`
which already calls `go tool asm`). All symbol data is cached per package
for the lifetime of the GOOBJ emission.
## 2026-08-30 non-amd64 JIT execution trampolines
**Status:** resolved for riscv64 (validated end to end under qemu-user)
and arm64 (fix in place, consistent with the observed frame convention);
open for loong64 until hardware validation.
**Root cause (found 2026-08-31).** The trampolines advanced SP past the
leave-address slot after loading it, while the assembled kernels read
their first argument at SP+8 per the frame convention (the amd64 path
already kept SP on that slot). Removing the advance fixed riscv64
immediately (plain and checked ABI tests pass under qemu-user); the
arm64 kernel's pre-fix trace showed exactly the same SP+8 reading. The
apparent arm64/loong64 "crashes in the JIT" turned out to be dominated
by an unrelated instability: the Go 1.26 and 1.27 runtimes crash under
qemu-user arm64 emulation (GC worker start, identical signature with the
JIT tests skipped, both qemu 7.2 and 10.2), and the Go loong64 runtime
does not start at all. `gasm verify` therefore keeps loong64 kernels on
the ground-truth path until hardware validation; the GOARCH-guarded
tests (`verify/jit_arch_test.go`, `verify/abi_arch_test.go`) are the
hardware validation entry point.
**State.** The per-architecture trampolines compile for all targets, the
kernels they execute are byte-for-byte correct against `go tool asm`, and
under `qemu-aarch64` the arm64 kernel demonstrably executes and stores its
result correctly. The failure is on the return path into Go code: arm64
and loong64 take a SIGSEGV after the kernel's RET (the Go-side unwind
through `leaveJIT` and its interposed ABIInternal wrapper is the suspect),
and riscv64 returns cleanly but with an untouched result area. amd64 is
unaffected (the checked trampoline saves and restores BP/R14 and the flow
is validated end to end).
**Evidence harness.** `verify/jit_arch_test.go` (plain call) and
`verify/abi_arch_test.go` (checked call) are GOARCH-guarded tests; build
the test binary per target (`GOARCH=arm64 go test -c -o v.test ./verify/`)
and run it under `qemu-aarch64-static` from the `verify/` directory. A
minimal reproducer pattern lives in the qemu exploration notes: verify
loads, the kernel executes, the fault follows the return.
**Fix direction.** Compare the amd64 checked trampoline (GLOBL/DATA raw
address, explicit SP/BP/R14 save-restore) against the arm64/riscv64/
loong64 `leaveJIT` unwind, in particular the interaction with the
ABIInternal wrapper that `reflect.ValueOf(leaveJIT).Pointer()` returns.
The plain-call path (no sentinels) fails the same way, so the checked
path is not the variable.
---
## 2026-08-29 tooling round
- `lint abi0-register-args`: flags kernels whose `// func` parameters are
never read from the FP frame. Motivated by a real latent bug: kernels
reading arguments from registers pass every test while the autogenerated
`F.abi0` wrapper happens to leave the caller's register values intact, and
break on a toolchain upgrade.
- `lint nonportable-register-name`: the RAX/EAX register spellings are a gasm
extension; go tool asm rejects them, so files using them only link through
the gasm goobj path.
- `lint unencodable-instruction`: a mnemonic in the architecture table that
`asm.Encodable` rejects is flagged at edit time instead of failing at
assembly time.
- `audit-instructions`: black-box diff of the encoder against go tool asm.
As of this round the tables fully overlap on names; the audit exists to
catch drift in both directions (future supersets and future gaps).
- `scaffold differential`: generates the direct-call differential skeleton
(two independent seed sets, output and in-place buffer comparison) that a
pipeline-level fuzz can never replace.
- `verify --args`: scalar arguments for `-call`, closing the repro gap where
only buffers could be supplied.
- `debug --script/--timeout/--cover`: headless debugging with a watchdog
armed before the ptrace attach (untracing sandboxes hang the attach), and
label-level block coverage for the "did my test ever enter that branch"
question.
- Superset policy remains: gasm may accept spellings and encodings go tool
asm lacks, but such kernels ship only via `gasm asm --format goobj`; the
audit reports the superset surface. The register-alias superset is warned
about by lint because the default `go build` path cannot consume it.
+88 -66
View File
@@ -4,79 +4,80 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
## Prerequisites ## Prerequisites
- **Go** 1.27+ with `toolchain go1.27.0` - **Go** 1.27.1, the exact version the `go` directive in `go.mod` declares
- **just**, the command runner; every task below is a just recipe - **just**, the command runner; every task below is a just recipe
- A Linux host on amd64, arm64, riscv64 or loong64: `gasm debug` needs ptrace
and the JIT checks of `gasm verify` need executable memory
- No external dependencies beyond the Go toolchain - No external dependencies beyond the Go toolchain
## Quick Start ## Setup
```sh ```sh
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
cd gasm-devkit cd gasm-devkit
just install # go mod download just build # compile bin/gasm, zero errors and zero warnings
just build # go vet + gofmt, must pass with zero output just gates # build, fmt-check, vet, test, race: the definition of done
just test # full suite, race detector, 80 % coverage gate
``` ```
## Just Recipes ## Recipes
### `just install` Every recipe in the `justfile`, and what it does.
`go mod download`. The only module dependency, `golang.org/x/arch`, is | Recipe | What it does |
used in tests only. |---|---|
| `just build` | compiles `bin/gasm` with `CGO_ENABLED=0` and stripped symbols; zero errors and zero warnings |
### `just build` | `just test` | the test gate: the suite with `-count=1`, the coverage profile and the 80 % floor |
| `just race` | the same suite under the race detector; the expensive one, so it runs once, inside `gates` |
Runs `go vet ./...` and checks `gofmt -l .` produces no output. This is | `just unit [packages] [run]` | fast, cached, scoped run for iterating: no race and no coverage, so an unchanged package reports instantly |
the minimum bar before any commit. | `just fuzz <target> <pkg> [fuzztime]` | time-boxed fuzz of one target; the package is required, because `go test -fuzz` refuses more than one |
| `just bench [packages]` | benchmarks (`-benchmem -count=5`); on an idle machine only |
| `just fmt` | formats the tree in place with `gofmt` |
| `just fmt-check` | zero diff; prints nothing when everything is formatted, which is the shape the CI step wants |
| `just vet` | both static gates: `go vet` and `go fix -diff` |
| `just gates` | `build`, `fmt-check`, `vet`, `test` and `race`, in that order: the definition of done |
| `just clean` | removes the build artefacts, `bin/` and `coverage.out` |
| `just install` | builds, then copies the binary into `bindir` (`~/.local/bin`); `gasm debug` needs an installed binary, because it spawns the debuggee from `$PATH` |
| `just uninstall` | removes the installed binary from `bindir` |
| `just run` | runs the CLI with `go run -buildvcs=true`; the recipe takes no arguments, so flags go through the package instead |
| `just dev` | the same as `run`; the project has no watcher to add |
| `just gen` | regenerates the `arch` instruction tables from the Go toolchain source; not a gate |
### `just test` ### `just test`
```sh ```sh
go test -race -count=1 ./... go test -count=1 -timeout 10m -coverprofile=coverage.out \
./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... \
./lint/... ./lsp/... ./parser/... ./token/... ./verify/...
``` ```
Plus a coverage run over the ten analysable packages (arch, asm, ast, The suite runs over the logic packages (`-count=1`, so no cached pass
format, lexer, lint, lsp, parser, token, verify; `debug` and `cmd/gasm` counts): arch, asm, ast, disasm, format, lexer, lint, lsp, parser,
need hardware or are CLI glue) and an `awk` gate that fails if total token, verify. `debug` traces a live process and `cmd/gasm` is thin CLI
coverage is below 80 %. glue, so both sit outside the sweep, and a thin `cmd/` in it would drag
the coverage total under the floor. The floor fails if the total is
below 80 %. CI runs the same command with the same ten-minute bound, so
the number is the same everywhere.
### `just fmt` ### `just run`
```sh ```sh
gofmt -w . just run
go run -buildvcs=true ./cmd/gasm lint kernel_amd64.s
go run -buildvcs=true ./cmd/gasm verify --ground-truth kernel_amd64.s
``` ```
Run after editing any Go source. The output must be idempotent. The flag on `go run` is there because it does not stamp the build otherwise,
which `--version` would then report as `(devel)`.
### `just run -- <args>`
Runs the CLI via `go run` with the version string stamped:
```sh
just run -- lint kernel_amd64.s
just run -- fmt -w kernel_amd64.s
just run -- verify --ground-truth kernel_amd64.s
```
### `just install-bin`
Installs the `gasm` binary into `$GOBIN` with the release version
embedded via `-ldflags "-X main.version=..."`.
### `just gen` ### `just gen`
Regenerates the architecture instruction tables in `arch/` by parsing Regenerates the architecture instruction tables in `arch/` by parsing the Go
the Go toolchain's own assembler source toolchain's own assembler source
(`$GOROOT/src/cmd/internal/obj/<arch>/anames.go`). Requires a Go (`$GOROOT/src/cmd/internal/obj/<arch>/anames.go`). Requires a Go
installation. Output is committed, with no runtime dependency on the installation. Output is committed, with no runtime dependency on the
toolchain. toolchain.
### `just uninstall` ## Running a single test
Removes `coverage.out`, the `gasm` binary, and `*.test` artefacts.
## Running Individual Tests
```sh ```sh
go test -run TestVexGroundTruth ./asm/ go test -run TestVexGroundTruth ./asm/
@@ -85,32 +86,53 @@ go test -run TestGOObjectLinkAndRun ./asm/
go test -run TestFuzzWideCopy ./verify/ go test -run TestFuzzWideCopy ./verify/
``` ```
## Debugger Note Add `-v` for the sub-test names, and `-race` when the change touches
concurrency. `-count=1` defeats the test cache when a result looks stale.
`gasm debug` spawns a child process from the binary on `$PATH`. It does ## Coverage
not work with `go run`; install first:
```sh ```sh
just install-bin just test
gasm debug --func decodeBlockAVX2 path/to/kernel_amd64.s go tool cover -func=coverage.out
``` ```
## Project Layout The `total:` line is the number that matters, and it stays at 80 percent or
more.
## Debugging the build
```sh
go build -gcflags='-m' ./... # inlining decisions
go build -gcflags='-S' ./... # what the compiler generated
go tool asm -S kernel_amd64.s # how the toolchain's assembler encodes a kernel
gasm dis kernel_amd64.s # what gasm makes of the same kernel
gasm tokens kernel_amd64.s # the token stream
gasm profile kernel_amd64.s # the basic blocks of each function
``` ```
cmd/gasm/ CLI entry point (subcommands)
token/ Lexical token kinds and positions `gasm verify --ground-truth` is the differential check that ties the two
lexer/ Hand-written scanner together: it compares gasm's bytes with `go tool asm`'s, with the relocation
ast/ Abstract syntax tree sites masked, so an encoding drift shows up as a byte difference rather than a
parser/ Line-oriented parser crash later.
arch/ Register and instruction tables (generated)
lint/ Static analysis rules ## Continuous integration
format/ Canonical formatter
lsp/ Language Server Protocol server Workflows live in `.gitea/workflows/` and run on the project's own runners:
asm/ Standalone assembler, encoder, object emitters Test on a push or pull request to `development`, race dispatched by hand, and
verify/ JIT execution, differential testing, ABI checks the release on a `v*` tag. They are written by hand rather than through
debug/ Interactive ptrace debugger (all four architectures) `just`, but they enforce the same set of gates minus the race detector, which
_gen/ Instruction table generator the shared runner cannot afford on a push; a green `just gates` locally is
testdata/ Test fixtures therefore the fastest way to a green pipeline.
docs/ Architecture, development, CLI reference
``` ## Releases
Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`,
which triggers the release workflow: it builds the portable Linux targets,
takes the notes from the matching `CHANGELOG.md` section and uploads the
assets.
The version is never injected. `gasm --version` prints what the
toolchain recorded in the build information: the tag on a tagged
checkout, a pseudo-version naming the commit below one, `+dirty` on a
dirty tree, and `(devel)` outside version control. There is no
`-ldflags "-X"` anywhere and no version constant in the source.
+62
View File
@@ -0,0 +1,62 @@
.TH GASM-ASM 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-asm \- assemble Plan 9 assembly without the Go toolchain
.SH SYNOPSIS
.B gasm asm [\-\-format raw|elf|goobj] [\-p pkg] [\-GOARCH arch] [\-o out] <file>
.SH DESCRIPTION
Assemble FILE without the Go toolchain: every TEXT function is encoded
to machine code and printed as a hex dump. Supported architectures:
amd64 (including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer,
FP, conditional select, CRC32 and MOV pseudo), riscv64 (RV64IMAFDC and
RVC) and loong64 (LoongArch base ISA).
.PP
With
.B \-o
the output is written to a file instead. The
.B \-\-format
flag selects what is written:
.B raw
(the default) concatenates the functions and the data section into one
self-consistent image;
.B elf
emits a relocatable object (.text/.data sections, a symbol table and
one PC32 relocation per static-symbol reference) that links with the
system toolchain;
.B goobj
emits the Go toolchain's own object format, which cmd/link consumes
directly (it requires
.BR \-p ,
the package path, and the installed Go toolchain).
.PP
Framed functions receive the stack-split guard and the trailing
morestack block, byte-identical to the toolchain's output, so split
functions link too.
.SH OPTIONS
.TP
.B \-\-format \fIraw|elf|goobj\fR
Output format; the default is raw.
.TP
.B \-p \fIpkg\fR
Package path for --format goobj, qualifying the exported symbols.
.TP
.B \-GOARCH \fIarch\fR
Target architecture: amd64, arm64, riscv64 or loong64; overrides the
file-name suffix, which is how the suffix-less majority of GOROOT's
files (cpu_x86.s, stub.s, ...) become assemblable.
.TP
.B \-o \fIfile\fR
Write the output to this file instead of a hex dump on stdout.
.SH EXIT STATUS
Exits 0 on success, 1 when parsing or assembly fails, and 2 on a usage
error.
.SH EXAMPLES
.nf
gasm asm \-o hello.bin hello_amd64.s raw image
gasm asm \-\-format elf \-o k.o k.s linkable ELF object
gasm asm \-\-format goobj \-p pkg/path \-o k.o k.s Go object for go build
gasm asm \-GOARCH amd64 cpu_x86.s arch override
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-dis (1),
.BR gasm\-verify (1)
+47
View File
@@ -0,0 +1,47 @@
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
.SH SYNOPSIS
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [amd64|arm64|riscv64|loong64]
.SH DESCRIPTION
Compare the gasm encoder for the given architecture (default amd64)
against
.B go tool asm
and print the diff: superset encodings (gasm-only, shippable via
.BR "gasm asm \-\-format goobj" ),
known-but-unencodable names (the backlog) and go-only names (feature
gaps). The Go side is probed black-box with a battery of bare
mnemonics, so the audit tracks whatever toolchain
.B go env GOROOT
provides.
.PP
With
.BR \-\-corpus ,
the audit changes shape: it assembles every
.I .s
file under the given directory (default GOROOT/src) with the gasm
encoder only, no toolchain probing. A file whose name carries a
recognisable _arch suffix is attempted for that architecture; a file
without one is attempted for all four, exactly as a GOARCH build would
compile it. The report gives the headline number (files that assemble
for every target architecture), the per-architecture pass rates and the
most common failure reasons, which drive the encodability backlog by
frequency rather than by table order. A run over GOROOT takes under a
second.
.SH OPTIONS
.TP
.B \-\-corpus [\fIdir\fR]
Assemble a corpus of .s files and report pass rates and failure
reasons.
.SH EXIT STATUS
The mnemonic-diff mode reports through its output and exits 0; a failed
probe or an unknown architecture exits non-zero.
.SH EXAMPLES
.nf
gasm audit\-instructions amd64
gasm audit\-instructions \-\-corpus
gasm audit\-instructions \-\-corpus "$(go env GOROOT)/src/crypto"
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-asm (1)
+113
View File
@@ -0,0 +1,113 @@
.TH GASM-DEBUG 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-debug \- interactive source-level debugger for JIT-assembled functions
.SH SYNOPSIS
.B gasm debug <file.s> \-\-func <name>
.SH DESCRIPTION
Interactive debugger for JIT-assembled functions. Launches the function
in a traced subprocess (ptrace), then provides a REPL for
single-stepping, breakpoints, register and memory inspection.
.PP
With
.B \-\-script
the REPL commands run from a file and the session ends: the headless
mode CI and scripts use.
.B \-\-cover
runs to completion with a breakpoint on every instruction and reports
which executed and how often, the label-level coverage view.
.SH REPL COMMANDS
.TP
.B break \fIlabel|addr\fR [\fBif \fIreg op val\fR]
Set a breakpoint, optionally conditional on a register comparison
(register against register or immediate).
.TP
.B delete \fIlabel|addr\fR
Remove a breakpoint.
.TP
.B info break
List all breakpoints.
.TP
.BR step " [" n ], " s
Single-step n instructions; the default is 1.
.TP
.BR next ", " n
Step over a CALL.
.TP
.BR finish ", " fin
Run until the function returns.
.TP
.BR continue ", " c
Run until a breakpoint, watchpoint or exit.
.TP
.BR disas " [" n ], " u
Disassemble n instructions at PC.
.TP
.B regs
Print general-purpose and vector registers.
.TP
.B where
Show the source line and nearest label at PC.
.TP
.B stack
Show the stack near RSP (return address and ABI0 args).
.TP
.BR bt ", " backtrace
Backtrace: current frame plus return address.
.TP
.B x [\fIaddr\fR] [\fIlen\fR]
Hex-dump memory; the defaults are the current PC and 64 bytes.
.TP
.B w \fIaddr val...\fR
Write bytes to memory.
.TP
.B set \fIreg value\fR
Set a register.
.TP
.B watch \fIaddr\fR [\fBr|w\fR] [\fIsize\fR]
Set a hardware watchpoint; writes are watched by default.
.TP
.B unwatch [\fIslot\fR]
Clear one watchpoint, or all without an argument.
.TP
.BR labels ", " l
List function labels and offsets.
.TP
.BR help ", " h ", " ?
Show command help.
.TP
.BR quit ", " q
Kill the debuggee and exit.
.SH OPTIONS
.TP
.B \-args \fIfile\fR
File containing the ABI0 argument block.
.TP
.B \-buf \fIspec\fR
Buffer specification: name:size:pattern[,name:size:pattern...] where
pattern is zero, ones, seq, or hex.
.TP
.B \-cover
Run to completion with a breakpoint on every instruction and report
which executed and how often.
.TP
.B \-func \fIname\fR
Function to debug.
.TP
.B \-script \fIfile\fR
Run REPL commands from a file (one per line) and exit; - reads stdin.
.TP
.B \-timeout \fIduration\fR
Kill the debuggee after this duration (e.g. 30s); for headless --script
runs.
.SH EXIT STATUS
Exits 0 when the scripted session completes and 1 when the debuggee
crashes or a check fails; the debugger is Linux-only.
.SH EXAMPLES
.nf
gasm debug \-\-func name k.s
gasm debug \-\-func name \-\-script cmds.txt \-\-timeout 30s k.s
gasm debug \-\-func name \-\-cover k.s
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-verify (1)
+35
View File
@@ -0,0 +1,35 @@
.TH GASM-DIFF 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-diff \- compare the machine code of two assembly files
.SH SYNOPSIS
.B gasm diff [\-GOARCH arch] <file1.s> <file2.s>
.SH DESCRIPTION
Compare the machine code produced by assembling two files. Shows which
functions differ and the byte-level differences. Useful for verifying
that two implementations produce identical code, or for tracking
encoding changes between Go assembler versions.
.PP
Functions are paired by exact name unless
.B \-\-map
says otherwise, so
.B \-\-map wideCopyAVX2=wideCopyAVX512
pairs two variants regardless of suffix.
.SH OPTIONS
.TP
.B \-GOARCH \fIarch\fR
Target architecture for both files: amd64, arm64, riscv64 or loong64;
overrides the file-name suffixes.
.TP
.B \-\-map \fIspec\fR
Comma-separated old=new pairs to match functions with different names.
.SH EXIT STATUS
Exits 0 when every paired function is identical and 1 when anything
differs; a usage error exits 2.
.SH EXAMPLES
.nf
gasm diff hello_amd64.s hello_amd64.s
gasm diff \-\-map wideCopyAVX2=wideCopyAVX512 avx2.s avx512.s
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-asm (1)
+36
View File
@@ -0,0 +1,36 @@
.TH GASM-DIS 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-dis \- disassemble machine code to instruction text
.SH SYNOPSIS
.B gasm dis [\-a arch] <file>
.SH DESCRIPTION
Disassemble machine code to instruction text, decoded through
golang.org/x/arch.
.PP
With a
.I .s
file, the file is assembled first and the listing follows the real
layout: one block per TEXT function, local labels printed at their
offsets. The architecture comes from the file-name suffix, or from
.BR \-a .
.PP
With any other file, or
.B \-
for standard input, the bytes are disassembled linearly and
.B \-a
selects the architecture (amd64, arm64, riscv64 or loong64).
.SH OPTIONS
.TP
.B \-a \fIarch\fR
Architecture for raw input: amd64, arm64, riscv64 or loong64.
.SH EXIT STATUS
Exits 0 on success, 1 when assembly or decoding fails, and 2 on a usage
error.
.SH EXAMPLES
.nf
gasm dis k.s assemble, then list each function
gasm dis \-a amd64 \- < dump.bin disassemble raw bytes from stdin
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-asm (1)
+52
View File
@@ -0,0 +1,52 @@
.TH GASM-FMT 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-fmt \- canonicalise the formatting of Plan 9 assembly sources
.SH SYNOPSIS
.B gasm fmt [\-w|\-l|\-d] [path...]
.SH DESCRIPTION
Canonicalise the formatting of Plan 9 assembly sources: indentation,
operand spacing, per-function mnemonic alignment and blank-line layout
(exactly one blank line before each label, TEXT and GLOBL block).
Formatting is idempotent and preserves every line, comments included.
.PP
With no paths, or a directory path, every
.I .s
file below it is reformatted in place and the changed files are listed,
the way
.B go fmt
does;
.B .
and
.B _
directories are skipped. Explicit file paths print to stdout unless
.B \-w
is given.
.PP
.B \-l
and
.B \-d
rewrite nothing:
.B \-l
prints the paths whose formatting differs from gasm's (empty output
means everything is formatted, which is what a CI check wants),
.B \-d
prints the diffs. They are mutually exclusive.
.SH OPTIONS
.TP
.B \-d
Print diffs instead of rewriting files.
.TP
.B \-l
List files whose formatting differs from gasm's.
.TP
.B \-w
Write the result to the source file.
.SH EXAMPLES
.nf
gasm fmt reformat every .s below here
gasm fmt \-w kernel_amd64.s canonicalise one file in place
gasm fmt \-l *.s list files whose formatting differs
gasm fmt \-d kernel_amd64.s print a unified diff instead
.fi
.SH SEE ALSO
.BR gasm (1)
+90
View File
@@ -0,0 +1,90 @@
.TH GASM-LINT 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-lint \- run the static checks over assembly files
.SH SYNOPSIS
.B gasm lint <file...>
.SH DESCRIPTION
Run the static checks over the given files and print diagnostics as
\fIfile:line:col: severity: message [code]\fR. The exit status is
non-zero when an error-severity diagnostic is found; warnings (e.g. the
register-clobber audit) do not affect it.
.PP
The checks are conservative: they report what can be proven wrong and
stay quiet otherwise, so a clean lint run is meaningful without
suppression lists.
.SH RULES
.TP
.B unknown-instruction
The mnemonic is not in the architecture's instruction table.
.TP
.B operand-count
The operand count disagrees with the instruction's declared arity.
.TP
.B undefined-label
A jump target names no label in the function.
.TP
.B duplicate-label
Two labels in one function share a name.
.TP
.B missing-ret
The function can fall off its end without a terminator.
.TP
.B missing-textflag-include
TEXT flags are used without including textflag.h.
.TP
.B abi-argsize
The declared frame or argument size disagrees with the
.B //\ function
signature.
.TP
.B unreachable-code
Code after RET and before the next label is dead; suppressed for
functions with PC-relative or register-indirect control flow.
.TP
.B register-clobber
A register the Go ABI fixes across calls is written without save and
restore, computed by liveness over the control-flow graph.
.TP
.B funcdata-pcdata
FUNCDATA and PCDATA indices are malformed.
.TP
.B unused-label
A label no jump reaches.
.TP
.B invalid-textflag
A TEXT flag combination the toolchain rejects.
.TP
.B stack-imbalance
The function does not restore the stack pointer on every path.
.TP
.B register-width-mismatch
An operand register has the wrong width for the instruction.
.TP
.B abi0-register-args
A call passes arguments in registers where ABI0 expects the stack
frame.
.TP
.B nonportable-register-name
A register spelling that does not exist on the target architecture.
.TP
.B unencodable-instruction
The mnemonic is known to the table but the encoder cannot assemble it
yet (amd64).
.TP
.B reserved-register-write
A write to the register the runtime reserves (arm64 R18).
.SH OPTIONS
.TP
.B \-disable \fIcodes\fR
Comma-separated rule codes to disable.
.SH EXIT STATUS
Exits 0 when no error-severity diagnostic is found, 1 otherwise, and 2
on a usage error.
.SH EXAMPLES
.nf
gasm lint kernel_amd64.s
gasm lint \-disable register-clobber,unused-label *.s
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-asm (1)
+24
View File
@@ -0,0 +1,24 @@
.TH GASM-LSP 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-lsp \- run the Plan 9 assembly language server
.SH SYNOPSIS
.B gasm lsp
.SH DESCRIPTION
Run the language server over standard input/output: JSON-RPC 2.0 with
Content-Length framing. Point an LSP-capable editor at the binary and
associate it with
.I .s
files; the target architecture is inferred from the file suffix
(_amd64.s, _arm64.s, _riscv64.s, _loong64.s).
.PP
Provides completion, hover, document symbols, push and pull
diagnostics, semantic-token highlighting, go-to-definition, find
references, rename, formatting, inlay hints, code actions, signature
help, document highlights, workspace symbol search, include document
links and folding ranges; definition, references and rename work across
every open document. Syntax highlighting is delivered as LSP semantic
tokens, so no editor-specific grammar is required.
.SH EXIT STATUS
Runs until the client closes the session; exits 0 on a clean shutdown.
.SH SEE ALSO
.BR gasm (1)
+20
View File
@@ -0,0 +1,20 @@
.TH GASM-PARSE 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-parse \- parse an assembly file and report syntax errors
.SH SYNOPSIS
.B gasm parse <file>
.SH DESCRIPTION
Parse FILE and report syntax errors on stderr. The parser is
error-tolerant and line-oriented: a malformed line becomes a diagnostic
and parsing continues, so one run reports every syntax error in the
file rather than the first.
.PP
On success, print how many declarations and TEXT functions the file
contains. FILE may be
.B \-
to read standard input.
.SH EXIT STATUS
Exits 0 when the file parses without errors and 1 otherwise.
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-tokens (1)
+16
View File
@@ -0,0 +1,16 @@
.TH GASM-PROFILE 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-profile \- show the basic-block structure of functions
.SH SYNOPSIS
.B gasm profile <file.s>
.SH DESCRIPTION
Show the basic-block structure of functions in an assembly file: each
function's labels, their offsets, and the block boundaries. This is
the static structure; for runtime execution counts, use
.BR "gasm verify \-fuzz" ,
which exercises the code paths.
.SH EXIT STATUS
Exits 0 on success and 1 when the file cannot be assembled.
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-verify (1)
+22
View File
@@ -0,0 +1,22 @@
.TH GASM-SCAFFOLD 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-scaffold \- generate a differential test skeleton for a kernel file
.SH SYNOPSIS
.B gasm scaffold differential <file.s>
.SH DESCRIPTION
Print a differential test skeleton for every
.B //\ func
signature in FILE. The test seeds random states, drives the kernel and
a portable reference (\fI<name>Portable\fR), and compares outputs
byte-for-byte. Write the reference bodies, place the file in the
kernel's package, and run it in CI.
.SH EXIT STATUS
Exits 0 when the skeleton is written to stdout and 1 when the file
cannot be parsed; a usage error exits 2.
.SH EXAMPLES
.nf
gasm scaffold differential kernel_amd64.s > kernel_differential_test.go
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-verify (1)
+19
View File
@@ -0,0 +1,19 @@
.TH GASM-TOKENS 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-tokens \- print the lexical token stream of an assembly file
.SH SYNOPSIS
.B gasm tokens <file>
.SH DESCRIPTION
Print the lexical token stream of FILE: position, token kind and text,
one token per line. FILE may be
.B \-
to read standard input.
.PP
This is the front end's raw view, for when the assembler's own
diagnostic is not enough: a mis-scanned operand or a swallowed comment
shows up here as the tokens the parser actually received.
.SH EXIT STATUS
Exits 0 on success and 1 when the file cannot be read.
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-parse (1)
+117
View File
@@ -0,0 +1,117 @@
.TH GASM-VERIFY 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm-verify \- JIT-assemble a file and run dynamic checks against it
.SH SYNOPSIS
.B gasm verify [\-smoke] [\-abi] [\-fuzz] [\-ground\-truth] [\-profile] [\-call] <file.s>
.SH DESCRIPTION
Assemble FILE, map it into executable memory and report the available
functions. This confirms the assembled image is self-consistent (no
unresolved external symbols) and executable, the prerequisite for
dynamic testing.
.PP
With
.BR \-smoke ,
each NOSPLIT function is called with a zeroed argument block to confirm
the JIT trampoline works end-to-end. This is safe only for functions
that tolerate nil pointers and zero lengths in their arguments.
.PP
With
.BR \-abi ,
each function is called with sentinel values in the registers the Go
ABI fixes across calls (the frame pointer and the goroutine pointer)
plus a canary below SP; violations are reported. JIT-based checks run
when the host matches the file's architecture (all but loong64, which
is ground-truth only for now).
.PP
With
.BR \-fuzz ,
each function with a
.B //\ func
signature is differentially fuzzed against the go-tool-asm version in a
subprocess (so a crash on a partial function is reported, not fatal).
.PP
With
.BR \-ground\-truth ,
the assembled machine code is compared byte-for-byte against
.B go tool asm
(relocation sites masked), reporting any encoding drift.
.PP
With
.BR \-profile ,
the static basic-block structure is listed for each function.
.PP
With
.BR \-call ,
a single function is invoked with user-supplied buffers
.RB ( \-buf )
instead of the smoke/abi/fuzz sweeps. Useful for partial functions
(e.g. decoders) that crash on random input but should succeed on valid
data.
.PP
With
.B \-save\-corpus
(and
.BR \-fuzz ),
every input that crashes or mismatches is written to the directory as
replayable JSON.
.B \-replay
re-runs saved entries against the kernel, one child process per entry,
so an input that crashed the original run crashes only the child: the
report says whether each entry reproduces.
.SH OPTIONS
.TP
.B \-abi
Run ABI-checking calls (sentinel registers and red zone).
.TP
.B \-abi\-n \fIn\fR
Number of ABI check iterations with varied inputs; the default is 100.
.TP
.B \-args \fIspec\fR
Scalar args for -call: name=value[,name=value] (decimal or 0x hex).
.TP
.B \-buf \fIspec\fR
Buffer spec for -call: name:size:pattern[,name:size:pattern] where
pattern is zero, ones, seq, or hex.
.TP
.B \-call \fIname\fR
Call a single function with -buf instead of the sweeps.
.TP
.B \-fuzz
Differential fuzz: JIT both the gasm and the go-tool-asm versions and
compare outputs.
.TP
.B \-ground\-truth
Compare machine code byte-for-byte against go tool asm.
.TP
.B \-n \fIn\fR
Number of fuzz iterations per function; the default is 1000.
.TP
.B \-profile
List basic-block structure per function.
.TP
.B \-repeat \fIn\fR
Number of times to repeat a -call invocation; the default is 1.
.TP
.B \-replay \fIdir\fR
Replay saved corpus entries (JSON files in this directory) against the
kernel.
.TP
.B \-save\-corpus \fIdir\fR
With -fuzz: write each failing input to this directory as replayable
JSON.
.TP
.B \-smoke
Call each NOSPLIT function with zeroed args.
.SH EXIT STATUS
Exits 0 when every requested check passes and 1 when any check fails;
a file that cannot be assembled exits 1 and a usage error exits 2.
.SH EXAMPLES
.nf
gasm verify \-\-call add \-\-args a=2,b=3 hello_amd64.s
gasm verify \-\-ground\-truth k.s
gasm verify \-\-fuzz \-n 500 k.s
.fi
.SH SEE ALSO
.BR gasm (1),
.BR gasm\-asm (1),
.BR gasm\-debug (1)
+96
View File
@@ -0,0 +1,96 @@
.TH GASM 1 "2026-09-19" "gasm 0.33.0" "User Commands"
.SH NAME
gasm \- developer tooling for Go's Plan 9 assembler
.SH SYNOPSIS
.B gasm
.I command
.RI [ arguments ]
.br
.B gasm
.BR \-h | \-\-help
.br
.B gasm
.BR \-V | \-\-version
.SH DESCRIPTION
.B gasm
bundles a lexer, parser, formatter, linter, standalone assembler and
language server for Plan 9 assembly into one self-contained binary. It
serves two purposes: it brings developer tooling to the
.I .s
files of Go programs, and it assembles Plan 9 assembly without the Go
toolchain at all, to raw images, linkable ELF objects with DWARF5 debug
sections, or the Go toolchain's own GOOBJ format, which
.B go build
consumes directly.
.PP
Four architectures are covered: amd64 (including VEX/AVX2 and
EVEX/AVX-512), arm64, riscv64 (RV64IMAFDC and RVC) and loong64. The
target architecture is inferred from the file-name suffix
(\fI_amd64.s\fR, \fI_arm64.s\fR, \fI_riscv64.s\fR, \fI_loong64.s\fR) or
named explicitly with \fB\-GOARCH\fR where the commands accept it.
.SH COMMANDS
.TP
.B gasm\-tokens(1)
Print the lexical token stream.
.TP
.B gasm\-parse(1)
Parse a file and report syntax errors.
.TP
.B gasm\-fmt(1)
Canonicalise formatting: gofmt for assembly.
.TP
.B gasm\-lint(1)
Run the static checks.
.TP
.B gasm\-asm(1)
Assemble \fI.s\fR files to machine code, raw images, ELF objects or GOOBJ.
.TP
.B gasm\-dis(1)
Disassemble machine code, raw bytes or an assembled \fI.s\fR file.
.TP
.B gasm\-verify(1)
JIT-assemble and run dynamic checks: smoke calls, ABI checks,
differential fuzzing, ground-truth comparison.
.TP
.B gasm\-debug(1)
Interactive source-level debugger.
.TP
.B gasm\-diff(1)
Compare the machine code of two files byte-for-byte.
.TP
.B gasm\-profile(1)
Show the basic-block structure of functions.
.TP
.B gasm\-audit\-instructions(1)
Diff the encoder against the Go toolchain's name table, or measure a
corpus of \fI.s\fR files.
.TP
.B gasm\-scaffold(1)
Generate a differential test skeleton for a kernel file.
.TP
.B gasm\-lsp(1)
Run the language server over standard input/output.
.TP
.B gasm version
Print the version, the same as \fB\-\-version\fR.
.SH GLOBAL FLAGS
.TP
.BR \-h ", " \-\-help
Show the command overview.
.TP
.BR \-V ", " \-\-version
Print the version the toolchain recorded at build time.
.SH EXIT STATUS
Exits 0 on success, 1 when a command fails, and 2 on a usage error. An
unknown command exits 2.
.SH SEE ALSO
.BR gasm\-asm (1),
.BR gasm\-fmt (1),
.BR gasm\-lint (1),
.BR gasm\-verify (1),
.BR gasm\-debug (1)
.PP
The full command reference, with worked examples and every flag, is in
docs/CLI.md of the repository
.UR https://sourcedock.dev/petrbalvin/gasm-devkit
.UE .
+96 -9
View File
@@ -50,12 +50,31 @@ func Source(src string) string {
} }
case len(line) >= 2 && line[1].Kind == token.Colon: case len(line) >= 2 && line[1].Kind == token.Colon:
inf.kind = kLabel inf.kind = kLabel
// Peel stacked labels exactly as the render pass does; the
// instruction after the last one is rendered at the
// function's alignment width, so its mnemonic counts here.
rest := line[2:]
for len(rest) >= 2 && rest[0].Kind == token.Ident && rest[1].Kind == token.Colon &&
!isDirective(rest[0].Text) {
rest = rest[2:]
}
if len(rest) > 0 && rest[0].Kind == token.Ident && !isDirective(rest[0].Text) {
inf.mnemLen = len(rest[0].Text)
if funcID >= 0 && inf.mnemLen > maxWidth[funcID] {
maxWidth[funcID] = inf.mnemLen
}
}
default: default:
inf.kind = kInstr inf.kind = kInstr
inf.funcID = funcID inf.funcID = funcID
inf.mnemLen = len(line[0].Text) // Only an identifier mnemonic takes the alignment width; a
if funcID >= 0 && inf.mnemLen > maxWidth[funcID] { // line starting with anything else renders unpadded, so its
maxWidth[funcID] = inf.mnemLen // length must not enter the width either.
if line[0].Kind == token.Ident {
inf.mnemLen = len(line[0].Text)
if funcID >= 0 && inf.mnemLen > maxWidth[funcID] {
maxWidth[funcID] = inf.mnemLen
}
} }
} }
} }
@@ -83,12 +102,35 @@ func Source(src string) string {
out = line[0].Text + " " + renderOps(line[1:]) out = line[0].Text + " " + renderOps(line[1:])
inBody = line[0].Text == "TEXT" inBody = line[0].Text == "TEXT"
case kLabel: case kLabel:
out = line[0].Text + ":" // Every label, and a trailing instruction, becomes its own
// A label may share its line with an instruction; emit the // output line: separate outLines keep the blank-line pass
// instruction on the following line. // honest about what it is looking at.
if rest := line[2:]; len(rest) > 0 { outs = append(outs, outLine{kind: kLabel, text: line[0].Text + ":"})
out += "\n" + renderInstr(rest, maxWidth[inf.funcID]) rest := line[2:]
for len(rest) >= 2 && rest[0].Kind == token.Ident && rest[1].Kind == token.Colon &&
!isDirective(rest[0].Text) {
outs = append(outs, outLine{kind: kLabel, text: rest[0].Text + ":"})
rest = rest[2:]
} }
// A label may share its line with an instruction; the canonical
// form puts the instruction on the following line. Trailing
// content that does not start an instruction (a stray operand
// token) stays on the label line: splitting it off would produce
// a line the parser rejects.
if len(rest) > 0 && rest[0].Kind == token.Ident && isDirective(rest[0].Text) {
// A bare directive cannot start a line of its own (the
// parser wants a symbol per line), so a directive sharing
// the label's line stays there.
outs[len(outs)-1].text += " " + strings.TrimRight(renderOps(rest), " \t")
} else if len(rest) > 0 && rest[0].Kind == token.Ident {
outs = append(outs, outLine{kind: kInstr, text: strings.TrimRight(renderInstr(rest, maxWidth[inf.funcID]), " \t")})
if strings.EqualFold(rest[0].Text, "RET") {
inBody = false
}
} else if len(rest) > 0 {
outs[len(outs)-1].text += " " + renderOps(rest)
}
continue
case kInstr: case kInstr:
out = renderInstr(line, maxWidth[inf.funcID]) out = renderInstr(line, maxWidth[inf.funcID])
// A RET ends the body for indentation purposes: comments that // A RET ends the body for indentation purposes: comments that
@@ -192,6 +234,13 @@ func renderInstr(line []token.Token, width int) string {
if ops == "" { if ops == "" {
return "\t" + mnem return "\t" + mnem
} }
// Alignment is a mnemonic convention: a line that does not start with
// an identifier (a stray operand token the parser tolerates) renders
// unpadded, so that no alignment width can depend on it and the output
// stays stable across passes.
if line[0].Kind != token.Ident {
return "\t" + mnem + " " + ops
}
if width < len(mnem) { if width < len(mnem) {
width = len(mnem) width = len(mnem)
} }
@@ -217,7 +266,19 @@ func renderPreproc(line []token.Token) string {
func renderOps(toks []token.Token) string { func renderOps(toks []token.Token) string {
var b strings.Builder var b strings.Builder
for i, t := range toks { for i, t := range toks {
if i > 0 && spaceBetween(toks[i-1], t) { sp := i > 0 && spaceBetween(toks[i-1], t)
// The accumulated text ending in '/' must never meet a '/' or '*':
// the pair would re-lex as a comment and the next pass would see a
// different line, whatever the token boundaries were.
if !sp && i > 0 && (t.Kind == token.Slash || t.Kind == token.Star) && strings.HasSuffix(b.String(), "/") {
sp = true
}
if sp {
b.WriteByte(' ')
} else if i > 0 && wouldMerge(toks[i-1], t) {
// The tight spelling would re-lex as something else ('/'
// before '*' opens a comment), which would make the next
// formatting pass see a different line.
b.WriteByte(' ') b.WriteByte(' ')
} }
b.WriteString(t.Text) b.WriteString(t.Text)
@@ -225,8 +286,27 @@ func renderOps(toks []token.Token) string {
return b.String() return b.String()
} }
// wouldMerge reports whether writing prev immediately before cur would
// re-lex as something other than those two tokens: a '/' before a '*' opens
// a comment, '>' before '>' shifts, and adjacent operators regroup.
func wouldMerge(prev, cur token.Token) bool {
var kinds []token.Kind
for _, t := range lexer.Tokenize(prev.Text + cur.Text) {
if t.Kind == token.EOF {
break
}
kinds = append(kinds, t.Kind)
}
return len(kinds) != 2 || kinds[0] != prev.Kind || kinds[1] != cur.Kind
}
// spaceBetween decides whether a single space separates prev and cur. // spaceBetween decides whether a single space separates prev and cur.
func spaceBetween(prev, cur token.Token) bool { func spaceBetween(prev, cur token.Token) bool {
// '/' beside '/' or '*' would form a comment opener in the output and
// make the next pass see a different line; keep them separated.
if prev.Kind == token.Slash && (cur.Kind == token.Slash || cur.Kind == token.Star) {
return true
}
switch cur.Kind { switch cur.Kind {
case token.RParen: case token.RParen:
return false return false
@@ -274,6 +354,13 @@ func splitLines(toks []token.Token) [][]token.Token {
if t.Kind == token.EOF { if t.Kind == token.EOF {
break break
} }
if t.Kind == token.Illegal {
// Illegal tokens carry no canonical spelling: the parser
// reports them as errors where they matter, and the formatter
// drops them so that a stray character cannot survive into the
// output and make the next pass render a different file.
continue
}
if t.Kind == token.Newline { if t.Kind == token.Newline {
lines = append(lines, cur) lines = append(lines, cur)
cur = nil cur = nil
+1 -1
View File
@@ -39,7 +39,7 @@ func TestGolden(t *testing.T) {
} }
// TestDocCommentIndent checks that a doc comment preceding a TEXT directive // TestDocCommentIndent checks that a doc comment preceding a TEXT directive
// sits at column 0 even when another function (ending in RET) precedes it — // sits at column 0 even when another function (ending in RET) precedes it;
// the RET must terminate the previous body for indentation purposes. // the RET must terminate the previous body for indentation purposes.
func TestDocCommentIndent(t *testing.T) { func TestDocCommentIndent(t *testing.T) {
in := "#include \"textflag.h\"\n" + in := "#include \"textflag.h\"\n" +
+47
View File
@@ -0,0 +1,47 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package format
import (
"os"
"path/filepath"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// FuzzFormatIdempotency hammers the formatter with arbitrary input. The
// contract: formatting twice equals formatting once, and input that parses
// cleanly still parses cleanly after formatting. The seed corpus carries the
// repository's kernels, so a plain `go test` run replays every seed as a
// regression case and CI exercises them without any fuzzing budget.
func FuzzFormatIdempotency(f *testing.F) {
for _, pattern := range []string{
"../testdata/*.s",
"../testdata/verify/*.s",
} {
files, _ := filepath.Glob(pattern)
for _, path := range files {
if b, err := os.ReadFile(path); err == nil {
f.Add(string(b))
}
}
}
f.Add("TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ AX, BX\n\tRET\n")
f.Add("TEXT ·f(SB),NOSPLIT,$0\n\tMOVQ AX,BX\n\n\n\tRET\n")
f.Add("garbage ### ???\n")
f.Fuzz(func(t *testing.T, src string) {
once := Source(src)
twice := Source(once)
if once != twice {
t.Fatalf("formatting is not idempotent:\nfirst: %q\nsecond: %q", once, twice)
}
if _, errs := parser.Parse("in.s", src); len(errs) == 0 {
if _, errs := parser.Parse("out.s", once); len(errs) > 0 {
t.Fatalf("formatted output of clean input does not parse: %v\n%s", errs[0], once)
}
}
})
}
@@ -0,0 +1,2 @@
go test fuzz v1
string("0:A:")
@@ -0,0 +1,2 @@
go test fuzz v1
string("$0/ *")
@@ -0,0 +1,2 @@
go test fuzz v1
string("A:TEXT")
@@ -0,0 +1,2 @@
go test fuzz v1
string("TEXT\n0:A:A0\nA 0")
@@ -0,0 +1,2 @@
go test fuzz v1
string("00/ /*")
@@ -0,0 +1,2 @@
go test fuzz v1
string("TEXT \n0:RET\n/*0")
@@ -0,0 +1,2 @@
go test fuzz v1
string("A:00")
@@ -0,0 +1,2 @@
go test fuzz v1
string("0:A\n0:")
@@ -0,0 +1,2 @@
go test fuzz v1
string("0\\")
@@ -0,0 +1,2 @@
go test fuzz v1
string("TEXT \n\" \nA\"")
@@ -0,0 +1,2 @@
go test fuzz v1
string("TEXT\nA:A0000\n0A")
@@ -0,0 +1,2 @@
go test fuzz v1
string("0> > >>")
@@ -0,0 +1,2 @@
go test fuzz v1
string("0:TEXT:")
@@ -0,0 +1,2 @@
go test fuzz v1
string("0:TEXT")
+1 -3
View File
@@ -1,7 +1,5 @@
module sourcedock.dev/petrbalvin/gasm-devkit module sourcedock.dev/petrbalvin/gasm-devkit
go 1.27 go 1.27.1
toolchain go1.27.0
require golang.org/x/arch v0.30.0 require golang.org/x/arch v0.30.0
+115 -40
View File
@@ -1,58 +1,133 @@
# Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org) # gasm-devkit.
# SPDX-License-Identifier: BSD-3-Clause #
# Replace gasm-devkit here and the values in the variable block. Everything below the block
# is the standard set and is identical in every repository; see the `justfile` skill.
binary := "gasm"
package := "./cmd/gasm"
# gasm-devkit — developer tooling for Go's Plan 9 assembler (GAsm). # What the test and bench recipes sweep. Scope this to the logic packages when a thin
# cmd/ drags the coverage floor down, for example "./internal/... ./pkg/...". Never name
# a directory the project does not have: a pattern that matches nothing is a setup
# failure, not an empty run.
packages := "./arch/... ./asm/... ./ast/... ./disasm/... ./format/... ./lexer/... ./lint/... ./lsp/... ./parser/... ./token/... ./verify/..."
version := "0.33.0" bindir := env_var_or_default("BINDIR", env_var("HOME") / ".local" / "bin")
# Where `install-man` puts the gzip-compressed pages (man1 below it). Exported because
# the Perl recipes read it from the environment.
export MANDIR := env_var_or_default("MANDIR", env_var("HOME") / ".local" / "share" / "man")
default: default:
@just --list @just --list
# Download module dependencies. # Compile. Zero errors, zero warnings.
install:
go mod download
# Vet + gofmt check — zero errors, zero warnings.
build: build:
go vet ./... CGO_ENABLED=0 go build -ldflags "-s -w" -o bin/{{binary}} {{package}}
@test -z "$(gofmt -l .)" || { echo "gofmt diff:"; gofmt -l .; exit 1; }
# Full test suite + race detector + 80 % coverage gate. # The test gate: the suite, no cache, the coverage floor.
# The coverage gate matches CI: it excludes packages that need hardware or
# are CLI glue (debug, cmd/gasm), so the number is identical locally and in CI.
test: test:
go test -race -count=1 ./... #!/usr/bin/env perl
go test -count=1 -coverprofile=coverage.out \ system(q{go}, q{test}, q{-count=1}, q{-timeout}, q{10m},
sourcedock.dev/petrbalvin/gasm-devkit/arch \ q{-coverprofile}, q{coverage.out}, qw({{packages}})) == 0
sourcedock.dev/petrbalvin/gasm-devkit/asm \ or die qq{the test suite failed\n};
sourcedock.dev/petrbalvin/gasm-devkit/ast \ open(my $c, q{-|}, q{go}, q{tool}, q{cover}, q{-func=coverage.out}) or die qq{cover: $!};
sourcedock.dev/petrbalvin/gasm-devkit/format \ my $total;
sourcedock.dev/petrbalvin/gasm-devkit/lexer \ while (my $l = <$c>) { $total = $1 if $l =~ m{^total:\s+\S+\s+([0-9.]+)%} }
sourcedock.dev/petrbalvin/gasm-devkit/lint \ close($c);
sourcedock.dev/petrbalvin/gasm-devkit/lsp \ die qq{no total line in coverage.out\n} unless defined $total;
sourcedock.dev/petrbalvin/gasm-devkit/parser \ printf qq{Total coverage: %s%%\n}, $total;
sourcedock.dev/petrbalvin/gasm-devkit/token \ exit($total < 80 ? 1 : 0);
sourcedock.dev/petrbalvin/gasm-devkit/verify
go tool cover -func=coverage.out | awk '/^total:/{gsub("%","",$3);if($3+0<80){print "coverage "$3"% < 80%";exit 1}print "coverage "$3"%"}'
# Format all Go sources. # The same suite under the race detector. The expensive one.
race:
go test -race -count=1 -timeout 10m {{packages}}
# Fast scoped run for iterating. This is the one that runs after every edit.
unit pkgs=packages run=".*":
go test {{pkgs}} -run '{{run}}'
# Time-boxed fuzz of one target in one package. The package is required; never a gate.
fuzz target pkg fuzztime="60s":
go test -run '^$' -fuzz '{{target}}' -fuzztime={{fuzztime}} {{pkg}}
# Benchmarks. On an idle machine only.
bench pkgs=packages:
go test -run '^$' -bench=. -benchmem -count=5 {{pkgs}}
# Format in place.
fmt: fmt:
gofmt -w . gofmt -w .
# Run the gasm CLI (pass args after --, e.g. `just run -- lint file.s`). # Zero diff. Prints nothing when everything is formatted.
run *ARGS: fmt-check:
go run -ldflags "-X main.version={{version}}" ./cmd/gasm {{ARGS}} #!/usr/bin/env perl
open(my $g, q{-|}, q{gofmt}, q{-l}, q{.}) or die qq{gofmt: $!};
my @bad = <$g>;
close($g);
print @bad;
exit(@bad ? 1 : 0);
# Install the gasm binary into $GOBIN (stamped with the release version). # Both static gates: go vet and go fix -diff.
install-bin: vet:
go install -ldflags "-X main.version={{version}}" ./cmd/gasm go vet ./...
go fix -diff ./...
# Regenerate the architecture instruction tables from the Go toolchain source. # The definition of done, in one command. Once per task, never per edit.
gates: build fmt-check vet test race
# Build artefacts only, not the installed binary.
clean:
rm -rf bin/ coverage.out
# Build, then copy the binary into bindir.
install: build
install -d "{{bindir}}"
install -m 755 bin/{{binary}} "{{bindir}}/{{binary}}"
# Remove the installed binary.
uninstall:
rm -f "{{bindir}}/{{binary}}"
# Install the man pages under docs/man into mandir/man1, gzip-compressed. Not a gate: a
# convenience for the person at the keyboard; man finds them through ~/.local/share/man.
install-man:
#!/usr/bin/env perl
my $out = $ENV{MANDIR} . q{/man1};
system(q{mkdir}, q{-p}, $out) == 0 or die qq{mkdir $out: $!\n};
for my $p (glob q{docs/man/*.1}) {
open(my $g, q{-|}, q{gzip}, q{-c}, $p) or die qq{gzip $p: $!\n};
my $content = do { local $/; <$g> };
close($g);
my $base = $p;
$base =~ s{docs/man/}{};
open(my $o, q{>}, qq{$out/$base.gz}) or die qq{write $out/$base.gz: $!\n};
print {$o} $content;
close($o);
print qq{$out/$base.gz\n};
}
# Remove the installed man pages.
uninstall-man:
#!/usr/bin/env perl
for my $p (glob q{docs/man/*.1}) {
my $base = $p;
$base =~ s{docs/man/}{};
my $f = $ENV{MANDIR} . q{/man1/} . $base . q{.gz};
if (-f $f) {
unlink($f) or die qq{unlink $f: $!\n};
print qq{removed $f\n};
}
}
# Run the program. The flag is there because `go run` does not stamp the build otherwise.
run:
go run -buildvcs=true {{package}}
# Run with watch or hot reload, where the project has one.
dev:
go run -buildvcs=true {{package}}
# Regenerate the architecture instruction tables from the Go toolchain source. Not a gate.
gen: gen:
go run _gen/gen.go go run _gen/gen.go
gofmt -w arch/ gofmt -w arch/
# Remove build artefacts.
uninstall:
rm -f coverage.out gasm
find . -name '*.test' -delete
+1 -1
View File
@@ -20,7 +20,7 @@ import (
// architecture's syntax (macros, addressing modes, branch aliases) against // architecture's syntax (macros, addressing modes, branch aliases) against
// production assembly. It is skipped when the toolchain source is absent. // production assembly. It is skipped when the toolchain source is absent.
// //
// The bar is zero parse errors and zero error-severity diagnostics — i.e. no // The bar is zero parse errors and zero error-severity diagnostics; i.e. no
// false "unknown instruction" / "undefined label" findings on code the real // false "unknown instruction" / "undefined label" findings on code the real
// assembler accepts. Advisory warnings are reported but not fatal, since they // assembler accepts. Advisory warnings are reported but not fatal, since they
// are heuristics that may legitimately differ across Go versions. // are heuristics that may legitimately differ across Go versions.
+29 -8
View File
@@ -209,10 +209,10 @@ func lintText(t *ast.Text, tab *arch.Table, archKnown bool, cfg Config, macros m
lastTerminal := false lastTerminal := false
hasMacro := false hasMacro := false
instrCount := 0 instrCount := 0
dead := false // inside a region unreachable from above dead := false // inside a region unreachable from above
reportedDead := false // the current dead region has already been reported reportedDead := false // the current dead region has already been reported
hasPCRel := referencesPC(t) // PC-relative jumps defeat reachability analysis hasPCRel := referencesPC(t) // PC-relative jumps defeat reachability analysis
hasIndirect := hasIndirectBranch(t) // register-indirect branches do too hasIndirect := hasIndirectBranch(t, tab) // register-indirect branches do too
// Unreachable-code analysis is only sound in functions whose control flow is // Unreachable-code analysis is only sound in functions whose control flow is
// fully label-resolvable: no PC-relative jumps, no register-indirect // fully label-resolvable: no PC-relative jumps, no register-indirect
// branches, and (file-level) no preprocessor conditionals. // branches, and (file-level) no preprocessor conditionals.
@@ -549,10 +549,12 @@ func referencesPC(t *ast.Text) bool {
} }
// hasIndirectBranch reports whether a function transfers control through a // hasIndirectBranch reports whether a function transfers control through a
// register (JALR/JR/JIRL/BR/BLR). Such targets are computed at runtime, so // register or a computed memory address: the RISC branch-register mnemonics
// reachability cannot be determined statically and the unreachable-code check is // (JALR/JR/JIRL/BR/BLR), or a JMP/CALL whose target is a register or memory
// suppressed for the whole function. // operand rather than a label or symbol. Such targets are computed at
func hasIndirectBranch(t *ast.Text) bool { // runtime, so reachability cannot be determined statically and the
// unreachable-code check is suppressed for the whole function.
func hasIndirectBranch(t *ast.Text, tab *arch.Table) bool {
for _, s := range t.Body { for _, s := range t.Body {
in, ok := s.(*ast.Instr) in, ok := s.(*ast.Instr)
if !ok { if !ok {
@@ -561,11 +563,30 @@ func hasIndirectBranch(t *ast.Text) bool {
switch strings.ToUpper(in.Mnemonic.Text) { switch strings.ToUpper(in.Mnemonic.Text) {
case "JALR", "JR", "JIRL", "BR", "BLR": case "JALR", "JR", "JIRL", "BR", "BLR":
return true return true
case "JMP", "CALL":
if indirectJumpTarget(in, tab) {
return true
}
} }
} }
return false return false
} }
// indirectJumpTarget reports whether the JMP/CALL operand addresses a
// register or a memory location rather than a label or a static symbol. The
// parser delivers a bare register and a bare label in the same shape, so
// register membership decides.
func indirectJumpTarget(in *ast.Instr, tab *arch.Table) bool {
if len(in.Operands) != 1 || in.Operands[0].Kind != ast.OpAddr {
return false
}
a := in.Operands[0].Addr
if a.Base != "" || a.Index != "" {
return true
}
return a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" && tab.IsRegister(a.Sym.Name)
}
// isMacroInvocation reports whether a mnemonic is a macro invocation rather // isMacroInvocation reports whether a mnemonic is a macro invocation rather
// than a machine instruction. No Plan 9 mnemonic contains an underscore, so an // than a machine instruction. No Plan 9 mnemonic contains an underscore, so an
// underscore is a reliable macro marker (the runtime headers define macros such // underscore is a reliable macro marker (the runtime headers define macros such
+38 -7
View File
@@ -8,6 +8,7 @@ import (
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser" "sourcedock.dev/petrbalvin/gasm-devkit/parser"
) )
@@ -48,7 +49,7 @@ func TestFixtureIsClean(t *testing.T) {
t.Fatalf("parse: %v", errs) t.Fatalf("parse: %v", errs)
} }
// The fixture mirrors the go-flac kernels, which write the Go ABI0 // The fixture mirrors the go-flac kernels, which write the Go ABI0
// scratch registers (BX, R13) without saving them — legal under Go's // scratch registers (BX, R13) without saving them; legal under Go's
// stack-based ABI, so the register-clobber audit stays silent and the // stack-based ABI, so the register-clobber audit stays silent and the
// fixture must lint entirely clean. // fixture must lint entirely clean.
diags := File(f, Config{Arch: arch.AMD64}) diags := File(f, Config{Arch: arch.AMD64})
@@ -186,8 +187,8 @@ done:
} }
} }
// TestEvexMaskingRecognised checks that masked EVEX forms — the .Z suffix and // TestEvexMaskingRecognised checks that masked EVEX forms; the .Z suffix and
// an explicit K operand — are recognised and exempt from operand-count // an explicit K operand; are recognised and exempt from operand-count
// checks. // checks.
func TestEvexMaskingRecognised(t *testing.T) { func TestEvexMaskingRecognised(t *testing.T) {
diags := lintSrc(t, ` diags := lintSrc(t, `
@@ -284,7 +285,7 @@ TEXT ·f(SB), NOSPLIT|NOFRAME|DUPOK, $0
} }
func TestStackImbalance(t *testing.T) { func TestStackImbalance(t *testing.T) {
// Function with frame size 16 but only SUB 8, SP — imbalance. // Function with frame size 16 but only SUB 8, SP; imbalance.
diags := lintSrc(t, ` diags := lintSrc(t, `
#include "textflag.h" #include "textflag.h"
TEXT ·f(SB), NOSPLIT, $16-0 TEXT ·f(SB), NOSPLIT, $16-0
@@ -297,7 +298,7 @@ TEXT ·f(SB), NOSPLIT, $16-0
} }
func TestStackBalanced(t *testing.T) { func TestStackBalanced(t *testing.T) {
// Function with frame size 16 and matching SUB/ADD — balanced. // Function with frame size 16 and matching SUB/ADD; balanced.
diags := lintSrc(t, ` diags := lintSrc(t, `
#include "textflag.h" #include "textflag.h"
TEXT ·f(SB), NOSPLIT, $16-0 TEXT ·f(SB), NOSPLIT, $16-0
@@ -311,7 +312,7 @@ TEXT ·f(SB), NOSPLIT, $16-0
} }
func TestRegisterWidthMismatch(t *testing.T) { func TestRegisterWidthMismatch(t *testing.T) {
// MOVQ with 32-bit register — mismatch. // MOVQ with 32-bit register; mismatch.
diags := lintSrc(t, ` diags := lintSrc(t, `
#include "textflag.h" #include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0 TEXT ·f(SB), NOSPLIT, $0
@@ -324,7 +325,7 @@ TEXT ·f(SB), NOSPLIT, $0
} }
func TestRegisterWidthCorrect(t *testing.T) { func TestRegisterWidthCorrect(t *testing.T) {
// MOVQ with 64-bit registers — correct. // MOVQ with 64-bit registers; correct.
diags := lintSrc(t, ` diags := lintSrc(t, `
#include "textflag.h" #include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0 TEXT ·f(SB), NOSPLIT, $0
@@ -495,3 +496,33 @@ TEXT ·f(SB), NOSPLIT, $0
t.Fatalf("amd64 must not be flagged: %+v", diags) t.Fatalf("amd64 must not be flagged: %+v", diags)
} }
} }
// TestHasIndirectBranchShape checks that a JMP/CALL through a register or
// memory suppresses reachability analysis, while a same-named label does not.
func TestHasIndirectBranchShape(t *testing.T) {
tab := arch.ForArch(arch.AMD64)
indirect := `TEXT ·f(SB), NOSPLIT, $0
JMP AX
RET
`
f, errs := parser.Parse("t_amd64.s", indirect)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if !hasIndirectBranch(f.Decls[0].(*ast.Text), tab) {
t.Error("JMP AX: indirect branch not detected")
}
label := `TEXT ·f(SB), NOSPLIT, $0
loop:
JMP loop
RET
`
f, errs = parser.Parse("t_amd64.s", label)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if hasIndirectBranch(f.Decls[0].(*ast.Text), tab) {
t.Error("JMP loop: label treated as an indirect branch")
}
}
+5 -5
View File
@@ -7,13 +7,13 @@ import "testing"
// TestRegisterClobber checks the register-clobber audit is calibrated to the // TestRegisterClobber checks the register-clobber audit is calibrated to the
// Go ABI (cmd/compile/abi-internal.md), not the platform ABI: Go's // Go ABI (cmd/compile/abi-internal.md), not the platform ABI: Go's
// stack-based ABI0 — which hand-written assembly uses — has no System V // stack-based ABI0, which hand-written assembly uses, has no System V
// style callee-saved registers, so argument and scratch registers may be // style callee-saved registers, so argument and scratch registers may be
// clobbered freely. Only the registers the ABI fixes across calls (the // clobbered freely. Only the registers the ABI fixes across calls (the
// frame pointer, the goroutine pointer, OS-reserved registers) are audited. // frame pointer, the goroutine pointer, OS-reserved registers) are audited.
func TestRegisterClobber(t *testing.T) { func TestRegisterClobber(t *testing.T) {
// amd64: BX, R12, R13 and R15 are argument/permanent-scratch registers in // amd64: BX, R12, R13 and R15 are argument/permanent-scratch registers in
// Go ABI0 — writing them unsaved is legal (a System V calibration would // Go ABI0; writing them unsaved is legal (a System V calibration would
// report all of these). // report all of these).
scratch := lintSrc(t, "#include \"textflag.h\"\n"+ scratch := lintSrc(t, "#include \"textflag.h\"\n"+
"TEXT ·f(SB), NOSPLIT, $0\n"+ "TEXT ·f(SB), NOSPLIT, $0\n"+
@@ -27,7 +27,7 @@ func TestRegisterClobber(t *testing.T) {
} }
// amd64: R14 (the goroutine pointer) in a NOSPLIT function without calls // amd64: R14 (the goroutine pointer) in a NOSPLIT function without calls
// is the runtime's own pattern — the ABI0 transition restores it — so it // is the runtime's own pattern (the ABI0 transition restores it), so it
// is not flagged. // is not flagged.
leaf := lintSrc(t, "#include \"textflag.h\"\n"+ leaf := lintSrc(t, "#include \"textflag.h\"\n"+
"TEXT ·f(SB), NOSPLIT, $0\n"+ "TEXT ·f(SB), NOSPLIT, $0\n"+
@@ -101,7 +101,7 @@ func TestRegisterClobber(t *testing.T) {
t.Fatalf("arm64 R18 write should be flagged: %+v", armReserved) t.Fatalf("arm64 R18 write should be flagged: %+v", armReserved)
} }
// riscv64: X27 holds the goroutine; X5–X7 are scratch. // riscv64: X27 holds the goroutine; X5-X7 are scratch.
riscScratch := lintSrcArch(t, "t_riscv64.s", "#include \"textflag.h\"\n"+ riscScratch := lintSrcArch(t, "t_riscv64.s", "#include \"textflag.h\"\n"+
"TEXT ·f(SB), NOSPLIT, $0\n"+ "TEXT ·f(SB), NOSPLIT, $0\n"+
"\tMOV X5, X6\n"+ "\tMOV X5, X6\n"+
@@ -117,7 +117,7 @@ func TestRegisterClobber(t *testing.T) {
t.Fatalf("unsaved riscv64 X27 write should be flagged: %+v", riscG) t.Fatalf("unsaved riscv64 X27 write should be flagged: %+v", riscG)
} }
// loong64: R22 holds the goroutine; R5–R19 are argument/scratch. // loong64: R22 holds the goroutine; R5-R19 are argument/scratch.
loongScratch := lintSrcArch(t, "t_loong64.s", "#include \"textflag.h\"\n"+ loongScratch := lintSrcArch(t, "t_loong64.s", "#include \"textflag.h\"\n"+
"TEXT ·f(SB), NOSPLIT, $0\n"+ "TEXT ·f(SB), NOSPLIT, $0\n"+
"\tMOVV R5, R6\n"+ "\tMOVV R5, R6\n"+
+41
View File
@@ -0,0 +1,41 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package parser
import (
"os"
"path/filepath"
"testing"
)
// FuzzParse hammers the parser with arbitrary input. The contract: no panic,
// and always a usable file, whether or not diagnostics were reported. The
// seed corpus carries the repository's kernels, so a plain `go test` run
// replays every seed as a regression case and CI exercises them without any
// fuzzing budget.
func FuzzParse(f *testing.F) {
for _, pattern := range []string{
"../testdata/*.s",
"../testdata/verify/*.s",
} {
files, _ := filepath.Glob(pattern)
for _, path := range files {
if b, err := os.ReadFile(path); err == nil {
f.Add(string(b))
}
}
}
f.Add("TEXT ·f(SB), NOSPLIT, $0\n\tRET\n")
f.Add("garbage ### ??? ::: \xff\xfe\n")
f.Add("#define A(x) x+1\nTEXT ·f(SB), $0\n\tA(2)\n\tRET\n")
f.Add("DATA t<>+0(SB)/8, $1\nGLOBL t<>(SB), RODATA, $8\n")
f.Add("TEXT ·f(SB), $0\n\tJMP (AX)\n\tCALL (BX)\n\tRET\n")
f.Fuzz(func(t *testing.T, src string) {
file, _ := Parse("fuzz.s", src)
if file == nil {
t.Fatal("Parse returned a nil file")
}
})
}
+47 -16
View File
@@ -438,25 +438,56 @@ func parseAddress(g []token.Token) ast.Address {
// Optional leading displacement before a '(' base group. A sign pushes // Optional leading displacement before a '(' base group. A sign pushes
// the parenthesis one token further out: -4(DX) has it at i+2. // the parenthesis one token further out: -4(DX) has it at i+2.
if isSignedNumber(g, i) { if isSignedNumber(g, i) {
paren := i + 1 j := i
if g[i].Kind == token.Minus || g[i].Kind == token.Plus { neg := false
paren = i + 2 if g[j].Kind == token.Minus {
neg = true
j++
} else if g[j].Kind == token.Plus {
j++
} }
if paren < len(g) && g[paren].Kind == token.LParen { if j < len(g) && g[j].Kind == token.Number {
neg := false v := parseInt(g[j].Text)
if g[i].Kind == token.Minus { j++
neg = true // A term may carry a *number factor: 0*8(base).
i++ for j+1 < len(g) && g[j].Kind == token.Star && g[j+1].Kind == token.Number {
} else if g[i].Kind == token.Plus { v *= parseInt(g[j+1].Text)
i++ j += 2
} }
if i < len(g) && g[i].Kind == token.Number { if neg {
addr.Offset = parseInt(g[i].Text) v = -v
addr.HasOff = true }
if neg { // Further +/- terms, each with its optional factor:
addr.Offset = -addr.Offset // 3*8+8(base), 8-4*2(base).
for {
termNeg := false
if j < len(g) && g[j].Kind == token.Minus {
termNeg = true
} else if j < len(g) && g[j].Kind == token.Plus {
} else {
break
} }
i++ if j+1 < len(g) && g[j+1].Kind == token.Number {
tv := parseInt(g[j+1].Text)
j += 2
for j+1 < len(g) && g[j].Kind == token.Star && g[j+1].Kind == token.Number {
tv *= parseInt(g[j+1].Text)
j += 2
}
if termNeg {
tv = -tv
}
v += tv
continue
}
break
}
// Commit only when the expression is followed by the base
// group; a bare number stays untouched for the caller.
if j < len(g) && g[j].Kind == token.LParen {
addr.Offset = v
addr.HasOff = true
i = j
} }
} }
} }
+3 -3
View File
@@ -149,7 +149,7 @@ func TestOperandStructure(t *testing.T) {
} }
} }
// MOVQ swin_base+0(FP), SI — the first MOVQ in the body. // MOVQ swin_base+0(FP), SI; the first MOVQ in the body.
var mov *ast.Instr var mov *ast.Instr
for _, s := range fn.Body { for _, s := range fn.Body {
if in, ok := s.(*ast.Instr); ok && in.Mnemonic.Text == "MOVQ" { if in, ok := s.(*ast.Instr); ok && in.Mnemonic.Text == "MOVQ" {
@@ -202,7 +202,7 @@ func TestAVX512Operands(t *testing.T) {
} }
} }
// VALIGND $15, Z9, Z0, Z1 — four operands. // VALIGND $15, Z9, Z0, Z1; four operands.
val := byMnem["VALIGND"] val := byMnem["VALIGND"]
if val == nil { if val == nil {
t.Fatal("VALIGND not found") t.Fatal("VALIGND not found")
@@ -224,7 +224,7 @@ func TestAVX512Operands(t *testing.T) {
t.Errorf("VMOVDQU32 dst = %+v, want 4(SI)(AX*1)", dst) t.Errorf("VMOVDQU32 dst = %+v, want 4(SI)(AX*1)", dst)
} }
// KTESTW K1, K1 — mask registers parse as bare names. // KTESTW K1, K1; mask registers parse as bare names.
kt := byMnem["KTESTW"] kt := byMnem["KTESTW"]
if kt == nil || len(kt.Operands) != 2 { if kt == nil || len(kt.Operands) != 2 {
t.Fatalf("KTESTW = %+v, want two operands", kt) t.Fatalf("KTESTW = %+v, want two operands", kt)
+5 -5
View File
@@ -5,8 +5,8 @@
// func add(a, b int64) int64 // func add(a, b int64) int64
TEXT ·add(SB), NOSPLIT, $0-24 TEXT ·add(SB), NOSPLIT, $0-24
MOV a+0(FP), X10 MOV a+0(FP), X10
MOV b+8(FP), X11 MOV b+8(FP), X11
ADD X11, X10, X10 ADD X11, X10, X10
MOV X10, ret+16(FP) MOV X10, ret+16(FP)
RET RET
+10 -10
View File
@@ -5,16 +5,16 @@
// func atomicAdd(ptr *int64, val int64) int64 // func atomicAdd(ptr *int64, val int64) int64
TEXT ·atomicAdd(SB), NOSPLIT, $0-24 TEXT ·atomicAdd(SB), NOSPLIT, $0-24
MOV a+0(FP), X10 MOV a+0(FP), X10
MOV b+8(FP), X11 MOV b+8(FP), X11
AMOADDD X11, (X10), X12 AMOADDD X11, (X10), X12
MOV X12, ret+16(FP) MOV X12, ret+16(FP)
RET RET
// func fpAdd(a, b float64) float64 // func fpAdd(a, b float64) float64
TEXT ·fpAdd(SB), NOSPLIT, $0-24 TEXT ·fpAdd(SB), NOSPLIT, $0-24
FLD a+0(FP), F10 FLD a+0(FP), F10
FLD b+8(FP), F11 FLD b+8(FP), F11
FADDD F10, F11, F12 FADDD F10, F11, F12
FSD F12, ret+16(FP) FSD F12, ret+16(FP)
RET RET
+13 -13
View File
@@ -5,22 +5,22 @@
// func readCSR(csr int64) int64 // func readCSR(csr int64) int64
TEXT ·readCSR(SB), NOSPLIT, $0-16 TEXT ·readCSR(SB), NOSPLIT, $0-16
MOV a+0(FP), X10 MOV a+0(FP), X10
CSRRS $0x300, X0, X11 CSRRS $0x300, X0, X11
MOV X11, ret+8(FP) MOV X11, ret+8(FP)
RET RET
// func setCSRBit(csr, bit int64) int64 // func setCSRBit(csr, bit int64) int64
TEXT ·setCSRBit(SB), NOSPLIT, $0-24 TEXT ·setCSRBit(SB), NOSPLIT, $0-24
MOV a+0(FP), X10 MOV a+0(FP), X10
MOV b+8(FP), X11 MOV b+8(FP), X11
CSRRS $0x304, X11, X12 CSRRS $0x304, X11, X12
MOV X12, ret+16(FP) MOV X12, ret+16(FP)
RET RET
// func writeCSR(val int64) int64 // func writeCSR(val int64) int64
TEXT ·writeCSR(SB), NOSPLIT, $0-16 TEXT ·writeCSR(SB), NOSPLIT, $0-16
MOV a+0(FP), X10 MOV a+0(FP), X10
CSRRW $0x305, X10, X11 CSRRW $0x305, X10, X11
MOV X11, ret+8(FP) MOV X11, ret+8(FP)
RET RET
+12 -12
View File
@@ -5,18 +5,18 @@
// func fma(a, b, c float64) float64 // func fma(a, b, c float64) float64
TEXT ·fma(SB), NOSPLIT, $0-32 TEXT ·fma(SB), NOSPLIT, $0-32
FLD a+0(FP), F10 FLD a+0(FP), F10
FLD b+8(FP), F11 FLD b+8(FP), F11
FLD c+16(FP), F12 FLD c+16(FP), F12
FMADDD F10, F11, F12, F13 FMADDD F10, F11, F12, F13
FSD F13, ret+24(FP) FSD F13, ret+24(FP)
RET RET
// func fms(a, b, c float64) float64 // func fms(a, b, c float64) float64
TEXT ·fms(SB), NOSPLIT, $0-32 TEXT ·fms(SB), NOSPLIT, $0-32
FLD a+0(FP), F10 FLD a+0(FP), F10
FLD b+8(FP), F11 FLD b+8(FP), F11
FLD c+16(FP), F12 FLD c+16(FP), F12
FMSUBD F10, F11, F12, F13 FMSUBD F10, F11, F12, F13
FSD F13, ret+24(FP) FSD F13, ret+24(FP)
RET RET
+22 -21
View File
@@ -6,31 +6,32 @@
// func casLoop(ptr *int64, old, new int64) bool // func casLoop(ptr *int64, old, new int64) bool
TEXT ·casLoop(SB), NOSPLIT, $0-32 TEXT ·casLoop(SB), NOSPLIT, $0-32
cas_retry: cas_retry:
MOV a+0(FP), X10 MOV a+0(FP), X10
LRD (X10), X11 LRD (X10), X11
MOV b+8(FP), X12 MOV b+8(FP), X12
BNE X11, X12, cas_fail BNE X11, X12, cas_fail
MOV c+16(FP), X13 MOV c+16(FP), X13
SCD X13, (X10), X14 SCD X13, (X10), X14
BNE X14, X0, cas_retry BNE X14, X0, cas_retry
ADDI X0, $1, X15 ADDI X0, $1, X15
MOV X15, ret+24(FP) MOV X15, ret+24(FP)
RET RET
cas_fail: cas_fail:
MOV X0, ret+24(FP) MOV X0, ret+24(FP)
RET RET
// func intToFloat(x int64) float64 // func intToFloat(x int64) float64
TEXT ·intToFloat(SB), NOSPLIT, $0-16 TEXT ·intToFloat(SB), NOSPLIT, $0-16
MOV a+0(FP), X10 MOV a+0(FP), X10
FCVTDL X10, F10 FCVTDL X10, F10
FSD F10, ret+8(FP) FSD F10, ret+8(FP)
RET RET
// func compare(a, b float64) bool // func compare(a, b float64) bool
TEXT ·compare(SB), NOSPLIT, $0-24 TEXT ·compare(SB), NOSPLIT, $0-24
FLD a+0(FP), F10 FLD a+0(FP), F10
FLD b+8(FP), F11 FLD b+8(FP), F11
FLTD F10, F11, X10 FLTD F10, F11, X10
MOV X10, ret+16(FP) MOV X10, ret+16(FP)
RET RET
+28 -27
View File
@@ -15,43 +15,44 @@ DATA mask24<>+4(SB)/4, $0x80050403
// func analyzeO1RangeAVX2(swin []int32, dstP []uint32, hist *[32]uint16) (partSum uint64, overflow bool) // func analyzeO1RangeAVX2(swin []int32, dstP []uint32, hist *[32]uint16) (partSum uint64, overflow bool)
TEXT ·analyzeO1RangeAVX2(SB), NOSPLIT, $0-65 TEXT ·analyzeO1RangeAVX2(SB), NOSPLIT, $0-65
MOVQ swin_base+0(FP), SI MOVQ swin_base+0(FP), SI
MOVQ dstP_base+24(FP), DI MOVQ dstP_base+24(FP), DI
MOVQ dstP_len+32(FP), BX MOVQ dstP_len+32(FP), BX
MOVQ hist+48(FP), R13 MOVQ hist+48(FP), R13
VPCMPEQD Y0, Y0, Y0 VPCMPEQD Y0, Y0, Y0
VPSLLD $31, Y0, Y0 VPSLLD $31, Y0, Y0
LEAQ (SI)(BX*4), R9 LEAQ (SI)(BX*4), R9
MOVQ BX, R10 MOVQ BX, R10
ANDQ $-8, R10 ANDQ $-8, R10
vec1: vec1:
CMPQ SI, R10 CMPQ SI, R10
JGE vec1done JGE vec1done
VMOVDQU (SI), Y1 VMOVDQU (SI), Y1
VMOVDQU 4(SI), Y2 VMOVDQU 4(SI), Y2
VPSUBD Y1, Y2, Y3 VPSUBD Y1, Y2, Y3
ADDQ $32, SI ADDQ $32, SI
JMP vec1 JMP vec1
vec1done: vec1done:
MOVQ AX, partSum+56(FP) MOVQ AX, partSum+56(FP)
MOVB AL, overflow+64(FP) MOVB AL, overflow+64(FP)
VZEROUPPER VZEROUPPER
RET RET
// func decodeFixedO1AVX512(samples []int32, residual []int32) // func decodeFixedO1AVX512(samples []int32, residual []int32)
TEXT ·decodeFixedO1AVX512(SB), NOSPLIT, $0-48 TEXT ·decodeFixedO1AVX512(SB), NOSPLIT, $0-48
MOVQ samples_base+0(FP), SI MOVQ samples_base+0(FP), SI
MOVQ residual_base+16(FP), DI MOVQ residual_base+16(FP), DI
VPBROADCASTD AX, Z15 VPBROADCASTD AX, Z15
VMOVDQU32 (DI)(AX*1), Z0 VMOVDQU32 (DI)(AX*1), Z0
VALIGND $15, Z9, Z0, Z1 VALIGND $15, Z9, Z0, Z1
VFMADD231PD Z14, Z12, Z10 VFMADD231PD Z14, Z12, Z10
VPCMPEQD Z0, Z3, K1 VPCMPEQD Z0, Z3, K1
KTESTW K1, K1 KTESTW K1, K1
VPSRAQ X31, Z8, Z8 VPSRAQ X31, Z8, Z8
VMOVDQU32 Z0, 4(SI)(AX*1) VMOVDQU32 Z0, 4(SI)(AX*1)
RET RET
+1 -1
View File
@@ -20,7 +20,7 @@ TEXT ·dirtyBP(SB), NOSPLIT, $0-16
RET RET
// func dirtyR14(a int64) int64 // func dirtyR14(a int64) int64
// Deliberately clobbers R14 (the goroutine pointer — a serious ABI violation). // Deliberately clobbers R14 (the goroutine pointer; a serious ABI violation).
TEXT ·dirtyR14(SB), NOSPLIT, $0-16 TEXT ·dirtyR14(SB), NOSPLIT, $0-16
MOVQ $0x5678, R14 MOVQ $0x5678, R14
MOVQ a+0(FP), AX MOVQ a+0(FP), AX
+30 -30
View File
@@ -13,55 +13,55 @@ TEXT ·add(SB), NOSPLIT, $0-24
// func sum(data []int64) int64 // func sum(data []int64) int64
// Sums all elements of the slice. // Sums all elements of the slice.
TEXT ·sum(SB), NOSPLIT, $0-32 TEXT ·sum(SB), NOSPLIT, $0-32
MOVQ data_base+0(FP), SI MOVQ data_base+0(FP), SI
MOVQ data_len+8(FP), CX MOVQ data_len+8(FP), CX
XORQ AX, AX XORQ AX, AX
TESTQ CX, CX TESTQ CX, CX
JZ sum_done JZ sum_done
sum_loop: sum_loop:
ADDQ (SI), AX ADDQ (SI), AX
ADDQ $8, SI ADDQ $8, SI
DECQ CX DECQ CX
JNZ sum_loop JNZ sum_loop
sum_done: sum_done:
MOVQ AX, ret+24(FP) MOVQ AX, ret+24(FP)
RET RET
// func wideCopy(dst, src []byte) // func wideCopy(dst, src []byte)
// Non-overlapping copy of min(len(dst), len(src)) bytes using 32-byte moves. // Non-overlapping copy of min(len(dst), len(src)) bytes using 32-byte moves.
TEXT ·wideCopy(SB), NOSPLIT, $0-48 TEXT ·wideCopy(SB), NOSPLIT, $0-48
MOVQ dst_base+0(FP), DI MOVQ dst_base+0(FP), DI
MOVQ dst_len+8(FP), BX MOVQ dst_len+8(FP), BX
MOVQ src_base+24(FP), SI MOVQ src_base+24(FP), SI
MOVQ src_len+32(FP), R8 MOVQ src_len+32(FP), R8
CMPQ BX, R8 CMPQ BX, R8
JLE wc_have_n JLE wc_have_n
MOVQ R8, BX MOVQ R8, BX
wc_have_n: wc_have_n:
CMPQ BX, $32 CMPQ BX, $32
JB wc_small JB wc_small
VMOVDQU (SI), Y0 VMOVDQU (SI), Y0
VMOVDQU Y0, (DI) VMOVDQU Y0, (DI)
VMOVDQU -32(SI)(BX*1), Y0 VMOVDQU -32(SI)(BX*1), Y0
VMOVDQU Y0, -32(DI)(BX*1) VMOVDQU Y0,-32(DI)(BX*1)
VZEROUPPER VZEROUPPER
RET RET
wc_small: wc_small:
TESTQ BX, BX TESTQ BX, BX
JZ wc_done JZ wc_done
wc_byte: wc_byte:
MOVB (SI), R8B MOVB (SI), R8B
MOVB R8B, (DI) MOVB R8B, (DI)
INCQ SI INCQ SI
INCQ DI INCQ DI
DECQ BX DECQ BX
JNZ wc_byte JNZ wc_byte
wc_done: wc_done:
RET RET
+32 -29
View File
@@ -5,53 +5,56 @@
// add returns a + b. // add returns a + b.
TEXT ·add(SB), NOSPLIT, $0-24 TEXT ·add(SB), NOSPLIT, $0-24
MOVD a+0(FP), R4 MOVD a+0(FP), R4
MOVD b+8(FP), R5 MOVD b+8(FP), R5
ADD R5, R4, R4 ADD R5, R4, R4
MOVD R4, ret+16(FP) MOVD R4, ret+16(FP)
RET RET
// arith exercises the register-register integer set. // arith exercises the register-register integer set.
TEXT ·arith(SB), NOSPLIT, $0-0 TEXT ·arith(SB), NOSPLIT, $0-0
ADD R4, R5, R6 ADD R4, R5, R6
SUB R7, R8, R9 SUB R7, R8, R9
AND R10, R11, R12 AND R10, R11, R12
ORR R12, R13, R14 ORR R12, R13, R14
EOR R14, R15, R16 EOR R14, R15, R16
CMP R16, R17 CMP R16, R17
ADD R4, R5 ADD R4, R5
SUB R6, R7 SUB R6, R7
RET RET
// branch exercises conditional and unconditional control flow. // branch exercises conditional and unconditional control flow.
TEXT ·branch(SB), NOSPLIT, $0-0 TEXT ·branch(SB), NOSPLIT, $0-0
BEQ done BEQ done
BNE skip BNE skip
BGE done BGE done
BLT done BLT done
BGT done BGT done
BLE done BLE done
skip: skip:
B loop B loop
loop: loop:
ADD R4, R5 ADD R4, R5
RET RET
done: done:
RET RET
// mov exercises the MOV pseudo-instruction. // mov exercises the MOV pseudo-instruction.
TEXT ·mov(SB), NOSPLIT, $0-16 TEXT ·mov(SB), NOSPLIT, $0-16
MOVD $0, R4 MOVD $0, R4
MOVD $1, R5 MOVD $1, R5
MOVD $42, R6 MOVD $42, R6
MOVD a+0(FP), R7 MOVD a+0(FP), R7
MOVD R7, ret+0(FP) MOVD R7, ret+0(FP)
MOVW $100, R8 MOVW $100, R8
RET RET
// frame exercises the prologue/epilogue of a function with a real frame. // frame exercises the prologue/epilogue of a function with a real frame.
TEXT ·frame(SB), NOSPLIT, $32-8 TEXT ·frame(SB), NOSPLIT, $32-8
MOVD arg+0(FP), R4 MOVD arg+0(FP), R4
ADD $1, R4, R4 ADD $1, R4, R4
MOVD R4, ret+0(FP) MOVD R4, ret+0(FP)
RET RET
+33 -30
View File
@@ -5,50 +5,53 @@
// add returns a + b. // add returns a + b.
TEXT ·add(SB), NOSPLIT, $0-24 TEXT ·add(SB), NOSPLIT, $0-24
MOVV a+0(FP), R4 MOVV a+0(FP), R4
MOVV b+8(FP), R5 MOVV b+8(FP), R5
ADDV R5, R4, R4 ADDV R5, R4, R4
MOVV R4, ret+16(FP) MOVV R4, ret+16(FP)
RET RET
// arith exercises the 3R integer and FP set. // arith exercises the 3R integer and FP set.
TEXT ·arith(SB), NOSPLIT, $0-0 TEXT ·arith(SB), NOSPLIT, $0-0
ADDV R4, R5, R6 ADDV R4, R5, R6
SUBV R7, R8, R9 SUBV R7, R8, R9
MULV R10, R11, R12 MULV R10, R11, R12
DIVV R13, R14, R15 DIVV R13, R14, R15
AND R16, R17, R18 AND R16, R17, R18
OR R18, R19, R20 OR R18, R19, R20
XOR R20, R21, R2 XOR R20, R21, R2
SLLV R2, R23, R24 SLLV R2, R23, R24
SRLV R24, R25, R26 SRLV R24, R25, R26
SRAV R26, R27, R28 SRAV R26, R27, R28
RET RET
// imm exercises the immediate forms. // imm exercises the immediate forms.
TEXT ·imm(SB), NOSPLIT, $0-0 TEXT ·imm(SB), NOSPLIT, $0-0
ADDV $42, R4, R5 ADDV $42, R4, R5
ADDV $-8, R6 ADDV $-8, R6
AND $0xff, R7, R8 AND $0xff, R7, R8
OR $1, R9, R10 OR $1, R9, R10
XOR $0, R11, R12 XOR $0, R11, R12
SGT $100, R13, R14 SGT $100, R13, R14
SLLV $4, R15, R16 SLLV $4, R15, R16
MOVV $0x12345, R17 MOVV $0x12345, R17
RET RET
// branch exercises conditional and unconditional control flow. // branch exercises conditional and unconditional control flow.
TEXT ·branch(SB), NOSPLIT, $0-0 TEXT ·branch(SB), NOSPLIT, $0-0
BEQ R4, R5, done BEQ R4, R5, done
BNE R6, R7, skip BNE R6, R7, skip
BLT R8, R9, done BLT R8, R9, done
BGE R10, R11, done BGE R10, R11, done
BLTU R12, R13, done BLTU R12, R13, done
BGEU R14, R15, done BGEU R14, R15, done
skip: skip:
JMP loop JMP loop
loop: loop:
JAL skip JAL skip
RET RET
done: done:
RET RET
+19 -16
View File
@@ -5,23 +5,26 @@
// branch exercises all conditional branch forms and jump chain folding. // branch exercises all conditional branch forms and jump chain folding.
TEXT ·branch(SB), NOSPLIT, $0-0 TEXT ·branch(SB), NOSPLIT, $0-0
BEQ done BEQ done
BNE skip BNE skip
BGE done BGE done
BLT done BLT done
BGT done BGT done
BLE done BLE done
BCS done BCS done
BCC done BCC done
BMI done BMI done
BPL done BPL done
BVS done BVS done
BVC done BVC done
BHI done BHI done
BLS done BLS done
skip: skip:
B loop B loop
loop: loop:
ADD R4, R5 ADD R4, R5
done: done:
RET RET
+14 -6
View File
@@ -5,31 +5,39 @@
TEXT ·branches(SB), NOSPLIT, $0 TEXT ·branches(SB), NOSPLIT, $0
ADDI $1, X10, X10 ADDI $1, X10, X10
BEQ X10, X11, beq_done BEQ X10, X11, beq_done
ADDI $2, X10, X10 ADDI $2, X10, X10
beq_done: beq_done:
BNE X10, X11, bne_done BNE X10, X11, bne_done
ADDI $3, X10, X10 ADDI $3, X10, X10
bne_done: bne_done:
BLT X10, X11, blt_done BLT X10, X11, blt_done
ADDI $4, X10, X10 ADDI $4, X10, X10
blt_done: blt_done:
BGE X10, X11, bge_done BGE X10, X11, bge_done
ADDI $5, X10, X10 ADDI $5, X10, X10
bge_done: bge_done:
BLTU X10, X11, bltu_done BLTU X10, X11, bltu_done
ADDI $6, X10, X10 ADDI $6, X10, X10
bltu_done: bltu_done:
BGEU X10, X11, bgeu_done BGEU X10, X11, bgeu_done
ADDI $7, X10, X10 ADDI $7, X10, X10
bgeu_done: bgeu_done:
RET RET
TEXT ·jumps(SB), NOSPLIT, $0 TEXT ·jumps(SB), NOSPLIT, $0
JMP done JMP done
ADDI $1, X10, X10 ADDI $1, X10, X10
done: done:
JAL X11, skip JAL X11, skip
ADDI $2, X10, X10 ADDI $2, X10, X10
skip: skip:
RET RET
+2 -2
View File
@@ -5,6 +5,6 @@
// caller exercises BL to an external symbol (produces a relocation). // caller exercises BL to an external symbol (produces a relocation).
TEXT ·caller(SB), NOSPLIT, $0-0 TEXT ·caller(SB), NOSPLIT, $0-0
BL other(SB) BL other(SB)
ADD R4, R5 ADD R4, R5
RET RET
+87 -87
View File
@@ -5,115 +5,115 @@
// fparith exercises the FP arithmetic set. // fparith exercises the FP arithmetic set.
TEXT ·fparith(SB), NOSPLIT, $0-0 TEXT ·fparith(SB), NOSPLIT, $0-0
FADDD F0, F1, F2 FADDD F0, F1, F2
FSUBD F3, F4, F5 FSUBD F3, F4, F5
FMULD F6, F7, F8 FMULD F6, F7, F8
FDIVD F9, F10, F11 FDIVD F9, F10, F11
FADDS F12, F13, F14 FADDS F12, F13, F14
FSUBS F15, F16, F17 FSUBS F15, F16, F17
FMULS F18, F19, F20 FMULS F18, F19, F20
FDIVS F21, F22, F23 FDIVS F21, F22, F23
FSQRTD F24, F25 FSQRTD F24, F25
FSQRTS F26, F27 FSQRTS F26, F27
FNEGD F28, F29 FNEGD F28, F29
FNEGS F30, F31 FNEGS F30, F31
FABSD F0, F1 FABSD F0, F1
FABSS F2, F3 FABSS F2, F3
FNMULD F4, F5, F6 FNMULD F4, F5, F6
FNMULS F7, F8, F9 FNMULS F7, F8, F9
FMIND F10, F11, F12 FMIND F10, F11, F12
FMAXD F13, F14, F15 FMAXD F13, F14, F15
FMINS F16, F17, F18 FMINS F16, F17, F18
FMAXS F19, F20, F21 FMAXS F19, F20, F21
RET RET
// fpfma exercises fused multiply-add. // fpfma exercises fused multiply-add.
TEXT ·fpfma(SB), NOSPLIT, $0-0 TEXT ·fpfma(SB), NOSPLIT, $0-0
FMADDD F0, F1, F2, F3 FMADDD F0, F1, F2, F3
FMSUBD F4, F5, F6, F7 FMSUBD F4, F5, F6, F7
FNMADDD F8, F9, F10, F11 FNMADDD F8, F9, F10, F11
FNMSUBD F12, F13, F14, F15 FNMSUBD F12, F13, F14, F15
FMADDS F16, F17, F18, F19 FMADDS F16, F17, F18, F19
FMSUBS F20, F21, F22, F23 FMSUBS F20, F21, F22, F23
FNMADDS F24, F25, F26, F27 FNMADDS F24, F25, F26, F27
FNMSUBS F28, F29, F30, F0 FNMSUBS F28, F29, F30, F0
RET RET
// fpconv exercises FP↔integer conversion and cross-precision. // fpconv exercises FP↔integer conversion and cross-precision.
// Syntax: FCVTZSD Fd, Rn (float→int: FP source first, int dest second) // Syntax: FCVTZSD Fd, Rn (float→int: FP source first, int dest second)
// SCVTFD Rn, Fd (int→float: int source first, FP dest second) // SCVTFD Rn, Fd (int→float: int source first, FP dest second)
TEXT ·fpconv(SB), NOSPLIT, $0-0 TEXT ·fpconv(SB), NOSPLIT, $0-0
FCVTSD F0, F1 FCVTSD F0, F1
FCVTDS F2, F3 FCVTDS F2, F3
FCVTZSD F4, R0 FCVTZSD F4, R0
FCVTZSS F5, R1 FCVTZSS F5, R1
FCVTZUD F6, R2 FCVTZUD F6, R2
FCVTZUS F7, R3 FCVTZUS F7, R3
SCVTFD R4, F8 SCVTFD R4, F8
SCVTFS R5, F9 SCVTFS R5, F9
UCVTFD R6, F10 UCVTFD R6, F10
UCVTFS R7, F11 UCVTFS R7, F11
SCVTFWD R0, F12 SCVTFWD R0, F12
SCVTFWS R1, F13 SCVTFWS R1, F13
UCVTFWD R2, F14 UCVTFWD R2, F14
UCVTFWS R3, F15 UCVTFWS R3, F15
FMOVS F14, R20 FMOVS F14, R20
FMOVS R21, F15 FMOVS R21, F15
FMOVD F16, R22 FMOVD F16, R22
FMOVD R23, F17 FMOVD R23, F17
RET RET
// fpcmp exercises FP compare and conditional compare. // fpcmp exercises FP compare and conditional compare.
// FCCMP syntax: FCCMP cond, Fn, Fm, $nzcv // FCCMP syntax: FCCMP cond, Fn, Fm, $nzcv
// FCSEL syntax: FCSEL cond, Fn, Fm, Fd // FCSEL syntax: FCSEL cond, Fn, Fm, Fd
TEXT ·fpcmp(SB), NOSPLIT, $0-0 TEXT ·fpcmp(SB), NOSPLIT, $0-0
FCMPS F0, F1 FCMPS F0, F1
FCMPD F2, F3 FCMPD F2, F3
FCMPS $0.0, F4 FCMPS $0.0, F4
FCMPD $0.0, F5 FCMPD $0.0, F5
FCCMPS EQ, F6, F7, $0 FCCMPS EQ, F6, F7, $0
FCCMPD NE, F8, F9, $0 FCCMPD NE, F8, F9, $0
FCSELS GE, F10, F11, F12 FCSELS GE, F10, F11, F12
FCSELD LT, F13, F14, F15 FCSELD LT, F13, F14, F15
RET RET
// frint exercises FP rounding. // frint exercises FP rounding.
TEXT ·frint(SB), NOSPLIT, $0-0 TEXT ·frint(SB), NOSPLIT, $0-0
FRINTND F0, F1 FRINTND F0, F1
FRINTNS F2, F3 FRINTNS F2, F3
FRINTPD F4, F5 FRINTPD F4, F5
FRINTPS F6, F7 FRINTPS F6, F7
FRINTMD F8, F9 FRINTMD F8, F9
FRINTMS F10, F11 FRINTMS F10, F11
FRINTZD F12, F13 FRINTZD F12, F13
FRINTZS F14, F15 FRINTZS F14, F15
FRINTAD F16, F17 FRINTAD F16, F17
FRINTAS F18, F19 FRINTAS F18, F19
FRINTXD F20, F21 FRINTXD F20, F21
FRINTXS F22, F23 FRINTXS F22, F23
FRINTID F24, F25 FRINTID F24, F25
FRINTIS F26, F27 FRINTIS F26, F27
FMOVD F0, F1 FMOVD F0, F1
FMOVS F2, F3 FMOVS F2, F3
RET RET
// condsel exercises conditional select and CRC32. // condsel exercises conditional select and CRC32.
TEXT ·condsel(SB), NOSPLIT, $0-0 TEXT ·condsel(SB), NOSPLIT, $0-0
CSEL EQ, R0, R1, R2 CSEL EQ, R0, R1, R2
CSINC NE, R3, R4, R5 CSINC NE, R3, R4, R5
CSINV GE, R6, R7, R8 CSINV GE, R6, R7, R8
CSNEG LT, R9, R10, R11 CSNEG LT, R9, R10, R11
CSET EQ, R12 CSET EQ, R12
CSETM NE, R13 CSETM NE, R13
CINC EQ, R14, R15 CINC EQ, R14, R15
CINV NE, R16, R17 CINV NE, R16, R17
CNEG GE, R19, R20 CNEG GE, R19, R20
CRC32B R0, R2 CRC32B R0, R2
CRC32H R3, R5 CRC32H R3, R5
CRC32W R6, R8 CRC32W R6, R8
CRC32X R9, R11 CRC32X R9, R11
CRC32CB R12, R14 CRC32CB R12, R14
CRC32CH R15, R0 CRC32CH R15, R0
CRC32CW R1, R3 CRC32CW R1, R3
CRC32CX R4, R6 CRC32CX R4, R6
RET RET
+106 -90
View File
@@ -6,138 +6,154 @@
// fp exercises the floating-point set: 3R arithmetic, 2R unary, compares // fp exercises the floating-point set: 3R arithmetic, 2R unary, compares
// into FCC, fused multiply-add and the register moves. // into FCC, fused multiply-add and the register moves.
TEXT ·fp(SB), NOSPLIT, $0-0 TEXT ·fp(SB), NOSPLIT, $0-0
ADDD F4, F5, F6 ADDD F4, F5, F6
SUBD F7, F8, F9 SUBD F7, F8, F9
MULD F9, F10, F11 MULD F9, F10, F11
DIVD F11, F12, F13 DIVD F11, F12, F13
MULF F13, F14, F15 MULF F13, F14, F15
ADDF F15, F16, F17 ADDF F15, F16, F17
SQRTD F17, F18 SQRTD F17, F18
SQRTF F18, F19 SQRTF F18, F19
ABSD F19, F20 ABSD F19, F20
NEGD F20, F21 NEGD F20, F21
MOVD F21, F22 MOVD F21, F22
CMPEQD F22, F23, FCC0 CMPEQD F22, F23, FCC0
CMPGTF F23, F24, FCC1 CMPGTF F23, F24, FCC1
CMPGED F24, F25, FCC2 CMPGED F24, F25, FCC2
FMADDD F0, F1, F2, F3 FMADDD F0, F1, F2, F3
FMSUBF F3, F4, F5, F6 FMSUBF F3, F4, F5, F6
FNMADDD F6, F7, F8, F9 FNMADDD F6, F7, F8, F9
FNMSUBF F9, F10, F11, F12 FNMSUBF F9, F10, F11, F12
FMAXD F12, F13, F14 FMAXD F12, F13, F14
FMINF F14, F15, F16 FMINF F14, F15, F16
FMAXAD F16, F17, F18 FMAXAD F16, F17, F18
FMINAF F18, F19, F20 FMINAF F18, F19, F20
FSCALEBF F20, F21, F22 FSCALEBF F20, F21, F22
FCOPYSGD F22, F23, F24 FCOPYSGD F22, F23, F24
MOVV F25, R25 MOVV F25, R25
MOVV R26, F27 MOVV R26, F27
MOVW R28, F29 MOVW R28, F29
MOVW F30, R31 MOVW F30, R31
RET RET
// mov forms: register moves, immediates (12/32/64-bit), memory with FP/SP // mov forms: register moves, immediates (12/32/64-bit), memory with FP/SP
// pseudo-registers and the register-indexed forms. // pseudo-registers and the register-indexed forms.
TEXT ·mov(SB), NOSPLIT, $0-16 TEXT ·mov(SB), NOSPLIT, $0-16
MOVV R4, R5 MOVV R4, R5
MOVW R6, R7 MOVW R6, R7
MOVB R8, R9 MOVB R8, R9
MOVBU R10, R11 MOVBU R10, R11
MOVHU R12, R13 MOVHU R12, R13
MOVWU R14, R15 MOVWU R14, R15
MOVV $42, R16 MOVV $42, R16
MOVV $0x12345, R17 MOVV $0x12345, R17
MOVV $0x100000, R18 MOVV $0x100000, R18
MOVW $-100, R19 MOVW $-100, R19
MOVV $0x123456789, R20 MOVV $0x123456789, R20
MOVV a+0(FP), R21 MOVV a+0(FP), R21
MOVV R23, b+8(FP) MOVV R23, b+8(FP)
MOVW c+16(FP), R24 MOVW c+16(FP), R24
MOVV (R24)(R25), R26 MOVV (R24)(R25), R26
MOVV R27, (R28)(R29) MOVV R27, (R28)(R29)
RET RET
// frame exercises the prologue/epilogue of a function with a real frame. // frame exercises the prologue/epilogue of a function with a real frame.
TEXT ·frame(SB), NOSPLIT, $32-8 TEXT ·frame(SB), NOSPLIT, $32-8
MOVV R4, R5 MOVV R4, R5
MOVV arg+0(FP), R6 MOVV arg+0(FP), R6
MOVV R7, local-8(SP) MOVV R7, local-8(SP)
MOVV local-8(SP), R8 MOVV local-8(SP), R8
MOVV R9, ret+0(FP) MOVV R9, ret+0(FP)
RET RET
// branches21 exercises the single-register and zero-register branch forms // branches21 exercises the single-register and zero-register branch forms
// with 21-bit offsets. // with 21-bit offsets.
TEXT ·branches21(SB), NOSPLIT, $0-0 TEXT ·branches21(SB), NOSPLIT, $0-0
BEQ R0, R4, l1 BEQ R0, R4, l1
BEQ R5, R0, l2 BEQ R5, R0, l2
BNE R0, R6, l3 BNE R0, R6, l3
BNE R7, R0, l4 BNE R7, R0, l4
BLTZ R8, l5 BLTZ R8, l5
BGEZ R9, l6 BGEZ R9, l6
BLEZ R10, l7 BLEZ R10, l7
BGTZ R11, l8 BGTZ R11, l8
JMP l9 JMP l9
l1: l1:
JMP l10 JMP l10
l2: l2:
JMP l11 JMP l11
l3: l3:
JMP l12 JMP l12
l4: l4:
JMP l13 JMP l13
l5: l5:
JMP l14 JMP l14
l6: l6:
JMP l15 JMP l15
l7: l7:
JMP l16 JMP l16
l8: l8:
JMP l16 JMP l16
l9: l9:
MOVV R1, R2 MOVV R1, R2
l10: l10:
LL (R12), R13 LL (R12), R13
LLV (R14), R15 LLV (R14), R15
SC R16, (R17) SC R16, (R17)
SCV R18, (R19) SCV R18, (R19)
RDTIMED R20, R21 RDTIMED R20, R21
SYSCALL SYSCALL
DBAR DBAR
RET RET
l11: l11:
JAL (R30) JAL (R30)
RET RET
l12: l12:
BSTRINSV $7, R4, $0, R5 BSTRINSV $7, R4, $0, R5
BSTRPICKV $63, R6, $32, R7 BSTRPICKV $63, R6, $32, R7
ALSLV $2, R8, R9, R10 ALSLV $2, R8, R9, R10
ADDV16 $65536, R11, R12 ADDV16 $65536, R11, R12
RET RET
l13: l13:
MOVV $0xffffffffffffffff, R13 MOVV $0xffffffffffffffff, R13
RET RET
l14: l14:
CPUCFG R14, R14 CPUCFG R14, R14
RET RET
l15: l15:
NOR R15, R16, R17 NOR R15, R16, R17
ORN R18, R19, R20 ORN R18, R19, R20
ANDN R21, R24, R25 ANDN R21, R24, R25
RET RET
l16: l16:
MOVB R26, (R27) MOVB R26, (R27)
MOVB (R28), R29 MOVB (R28), R29
RET RET
// sbdata loads and stores a static symbol with relocations (the relocation // sbdata loads and stores a static symbol with relocations (the relocation
// fields are masked before comparison). // fields are masked before comparison).
GLOBL ·table(SB), RODATA, $16 GLOBL ·table(SB), RODATA, $16
DATA ·table+0(SB)/8, $0x1122334455667788 DATA ·table+0(SB)/8, $0x1122334455667788
DATA ·table+8(SB)/8, $0x8877665544332211 DATA ·table+8(SB)/8, $0x8877665544332211
TEXT ·sbdata(SB), NOSPLIT, $0-0 TEXT ·sbdata(SB), NOSPLIT, $0-0
MOVV $·table(SB), R4 MOVV $·table(SB), R4
MOVV ·table(SB), R5 MOVV ·table(SB), R5
MOVV R6, ·table+8(SB) MOVV R6, ·table+8(SB)
RET RET
+18
View File
@@ -0,0 +1,18 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Indirect control flow: JMP/CALL through a register or memory, byte-compared
// against go tool asm. A CALL in the body also exercises the toolchain's
// forced base-pointer frame on a frameless function.
#include "textflag.h"
// func f()
TEXT ·f(SB), NOSPLIT, $0
JMP AX
CALL AX
JMP (BX)
CALL (BX)
JMP 8(BX)
JMP R8
RET
+14
View File
@@ -0,0 +1,14 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Indirect control flow: JMP (R0) and CALL (R0) lower to BR/BLR, the only
// indirect-branch spellings the toolchain accepts (the raw BR/BLR mnemonics
// stay a gasm superset).
#include "textflag.h"
// func f()
TEXT ·f(SB), NOSPLIT, $0
JMP (R0)
CALL (R0)
RET
+14
View File
@@ -0,0 +1,14 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Indirect control flow: JMP (R4) and JAL (R5) are the toolchain's spellings
// for jirl; the raw JIRL instruction is deliberately absent, because the Go
// loong64 assembler deletes it and a parity kernel could not hold it.
#include "textflag.h"
// func f()
TEXT ·f(SB), NOSPLIT, $0-0
JMP (R4)
JAL (R5)
RET
+19
View File
@@ -0,0 +1,19 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Indirect control flow: JMP (X5) lowers to JALR X0, 0(X5), the trampoline
// form JALR rd, offset(rs1) encodes with the destination first, and a linking
// JALR through X1 is the toolchain's only indirect call.
#include "textflag.h"
// func f()
TEXT ·leaf(SB), NOSPLIT, $0-0
JMP (X5)
JALR X0, 0(X6)
RET
// func g()
TEXT ·calls(SB), NOSPLIT, $0-0
JALR X1, 0(X8)
RET
+1 -1
View File
@@ -7,6 +7,6 @@ TEXT ·largeimm(SB), NOSPLIT, $0
ADDI $2048, X5 ADDI $2048, X5
ADDI $4095, X5, X6 ADDI $4095, X5, X6
ANDI $4095, X5, X6 ANDI $4095, X5, X6
ORI $-4096, X5, X6 ORI $-4096, X5, X6
XORI $0x12345, X5, X6 XORI $0x12345, X5, X6
RET RET
+8 -8
View File
@@ -4,14 +4,14 @@
#include "textflag.h" #include "textflag.h"
TEXT ·ldst(SB), NOSPLIT, $0 TEXT ·ldst(SB), NOSPLIT, $0
LD (X8), X9 LD (X8), X9
SD X9, (X8) SD X9, (X8)
LW (X8), X9 LW (X8), X9
SW X9, (X8) SW X9, (X8)
LD 8(X2), X10 LD 8(X2), X10
SD X10, 16(X2) SD X10, 16(X2)
LW 4(X2), X11 LW 4(X2), X11
SW X11, 8(X2) SW X11, 8(X2)
RET RET
TEXT ·addi4spn(SB), NOSPLIT, $0 TEXT ·addi4spn(SB), NOSPLIT, $0
+36
View File
@@ -0,0 +1,36 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// GOROOT-derived shapes: the MOV width suffixes for narrow loads and the
// branch-zero pseudos. Deliberately absent: the immediate ALU aliases
// (AND/SUB $imm) and the FP memory forms, whose RVC compression the encoder
// does not reproduce yet, so a parity kernel could not hold them.
#include "textflag.h"
// func mix(x int64, y int64) int64
TEXT ·mix(SB), NOSPLIT, $0-24
MOV x+0(FP), X5
MOVWU 0(X5), X11
MOVB 1(X5), X12
MOVBU 2(X5), X13
ADD X11, X12, X14
ADD X13, X14, X15
MOV X15, ret+16(FP)
RET
// func branchy(n int64) int64
TEXT ·branchy(SB), NOSPLIT, $0-16
MOV n+0(FP), X5
BEQZ X5, zero
BNEZ X5, one
BLTZ X5, zero
BGEZ X5, one
zero:
MOV $0, X6
MOV X6, ret+8(FP)
RET
one:
MOV $1, X6
MOV X6, ret+8(FP)
RET

Some files were not shown because too many files have changed in this diff Show More