Compare commits

...
33 Commits
Author SHA1 Message Date
petrbalvin 4d01bb3ecf build: rename the module to sourcedock.dev/petrbalvin/gasm-sdk
Test / test (push) Successful in 4m18s
2026-09-26 11:08:43 +02:00
petrbalvin 332c63e440 ci: compile-gate FreeBSD in the test pipeline
Test / test (push) Successful in 4m10s
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:49 +02:00
petrbalvin b306c210c6 feat(debug): port the debugger to FreeBSD
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:40 +02:00
petrbalvin b9015e1c2e fix(verify): make the executable mapping build on FreeBSD
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:31 +02:00
petrbalvin 26c5008136 fix(cmd): honour //go:build in the corpus audit
Test / test (push) Successful in 2m32s
Assisted-by: GLM 5.3 Flash
2026-09-23 21:03:16 +02:00
petrbalvin 74d6b90d69 fix(asm): read the arm64 move-wide immediate as an unsigned pattern
Assisted-by: GLM 5.3 Flash
2026-09-23 21:03:03 +02:00
petrbalvin 7b11c62f53 fix(asm): resolve negative numeric PC-relative jumps
Assisted-by: GLM 5.3 Flash
2026-09-23 21:02:50 +02:00
petrbalvin 8eed54b3da feat(lsp): quick fixes for the textflag include and the argument area
Test / test (push) Successful in 2m56s
Assisted-by: GLM 5.3 Flash
2026-09-23 20:23:35 +02:00
petrbalvin 4be16dcdf5 feat(lsp): resolve symbols across workspace files
Assisted-by: GLM 5.3 Flash
2026-09-23 20:21:58 +02:00
petrbalvin ded9cabdf4 fix(lsp): apply each rename edit to its own document
Assisted-by: GLM 5.3 Flash
2026-09-23 20:19:21 +02:00
petrbalvin cf6bc6987e fix(ci): pass the upload file to curl, not its interpolation
Test / test (push) Successful in 2m26s
Release / gates (push) Successful in 2m25s
Release / build (amd64, linux) (push) Successful in 1m15s
Release / build (arm64, linux) (push) Successful in 1m16s
Release / build (loong64, linux) (push) Successful in 1m16s
Release / build (riscv64, linux) (push) Successful in 1m16s
Release / release (push) Successful in 34s
2026-09-22 01:31:49 +02:00
petrbalvin ff7b1452b1 docs: name 0.35.0 as the supported release
Test / test (push) Successful in 2m33s
Release / gates (push) Successful in 2m29s
Release / build (amd64, linux) (push) Successful in 1m18s
Release / build (arm64, linux) (push) Successful in 1m20s
Release / build (loong64, linux) (push) Successful in 1m17s
Release / build (riscv64, linux) (push) Successful in 1m26s
Release / release (push) Failing after 35s
2026-09-22 00:52:56 +02:00
petrbalvin 517c1cea25 chore: prepare release v0.35.0
Test / test (push) Successful in 2m33s
Release / gates (push) Failing after 46s
Release / build (amd64, linux) (push) Skipped
Release / build (arm64, linux) (push) Skipped
Release / build (loong64, linux) (push) Skipped
Release / build (riscv64, linux) (push) Skipped
Release / release (push) Skipped
2026-09-22 00:44:10 +02:00
petrbalvin a3e3010e0f fix(cmd): resolve the runtime header test GOROOT from the go command
Test / test (push) Successful in 2m39s
2026-09-21 22:46:07 +02:00
petrbalvin 057c4eb545 docs: complete the release delta in the changelog and readme 2026-09-21 22:45:56 +02:00
petrbalvin f720381d43 feat(asm): the segment-absolute and crash-store forms GOROOT writes
Test / test (push) Failing after 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-21 22:19:53 +02:00
petrbalvin 2c9042d62c feat(asm): PCALIGN alignment on amd64
Assisted-by: GLM 5.3 Flash
2026-09-21 22:00:30 +02:00
petrbalvin 82ef289d3a feat(asm): the immediate multiply and arm64 indirect branches GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:50:11 +02:00
petrbalvin 7246b0e002 feat(asm): the TLS access pair in the toolchain's one-instruction form
Assisted-by: GLM 5.3 Flash
2026-09-21 21:35:15 +02:00
petrbalvin 8cfd40aac8 feat(asm): the operand forms and defines GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:17:34 +02:00
petrbalvin 5382c9a8e4 feat(audit): list every corpus failure per architecture 2026-09-21 21:17:34 +02:00
petrbalvin 53de91b2df docs(asm): describe the four target architectures
Test / test (push) Failing after 2m23s
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin 8a36af7c7d docs(asm): generate the instruction appendices
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin e9789ce3f4 chore(arch): regenerate the instruction tables 2026-09-21 20:15:55 +02:00
petrbalvin 837231c068 docs(asm): open the assembly language reference
Assisted-by: GLM 5.3 Flash
2026-09-21 19:49:04 +02:00
petrbalvin 95025be1bc docs(changelog): describe the encoder entries by content
Test / test (push) Failing after 2m33s
2026-09-21 19:20:01 +02:00
petrbalvin 03a964bb2d docs(goobj): document the GOOBJ object file format 2026-09-21 19:19:53 +02:00
petrbalvin 123a16e346 docs(readme): state the documentation goal 2026-09-21 18:35:27 +02:00
petrbalvin 9701812bee docs: changelog for the completeness waves
Test / test (push) Failing after 3m6s
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin 29ac03468e feat(amd64): floating-point immediates through a synthesised pool
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin bfb7701db1 feat(amd64): emit the quad-register EVEX families
Assisted-by: GLM 5.3 Flash
2026-09-21 02:02:19 +02:00
petrbalvin e8b6ff5d7c fix(parser): fold a signed parenthesised displacement expression
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-21 00:45:33 +02:00
petrbalvin 1456907000 feat(riscv64,loong64): operand tail, float DATA and honest port classification
Assisted-by: GLM 5.3 Flash
2026-09-21 00:44:47 +02:00
163 changed files with 12465 additions and 474 deletions
+4 -1
View File
@@ -342,7 +342,10 @@ jobs:
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
q{-H}, q{Content-Type: application/octet-stream},
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
# The @ must not sit inside a qq{} string: there it starts an
# array interpolation and the upload body collapses to empty,
# which Gitea stores as a 201-created zero-byte attachment.
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
my $code = <$curl>;
+22
View File
@@ -56,6 +56,28 @@ jobs:
- name: Build
run: go build ./...
- name: FreeBSD build (amd64)
# The debugger's ptrace surface and the JIT substrate are the two
# FreeBSD-portable layers the tree carries; the forge has no FreeBSD
# runner, so a push can only compile-gate them. Running the ptrace
# suite needs real FreeBSD hardware.
env:
GOOS: freebsd
GOARCH: amd64
run: go build ./...
- name: FreeBSD build (arm64)
env:
GOOS: freebsd
GOARCH: arm64
run: go build ./...
- name: FreeBSD build (riscv64)
env:
GOOS: freebsd
GOARCH: riscv64
run: go build ./...
- name: Format
run: |
perl -e '
+168 -10
View File
@@ -1,6 +1,6 @@
# Changelog
All notable changes to gasm-devkit are documented here.
All notable changes to gasm-sdk are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
@@ -9,6 +9,87 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
- **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on
amd64, arm64 and riscv64: the same interactive surface as on Linux —
breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
register file), single-stepping, register and memory access — behind the
kernel's own ptrace requests, with tracee memory through `PT_IO` and
stop reports through `PT_LWPINFO`. The JIT substrate maps executable
memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD
too. The pipeline compile-gates all three architectures; live
validation awaits a FreeBSD machine.
- **Workspace-wide navigation in the language server.** `gasm lsp` indexes
the `.s` files under the workspace root beyond the documents the editor
has open, so go-to-definition, find references and workspace symbol search
reach files that were never opened. An open buffer always shadows its
disk copy, and watched-file events together with a per-query freshness
check keep the index current.
- **Quick fixes for the textflag include and the argument area.** The
`missing-textflag-include` warning offers to add the include after the
last one in the file, and the `abi-argsize` warning offers to set the
TEXT argument area to the size the `// func` signature implies, computed
by the new `lint.ExpectedArgSize`.
### Changed
- **The corpus audit assembles like the build.** A file's `//go:build`
constraint decides which target architectures attempt it: cpu_x86.s is
an x86 build alone, and the msan and goexperiment.runtimesecret trees
are compiled by no supported build, so they leave the measured set
instead of failing it. The headline now reads "assemble for every
applicable target": every real-code GOROOT assembly file, the tree
without testdata, assembles for all four architectures (250 of 250,
100 %); over the whole tree including testdata the measure is 271 of
322 (84.2 %).
- **The module moves to `sourcedock.dev/petrbalvin/gasm-sdk`.** The
repository and the module rename together with the product, now the
GAsm Software Development Kit. Fresh installs become
`go install sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm@latest`, and
installs pinned to the old `gasm-sdk` path stop resolving once the
repository takes the new name: reinstall from the new path. The
binary stays `gasm`.
### Fixed
- **Rename edits land in their own documents.** A rename collected the
ranges of every reference across the open documents but applied them all
to the document that started it, so renaming a symbol used in a second
file moved that file's text into the first. Each edit now applies to the
document it was collected in.
- **Negative numeric PC-relative jumps.** `JMP -3(PC)`, the shape the
runtime's exit loops write (sys_linux_amd64.s, sys_netbsd_amd64.s),
resolved to nothing: only the forward forms counted. A negative count
now walks the same instruction statements backwards, labels excluded,
byte-identical with the toolchain.
- **The arm64 move-wide family reads its immediate as an unsigned
pattern.** `MOVK $(40000<<48)` folds to a negative int64 and was
rejected; the toolchain picks the 16-bit lane from the 64-bit bit
pattern, so the encoder now does the same, and a zero immediate is
rejected where the toolchain rejects it.
## [0.35.0] - 2026-09-22
### Added
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
itself when an assembly file includes it: the Go files beside the source
are type-checked for the target architecture and the constants and field
offsets become assembler defines, so package-context files assemble with
no compiler and no `go build` in the loop. `-GOOS` selects the
type-checking GOOS for GOOS-specific files, and the corpus audit derives
the GOOS from the file name.
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
every architecture (amd64 carried them already), so standalone ELF
objects link on all four targets.
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
prints every failing file with its failure reason, per architecture,
instead of one representative file per reason.
- **DATA with symbol values and relaxed symbol spellings.** DATA
initialisers accept `$symbol(SB)` values, laid down as an absolute
relocation at the data field (GOOBJ on all four architectures and ELF
on all four as of this release), and U+2215 is accepted inside symbol
package paths.
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
`gasm audit-instructions` now preprocess assembly the way the
toolchain does: object and parameterised `#define` macros expand at
@@ -19,7 +100,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
language server keep reading the raw file.
- **The GOROOT instruction wave, part 1.** The encoder now covers the
- **Encoder coverage: the instruction families GOROOT's real code
uses.** The encoder now covers the
instruction families GOROOT's real code uses that gasm lacked,
byte-verified against `go tool asm`: on amd64 the carry ALU, the
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
@@ -35,14 +117,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
riscv64 `VSETVLI` with an immediate length now canonicalises to
`vsetivli` as the toolchain does.
- **The corpus audit measures honestly.** Files named for Go ports gasm
does not target (arm, 386, s390x, ...) are no longer attempted for the
four supported architectures (no supported build compiles them), and
the headline rate is reported over attemptable files: 136 of 433 on
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
127 that the previous release measured. The probe battery that
decides encodability gained the operand shapes the new families use.
-
- **Encoder coverage: quad-register AVX-512 and floating-point
immediates.** The encoder gains the
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
floating-point immediates on the SSE scalar moves and arithmetic
(the constant lands in a synthesised read-only pool, a positive zero
collapses to XORPS exactly as the toolchain does), accept-and-ignore
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
operands for the legacy SSE moves, and the pooled 64-bit immediate
materialisation on riscv64. The parser carries bracketed register
ranges, index-only VSIB memory operands and bare trailing immediates;
macro substitution reaches parameters used with element suffixes
(`A.S4`), and `;` separates statements in plain files.
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
roles the ABI fixes, addressing, operand order with every special form,
constants and materialisation, alignment, fences and the relocations
each target emits. An instruction inventory appendix per architecture
is generated from the toolchain's own tables by `just gen`, and the
regenerated tables recognise 147 more mnemonics than the previous
release carried (arm64 107, riscv64 31, loong64 9).
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
opens the complete language reference with its common core: the lexicon,
statement structure and constant expressions, the operand grammar with
the pseudo-registers and symbol naming, the directives and the function
flag vocabulary, preprocessing with `#define` and `#include`, and the
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
runtime contract). Every claim is verified against `go tool asm` of
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
generated instruction appendices follow.
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
documents the Go object file format in full: both containers, the 96
byte header and all 19 blocks, every structure with its byte
offsets, symbol kinds and flag bits, all 106 relocation types with
the weak variants, aux symbols, the FuncInfo payload, the pc-value
table encoding, the content hashes and the builtin table, all
verified byte for byte against objects produced by Go 1.27.1's own
tools.
### Changed
- **The corpus audit measures like a build.** Files named for a Go port
gasm does not target (arm, 386, s390x, ...) are never attempted, because
no supported build compiles them; the GOOS comes from the file name; and
each target's go_asm.h is generated on the fly. The headline is reported
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
for every target architecture and 295 of 303 on real code (97.4 %),
against 108 of 627 over all files (17.2 %) that the previous release
measured.
### Fixed
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
own instruction counting and fold jump-to-jump chains exactly as its
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
assemble to the toolchain's RIP-relative LEA with an R_PCREL
relocation; negated constant expressions in operands (`ADJSP
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
one-instruction form (the bare `MOVQ TLS, r` load nops out and
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
the zero-immediate store (`MOVD $0, mem` through the zero register,
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
aligns on amd64, padding with the toolchain's greedy
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
AX` and the store direction) and the absolute crash-store
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
GOROOT corpus measure moves to 291 of 353 files assembling for every
target architecture (82.4 %), 97.4 % of the real-code corpus, from
70.8 % and 82.2 %.
- **Tool corrections across the pipeline.** The formatter keeps square
brackets in SIMD operands, statement separators and canonical macro
bodies; the linter drops false positives on shift counts, SETcc
spellings and ABIInternal references; the lexer treats a trailing
carriage return as a line end so comment text stays idempotent; and
arm64 rejects bare BTI with a diagnostic while accepting the full
family.
## [0.34.0] - 2026-09-20
+4 -4
View File
@@ -1,6 +1,6 @@
# Contributing
Contributions to **gasm-devkit** are governed by the Contributor terms
Contributions to **gasm-sdk** are governed by the Contributor terms
below; submitting one means you accept them.
## Contributor terms
@@ -29,8 +29,8 @@ compiler (gcc), because `just gates` includes `just race` and the race
detector needs cgo.
```sh
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
cd gasm-devkit
git clone https://sourcedock.dev/petrbalvin/gasm-sdk.git
cd gasm-sdk
just build
just gates
```
@@ -126,7 +126,7 @@ tag, where it would double the time and the memory a shared runner cannot spare.
## Reporting bugs
Open an issue at `https://sourcedock.dev/petrbalvin/gasm-devkit/issues` with the
Open an issue at `https://sourcedock.dev/petrbalvin/gasm-sdk/issues` with the
version, the operating system and architecture, the exact command, the full output,
and the expected against the actual behaviour.
+66 -20
View File
@@ -1,6 +1,6 @@
# Plan 9 assembly tooling, inside and outside Go
# GAsm: Software Development Kit for Plan 9 Assembly
> **Warning: this is an experiment.** gasm-devkit is under active
> **Warning: this is an experiment.** gasm-sdk is under active
> development and is not stable. The version is 0.x.x: commands, flags,
> output formats and behaviour can change without warning at any time.
> A 1.0.0 release is light years away. Nothing in this document is a
@@ -13,7 +13,7 @@
there is no formatter, no linter and no debugger for `.s` files, and no
assembler that works without a Go installation. Developers write
assembly blind, validate it by benchmark, and debug it by print
statement. gasm-devkit is the missing toolkit: a single, self-contained
statement. gasm-sdk is the missing toolkit: a single, self-contained
binary, `gasm`, that serves both purposes.
- **Help develop Plan 9 assembly.** Formatting, linting, disassembly,
@@ -58,7 +58,7 @@ Plan 9 (Go): MOVQ AX, total-16(SP)
The same lines, but only one of them tells you what the number is for.
The syntax is uppercase, regular and boring, which is the highest
compliment a language for machine code can earn. gasm-devkit exists
compliment a language for machine code can earn. gasm-sdk exists
to give that syntax the tooling it deserves.
## Features
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
GOOBJ format, which needs the installed toolchain and which `go build`
consumes in place of the toolchain's output. Framed functions get the
stack-split guard and the morestack block, byte-identical to the
toolchain's, so split functions link too.
toolchain's, so split functions link too. The assembler preprocesses
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
offsets after assembling, or disassembles raw bytes from a file or stdin.
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
@@ -91,13 +94,16 @@ to give that syntax the tooling it deserves.
- **Debugger.** `gasm debug` is a source-level ptrace debugger with
breakpoints (optionally conditional), hardware watchpoints, register and
memory inspection, and headless script runs that report instruction and
label coverage.
label coverage; it runs on Linux (all four architectures) and FreeBSD
(amd64, arm64, riscv64).
- **Language server.** `gasm lsp` serves completion, hover, document symbols,
push and pull diagnostics, semantic-token highlighting, go-to-definition,
find references, rename, formatting, inlay hints, code actions, signature
help, document highlights, workspace symbol search, #include document
links and folding ranges over stdio; definition, references and rename
work across every open document.
work across every open document and the indexed workspace files beyond
them, and the quick fixes add a missing textflag.h include and set the
argument area from the // func signature.
- **Comparators and audits.** `gasm diff` compares the machine code of two
assembly files byte-for-byte, `gasm profile` shows basic-block structure,
`gasm audit-instructions` diffs the encoder against the installed toolchain,
@@ -110,9 +116,9 @@ Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
"Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
@@ -124,10 +130,13 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
(31.4 %) assembling for every target architecture today (files named for
other Go ports are counted but never attempted), with the top failure
reasons per architecture; the number moves with every release.
`gasm audit-instructions --corpus` reports every real-code GOROOT assembly
file (the tree without testdata) assembling for every target its build
admits: 250 of 250, 100 %. Over the whole tree including testdata the
measure is 271 of 322 attemptable (84.2 %); files named for other Go ports
are counted but never attempted, and `//go:build` constraints decide which
targets attempt a file at all, exactly as the build does. The number moves
with every release.
### Validation status
@@ -142,7 +151,7 @@ actually been executed.
|---|---|---|
| Encoding: byte-for-byte against `go tool asm` | native hardware | native hardware (the toolchain cross-assembles any GOARCH on any host) |
| Execution: JIT calls, ABI checks, differential fuzzing | native hardware | qemu-user emulation |
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more |
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64, riscv64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
Consequences, stated plainly. An emulator is a model of a CPU, not the
CPU: instruction semantics are implemented in software and can differ
@@ -157,6 +166,39 @@ been compiled and read, never executed. Its architecture-neutral units
run under `go test ./...`, which the race workflow and a manual run
perform; the default `just test` gate does not sweep `./debug/...`.
## The documentation goal
The toolkit is the primary goal. The secondary one is documentation: a
specification of the Plan 9 assembly language and of the GOOBJ object
format that is 100 % complete, detailed enough to implement against,
and written to a professional standard. These are the two subjects this
project works with every day, and they are the two for which no usable
documentation exists.
Go documents the language on a single page, "A Quick Guide to Go's
Assembler", which carries no section for loong64, one of the four
architectures gasm supports, and covers a fraction of what each
assembler accepts. What exists beyond it lives as comments inside the
toolchain's internal source: per-architecture reference manuals for
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
maintainers rather than for an outside reader, and none at all for
amd64. GOOBJ fares worst of all. The format that `go build` consumes
has no specification anywhere: it is described by a comment in an
internal package, it is not a stable interface, and it can change with
any toolchain release.
The gap is therefore filled the only way it can be filled: by reverse
engineering the toolchain itself, the same work the encoders already
perform. Most of the documentation can come from nowhere else, and it
is written as that knowledge is produced during development. It is
verified the way the code is verified: an encoding documented here is
one that differential tests against `go tool asm` confirm
byte-for-byte, and a format field documented here is one the linker
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
specifies the object file format completely, and
[docs/asm/README.md](docs/asm/README.md) opens the language reference
with its common core. The per-architecture pages follow.
## Direction
The plan, in the order it is being worked:
@@ -179,9 +221,11 @@ The plan, in the order it is being worked:
toolchain itself does not support; through ELF, Plan 9 assembly becomes
usable outside Go entirely.
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
architectures and is where the binary builds. FreeBSD follows: the
JIT's executable-memory mapping and the ptrace debugger layer are the
two pieces of porting work. Other unix systems may follow those two.
architectures and is where the binary builds. FreeBSD follows on amd64,
arm64 and riscv64: the JIT's executable-memory mapping and the ptrace
debugger layer are ported (the debugger's live validation awaits a
FreeBSD machine, as the validation status states). Other unix systems
may follow those two.
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
No others are planned.
@@ -189,11 +233,11 @@ The plan, in the order it is being worked:
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
linux/loong64 are on the
[releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases).
[releases page](https://sourcedock.dev/petrbalvin/gasm-sdk/releases).
From source (Go 1.27.1):
```sh
go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest
go install sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm@latest
```
Or from a repository checkout:
@@ -282,6 +326,8 @@ recipe.
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [CHANGELOG.md](CHANGELOG.md): release history
+1 -1
View File
@@ -7,7 +7,7 @@ releases do not receive them.
| Version | Supported |
|---|---|
| 0.34.0 | yes |
| 0.35.0 | yes |
| older releases | no |
## Reporting a vulnerability
+105 -7
View File
@@ -5,9 +5,13 @@
// toolchain's own assembler source. Go's Plan 9 assembler defines the exact,
// complete set of mnemonics it accepts for each architecture in
// $GOROOT/src/cmd/internal/obj/<arch>/anames.go; this tool extracts those
// names so gasm-devkit supports every instruction the real assembler does,
// names so gasm-sdk supports every instruction the real assembler does,
// with no hand-maintained (and therefore inevitably incomplete) lists.
//
// The same data feeds the generated instruction appendices of the assembly
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
// cannot drift from the tables it documents.
//
// Usage (via the justfile):
//
// just gen
@@ -26,9 +30,12 @@ import (
"path/filepath"
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
)
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
// archDirs maps a gasm-sdk architecture name to its obj sub-directory.
var archDirs = []struct {
arch string
sub string
@@ -39,11 +46,30 @@ var archDirs = []struct {
{"loong64", "loong64"},
}
// docPages maps an architecture to its generated appendix in the language
// reference. The amd64 page carries a per-mnemonic encodability column,
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
// other targets have no single cheap predicate, so their pages carry the
// inventory and point at the live measurement instead.
var docPages = []struct {
arch arch.Arch
title string
file string
anames string
encodable bool
}{
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
}
func main() {
goroot := strings.TrimSpace(runGoEnvGOROOT())
if goroot == "" {
fatal("could not determine GOROOT")
}
version := strings.TrimSpace(runGoEnv("GOVERSION"))
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
@@ -57,16 +83,24 @@ func main() {
}
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
names := map[string][]string{}
for _, a := range archDirs {
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
names, err := extractInstrs(path)
names[a.arch], err = extractInstrs(path)
if err != nil {
fatal("extract %s: %v", a.arch, err)
}
if err := writeGen(a.arch, a.sub, names); err != nil {
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
fatal("write %s: %v", a.arch, err)
}
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
}
for _, p := range docPages {
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
fatal("write %s: %v", p.file, err)
}
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
}
}
@@ -85,7 +119,7 @@ func filterCommon(names []string) []string {
// writeCommon emits arch/common_gen.go.
func writeCommon(names []string) error {
var b strings.Builder
b.WriteString("// Code generated by gasm-devkit _gen; DO NOT EDIT.\n")
b.WriteString("// Code generated by gasm-sdk _gen; DO NOT EDIT.\n")
b.WriteString("// Source: cmd/internal/obj/util.go from the Go toolchain.\n")
b.WriteString("//\n")
b.WriteString("// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)\n")
@@ -156,7 +190,7 @@ func stringLit(elt ast.Expr) string {
// writeGen emits arch/<arch>_gen.go.
func writeGen(arch, sub string, names []string) error {
var b strings.Builder
b.WriteString("// Code generated by gasm-devkit _gen; DO NOT EDIT.\n")
b.WriteString("// Code generated by gasm-sdk _gen; DO NOT EDIT.\n")
b.WriteString("// Source: cmd/internal/obj/" + sub + "/anames.go from the Go toolchain.\n")
b.WriteString("//\n")
b.WriteString("// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)\n")
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
}
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
// the language reference for one architecture: every mnemonic the toolchain
// accepts, with the curated summary where the architecture table carries one
// and, on amd64, a per-mnemonic encodability column.
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
table := arch.ForArch(a)
instrs := table.Instructions()
var b strings.Builder
b.WriteString("# " + title + ": instruction inventory\n\n")
b.WriteString("Generated by gasm-sdk's `_gen` from the Go toolchain's instruction table\n")
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
encodableCount := 0
if encodable {
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
b.WriteString("`gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
b.WriteString("|---|---|---|\n")
for _, in := range instrs {
ok := asm.Encodable(in.Name)
if ok {
encodableCount++
}
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
} else {
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
b.WriteString("encodability is decided per operand shape, and the live measured\n")
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | Notes |\n")
b.WriteString("|---|---|\n")
for _, in := range instrs {
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
}
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
}
// yesNo renders a boolean as the word the appendix tables use.
func yesNo(v bool) string {
if v {
return "yes"
}
return "no"
}
func runGoEnvGOROOT() string {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
return string(out)
}
// runGoEnv runs `go env` for a single variable.
func runGoEnv(name string) string {
out, err := exec.Command("go", "env", name).Output()
if err != nil {
return ""
}
return string(out)
}
func fatal(format string, args ...any) {
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
os.Exit(1)
+1 -1
View File
@@ -1,4 +1,4 @@
// Code generated by gasm-devkit _gen; DO NOT EDIT.
// Code generated by gasm-sdk _gen; DO NOT EDIT.
// Source: cmd/internal/obj/x86/anames.go from the Go toolchain.
//
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
+108 -1
View File
@@ -1,4 +1,4 @@
// Code generated by gasm-devkit _gen; DO NOT EDIT.
// Code generated by gasm-sdk _gen; DO NOT EDIT.
// Source: cmd/internal/obj/arm64/anames.go from the Go toolchain.
//
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
"REVW",
"ROR",
"RORW",
"RPRFM",
"SB",
"SBC",
"SBCS",
"SBCSW",
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
"UXTH",
"UXTHW",
"UXTW",
"VABS",
"VADD",
"VADDP",
"VADDV",
"VAND",
"VBCAX",
"VBIC",
"VBIF",
"VBIT",
"VBSL",
"VCLS",
"VCLZ",
"VCMEQ",
"VCMGE",
"VCMGT",
"VCMHI",
"VCMHS",
"VCMLE",
"VCMLT",
"VCMTST",
"VCNT",
"VDUP",
"VEOR",
"VEOR3",
"VEXT",
"VFABS",
"VFADD",
"VFADDP",
"VFCMEQ",
"VFCMGE",
"VFCMGT",
"VFCMLE",
"VFCMLT",
"VFCVTL",
"VFCVTL2",
"VFCVTN",
"VFCVTN2",
"VFCVTZS",
"VFCVTZU",
"VFDIV",
"VFMAX",
"VFMAXNM",
"VFMAXNMP",
"VFMAXNMV",
"VFMAXP",
"VFMAXV",
"VFMIN",
"VFMINNM",
"VFMINNMP",
"VFMINNMV",
"VFMINP",
"VFMINV",
"VFMLA",
"VFMLS",
"VFMUL",
"VFNEG",
"VFRINTM",
"VFRINTN",
"VFRINTP",
"VFRINTZ",
"VFSQRT",
"VFSUB",
"VLD1",
"VLD1R",
"VLD2",
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
"VLD3R",
"VLD4",
"VLD4R",
"VMLA",
"VMLS",
"VMOV",
"VMOVD",
"VMOVI",
"VMOVQ",
"VMOVS",
"VMUL",
"VNEG",
"VNOT",
"VORN",
"VORR",
"VPMULL",
"VPMULL2",
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
"VREV16",
"VREV32",
"VREV64",
"VSCVTF",
"VSHADD",
"VSHL",
"VSHRN",
"VSHRN2",
"VSLI",
"VSMAX",
"VSMAXP",
"VSMAXV",
"VSMIN",
"VSMINP",
"VSMINV",
"VSMLAL",
"VSMLAL2",
"VSMLSL",
"VSMLSL2",
"VSMULL",
"VSMULL2",
"VSQABS",
"VSQADD",
"VSQNEG",
"VSQSHL",
"VSQSUB",
"VSQXTN",
"VSQXTN2",
"VSQXTUN",
"VSQXTUN2",
"VSRHADD",
"VSRI",
"VSRSHR",
"VSSHL",
"VSSHLL",
"VSSHLL2",
"VSSHR",
"VST1",
"VST2",
"VST3",
"VST4",
"VSUB",
"VSXTL",
"VSXTL2",
"VTBL",
"VTBX",
"VTRN1",
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
"VUADDLV",
"VUADDW",
"VUADDW2",
"VUCVTF",
"VUHADD",
"VUMAX",
"VUMAXP",
"VUMAXV",
"VUMIN",
"VUMINP",
"VUMINV",
"VUMLAL",
"VUMLAL2",
"VUMLSL",
"VUMLSL2",
"VUMULL",
"VUMULL2",
"VUQADD",
"VUQSHL",
"VUQSUB",
"VUQXTN",
"VUQXTN2",
"VURHADD",
"VUSHL",
"VUSHLL",
"VUSHLL2",
"VUSHR",
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
"VUZP1",
"VUZP2",
"VXAR",
"VXTN",
"VXTN2",
"VZIP1",
"VZIP2",
"WFE",
+1 -1
View File
@@ -1,4 +1,4 @@
// Code generated by gasm-devkit _gen; DO NOT EDIT.
// Code generated by gasm-sdk _gen; DO NOT EDIT.
// Source: cmd/internal/obj/util.go from the Go toolchain.
//
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
+10 -1
View File
@@ -1,4 +1,4 @@
// Code generated by gasm-devkit _gen; DO NOT EDIT.
// Code generated by gasm-sdk _gen; DO NOT EDIT.
// Source: cmd/internal/obj/loong64/anames.go from the Go toolchain.
//
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
"FNMADDF",
"FNMSUBD",
"FNMSUBF",
"FRINTD",
"FRINTF",
"FSCALEBD",
"FSCALEBF",
"FSEL",
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
"FTINTWF",
"JIRL",
"LL",
"LLACQV",
"LLACQW",
"LLV",
"LLW",
"LU12IW",
"LU32ID",
"LU52ID",
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
"ROTR",
"ROTRV",
"SC",
"SCQ",
"SCRELV",
"SCRELW",
"SCV",
"SCW",
"SGT",
"SGTU",
"SLL",
+32 -1
View File
@@ -1,4 +1,4 @@
// Code generated by gasm-devkit _gen; DO NOT EDIT.
// Code generated by gasm-sdk _gen; DO NOT EDIT.
// Source: cmd/internal/obj/riscv/anames.go from the Go toolchain.
//
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
"CLD",
"CLDSP",
"CLI",
"CLMUL",
"CLMULH",
"CLMULR",
"CLUI",
"CLW",
"CLWSP",
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
"CSDSP",
"CSLLI",
"CSRAI",
"CSRC",
"CSRCI",
"CSRLI",
"CSRR",
"CSRRC",
"CSRRCI",
"CSRRS",
"CSRRSI",
"CSRRW",
"CSRRWI",
"CSRS",
"CSRSI",
"CSRW",
"CSRWI",
"CSUB",
"CSUBW",
"CSW",
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
"ORCB",
"ORI",
"ORN",
"PAUSE",
"RDCYCLE",
"RDINSTRET",
"RDTIME",
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
"VADDVI",
"VADDVV",
"VADDVX",
"VANDNVV",
"VANDNVX",
"VANDVI",
"VANDVV",
"VANDVX",
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
"VASUBUVX",
"VASUBVV",
"VASUBVX",
"VBREV8V",
"VBREVV",
"VCLMULHVV",
"VCLMULHVX",
"VCLMULVV",
"VCLMULVX",
"VCLZV",
"VCOMPRESSVM",
"VCPOPM",
"VCPOPV",
"VCTZV",
"VDIVUVV",
"VDIVUVX",
"VDIVVV",
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
"VREMUVX",
"VREMVV",
"VREMVX",
"VREV8V",
"VRGATHEREI16VV",
"VRGATHERVI",
"VRGATHERVV",
"VRGATHERVX",
"VROLVV",
"VROLVX",
"VRORVI",
"VRORVV",
"VRORVX",
"VRSUBVI",
"VRSUBVX",
"VS1RV",
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
"VWMULVX",
"VWREDSUMUVS",
"VWREDSUMVS",
"VWSLLVI",
"VWSLLVV",
"VWSLLVX",
"VWSUBUVV",
"VWSUBUVX",
"VWSUBUWV",
+1 -1
View File
@@ -11,7 +11,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestGOObjectAARCH64Structure checks the basic structure of the emitted
+38 -15
View File
@@ -9,7 +9,7 @@ import (
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// assembleARM64 assembles an AArch64 (arm64) TEXT function body into machine
@@ -637,6 +637,20 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
// The bare spelling BL R9 is the same indirect branch: the parser reads
// a bare identifier as a symbol, and one named for a register is an
// indirect branch through it, which the toolchain accepts alongside the
// parenthesised form (BL (R3) and BL R3 both encode BLR R3).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" && op.Addr.Base == "" && op.Addr.Index == "" {
if rn := arm64RegNum(op.Addr.Sym.Name); rn >= 0 {
opc := uint32(0) // BR
if link {
opc = 1 // BLR
}
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
}
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
// relocation (R_CALLARM64 either way).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
@@ -1454,6 +1468,15 @@ func encodeARM64Mov(instr *ast.Instr, mnem string, wb string, fi arm64FrameInfo,
}
return encodeARM64SBAddr(src.Imm.Sym, rd, relocs), nil
}
// Immediate → memory: only storing zero is encodable (the ZR
// register); the toolchain rejects any other immediate-to-memory
// combination ("illegal combination").
if isMemOperand(dst) {
if arm64Imm64(src) != 0 {
return nil, fmt.Errorf("%s: illegal combination: an immediate store must be zero", mnem)
}
return encodeARM64MemOp(mnem, dst, 31, false, fi, "")
}
rd := arm64RegNum(operandRegName(dst))
if rd < 0 {
return nil, fmt.Errorf("%s $imm: invalid destination register", mnem)
@@ -3046,29 +3069,29 @@ func encodeARM64MoveWide(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte
// base, so MOVZ and MOVN come along for free.
opc := baseOp >> 29 & 3
sf := baseOp >> 31 & 1
v := arm64Imm64(ops[0])
if v < 0 {
return nil, fmt.Errorf("%s: negative immediate %d", mnem, v)
// The toolchain's optab case 33, shared by the whole family in both
// widths: the immediate is one unsigned 64-bit pattern (a high-lane
// constant such as $(40000<<48) arrives negative through int64
// folding), it must occupy exactly one 16-bit lane, zero is rejected,
// and the W forms cannot reach the top half.
u := uint64(arm64Imm64(ops[0]))
if u == 0 {
return nil, fmt.Errorf("%s: zero immediate cannot be handled", mnem)
}
hw := -1
for i := range 4 {
if v>>(uint(i)*16)&0xFFFF != 0 {
hw = i
for lane := range 4 {
if u&^(uint64(0xFFFF)<<(lane*16)) == 0 {
hw = lane
break
}
}
if hw < 0 {
hw = 0 // zero: every chunk is zero, hw = 0 carries it
}
for i := hw + 1; i < 4; i++ {
if v>>(uint(i)*16)&0xFFFF != 0 {
return nil, fmt.Errorf("%s: immediate %d does not fit one 16-bit chunk", mnem, v)
}
return nil, fmt.Errorf("%s: immediate %#x does not fit one 16-bit chunk", mnem, u)
}
if sf == 0 && hw > 1 {
return nil, fmt.Errorf("%s: immediate %d out of range for the 32-bit form", mnem, v)
return nil, fmt.Errorf("%s: immediate %#x out of range for the 32-bit form", mnem, u)
}
return a64wordLE(a64MoveWide(sf, opc, uint32(hw), uint32(v>>uint(hw*16)&0xFFFF), uint32(rd))), nil
return a64wordLE(a64MoveWide(sf, opc, uint32(hw), uint32(u>>uint(hw*16)&0xFFFF), uint32(rd))), nil
}
// ---- Bitfield/EXTR encoding ----
+1 -1
View File
@@ -33,7 +33,7 @@ import (
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// arm64RegNum returns the 5-bit register number for an AArch64 register name:
+37 -2
View File
@@ -7,8 +7,8 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
func TestArm64LDRSTREncoding(t *testing.T) {
@@ -1051,6 +1051,41 @@ func TestArm64MOVK(t *testing.T) {
}
}
// TestArm64MOVKHighLane pins the shifted high-lane immediate the arm64 test
// kernels write: $(40000<<48) folds to a negative int64, and the toolchain
// reads the value as an unsigned 64-bit pattern when it picks the lane.
func TestArm64MOVKHighLane(t *testing.T) {
got := arm64Words(t, "\tMOVK $(40000<<48), R0\n\tMOVK $0x9c40000000000000, R1\n")
want := []uint32{
0xf2f38800, // MOVK $(40000<<48), R0 (go tool asm: f2f38800)
0xf2f38801, // MOVK hw=3
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64MoveWideZeroImmediate pins the toolchain's rejection of a zero
// immediate in the move-wide family (optab case 33: "zero shifts cannot be
// handled"): every lane is zero, so no hw field can carry it.
func TestArm64MoveWideZeroImmediate(t *testing.T) {
for _, mnem := range []string{"MOVK", "MOVZ", "MOVN"} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\t"+mnem+" $0, R0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("%s: parse: %v", mnem, errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s $0: expected error, got nil", mnem)
}
}
}
// TestArm64LoadImm64 tests 64-bit immediate loading.
func TestArm64LoadImm64(t *testing.T) {
src := `#include "textflag.h"
+1 -1
View File
@@ -53,7 +53,7 @@ package asm
import (
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// arm64FrameInfo holds the frame layout derived from a TEXT directive.
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// parseArm64File is a helper assembling one arm64 source file.
+500 -44
View File
@@ -5,9 +5,10 @@ package asm
import (
"fmt"
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// Assemble encodes the body of a TEXT function into x86-64 machine code,
@@ -28,7 +29,7 @@ import (
// emitted: the bytes match go tool asm only for NOSPLIT functions or
// zero-frame leaves, where the toolchain emits no guard either.
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
code, _, labels, _, _, err := assemble(t, nil)
code, _, labels, _, _, _, err := assemble(t, nil)
return code, labels, err
}
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
// rejects SB operands outright (single-function assembly cannot resolve
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
// file defines is recorded as an external relocation instead of failing
// the object-file emitters resolve it at link time.
// the object-file emitters resolve it at link time. goos selects the TLS
// access form: the empty default behaves as linux.
type linkInfo struct {
symbols map[string]bool
allowExternal bool
goos string
}
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
// keep the two-instruction form; shared linux does too, which gasm's raw
// path does not model and therefore does not select.
func (l *linkInfo) tlsOneInsn() bool {
switch l.goos {
case "", "linux", "freebsd":
return true
}
return false
}
// sbPatch is a function-relative static-symbol relocation: the disp32 field
@@ -66,9 +82,9 @@ type spadjStep struct {
// assemble encodes a TEXT body, returning the machine code, the static-symbol
// patch sites (for the file-level layout to resolve), the label table and the
// stack-adjustment boundaries.
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
if err := checkAdjspBalance(t); err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
fi := computeFrame(t)
chain := jumpChain(t)
@@ -85,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
// outgrows the short form.
long := make([]bool, len(t.Body))
sizes := make([]int, len(t.Body))
numTargets := make([]int, len(t.Body))
for i := range numTargets {
numTargets[i] = -1
}
offsets := map[string]int{}
pcs := make([]int, len(t.Body))
var guardJBlong, guardJBElong, moreJMPlong bool
poolSeen := map[string]bool{}
var poolList []floatPoolEntry
for {
guard := fi.guardLen(guardJBlong, guardJBElong)
pos := guard + len(fi.prologue)
for i := range numTargets {
numTargets[i] = -1
}
idxAtPc := map[int]int{}
for i, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
offsets[s.Name.Text] = pos
case *ast.Instr:
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
// The alignment pseudo-statement: its size is the
// padding to the next boundary at this very position,
// filled with NOPs at emission.
pad, err := pcAlignPad(pcAlignValue(s), pos)
if err != nil {
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
}
sizes[i] = pad
pcs[i] = pos
pos += pad
continue
}
sz, err := instrSize(s, fi, long[i], link)
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
sizes[i] = sz
pcs[i] = pos
idxAtPc[pos] = i
pos += sz
}
}
bodyLen := pos - (guard + len(fi.prologue))
// Expand any short jump whose displacement no longer fits rel8.
changed := false
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
// emission pass reads the same table after the loop converges. A
// target that is itself an unconditional local JMP is chased to the
// ultimate target: the toolchain's brloop pass collapses branch-to-
// branch chains before it encodes, so matching its bytes requires
// the same redirection.
for i := range numTargets {
numTargets[i] = -1
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if len(s.Operands) == 1 {
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
numTargets[i] = target
}
}
}
}
for i := range numTargets {
if numTargets[i] < 0 {
continue
}
tgt := numTargets[i]
for hop := 0; hop < len(t.Body); hop++ {
idx, ok := idxAtPc[tgt]
if !ok {
break
}
in, ok := t.Body[idx].(*ast.Instr)
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
break
}
if name, isLabel := labelName(in.Operands[0]); isLabel {
tgt = offsets[resolve(name)]
continue
}
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
next, okT := pcJumpTarget(t, idx, n, pcs)
if !okT {
break
}
tgt = next
continue
}
break // JMP through a register or memory: the chain ends
}
numTargets[i] = tgt
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if numTargets[i] >= 0 && !long[i] {
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
if !fits8(rel) {
long[i] = true
changed = true
}
}
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
@@ -229,12 +334,18 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{pos + epi, 0},
)
}
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
for _, entry := range pool {
if !poolSeen[entry.name] {
poolSeen[entry.name] = true
poolList = append(poolList, entry)
}
}
if len(code) != sizes[i] {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
}
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
for k := range ps {
@@ -272,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
pos += len(suffix)
}
_ = pos
return out, patches, offsets, steps, lines, nil
return out, patches, offsets, steps, lines, poolList, nil
}
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
@@ -419,6 +530,99 @@ func computeFrame(t *ast.Text) frameInfo {
return fi
}
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
// the second instruction boundary after the branch.
func pcJumpOffset(op *ast.Operand) (int, bool) {
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
return 0, false
}
return int(op.Addr.Offset), true
}
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
// instruction statements after the jump itself (N = 0 is the jump's own
// address, the classic park loop), and the target is the start of the Nth
// one. A negative N counts the same way backwards, before the jump: the
// exit loops write JMP -3(PC) to land three instructions earlier. Labels
// count not, in either direction. It reports false when the count runs
// past the end of the function, or before its first instruction.
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
if n == 0 {
return pcs[j], true
}
if n < 0 {
seen := 0
for k := j - 1; k >= 0; k-- {
if _, ok := t.Body[k].(*ast.Instr); !ok {
continue
}
seen--
if seen == n {
return pcs[k], true
}
}
return 0, false
}
seen := 0
for k := j + 1; k < len(t.Body); k++ {
if _, ok := t.Body[k].(*ast.Instr); !ok {
continue
}
seen++
if seen == n {
return pcs[k], true
}
}
return 0, false
}
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
// toolchain's asm6.go nop table); longer padding repeats the largest that
// fits, greedy from the end.
var x86Nops = [][]byte{
{0x90},
{0x66, 0x90},
{0x0F, 0x1F, 0x00},
{0x0F, 0x1F, 0x40, 0x00},
{0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
}
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
// the toolchain's fillnop.
func fillNOPs(p []byte) {
for len(p) > 0 {
m := min(len(p), len(x86Nops))
copy(p[:m], x86Nops[m-1])
p = p[m:]
}
}
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
// alignment must be a power of two in [8, 2048] and the padding runs to the
// next boundary (zero when the position is already aligned).
func pcAlignPad(align, pos int) (int, error) {
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
}
if lob := pos & (align - 1); lob != 0 {
return align - lob, nil
}
return 0, nil
}
// pcAlignValue reads a PCALIGN statement's alignment operand.
func pcAlignValue(s *ast.Instr) int {
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
return int(s.Operands[0].Imm.Val)
}
return 0 // rejected by pcAlignPad's range check
}
// hasCall reports whether the function body contains a CALL instruction.
func hasCall(t *ast.Text) bool {
for _, stmt := range t.Body {
@@ -593,7 +797,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
}
return jumpSize(mnem, long), nil
}
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
if err != nil {
return 0, err
}
@@ -628,9 +832,21 @@ func jumpSize(mnem string, long bool) int {
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
// jump label through the jump-to-jump chain before the offset lookup.
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnem := strings.ToUpper(s.Mnemonic.Text)
if mnem == "PCALIGN" {
// The layout pass already accounted the padding; emit the same
// amount of NOP bytes for the statement's own position.
pad, err := pcAlignPad(pcAlignValue(s), pc)
if err != nil {
return nil, nil, nil, err
}
out := make([]byte, pad)
fillNOPs(out)
return out, nil, nil, nil
}
var prefix []byte
if mnem == "RET" && fi.useFP {
prefix = fi.epilogue
@@ -638,6 +854,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
var code []byte
var ps []sbPatch
var pool []floatPoolEntry
var err error
if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
@@ -646,7 +863,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
// or the linker.
code, ps, err = encodeSBCall(s, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
for i := range ps {
ps[i].kind = RelCall
@@ -656,23 +873,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, nil, nil
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
return append(prefix, code...), nil, nil
return append(prefix, code...), nil, nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
} else {
code, ps, err = encodeNormal(s, fi, link)
code, ps, pool, err = encodeNormal(s, fi, link)
}
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
// Anchor the patch fields at function-relative positions: off indexes the
// disp32 field, after is the address just past the instruction.
@@ -681,49 +898,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, pool, nil
}
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
code, err := encodeBookkeeping(mnemUpper, s)
if err != nil {
return nil, nil, nil, err
}
return code, nil, nil, nil
}
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
// never as a 64-bit absolute immediate (verified against go tool asm).
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
s.Operands[0].Imm.Sym.Pseudo == "SB" {
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
e := &enc{}
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
}
return e.out, ps, nil, nil
}
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
// it out on the one-instruction TLS systems (linux and freebsd, not
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
// the macro GOROOT's go_tls.h defines, expands to exactly this
// statement, and the toolchain's pairing pass removes it whenever the
// following instruction's (TLS*1) index folds.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
return encodeTLSBaseLoad(s, fi, link)
}
_, size := splitSize(mnemUpper)
if size == 0 {
size = 8
}
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, size, fi, link)
o, err := operandFromAST(mnemUpper, op, size, fi, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
if p.tls {
ps[i].kind = RelTLSLE
}
}
return e.out, ps, nil
return e.out, ps, e.floatPoolList(), nil
}
// isBareTLS reports whether the operand is the bare TLS pseudo-register
// load source, the expansion of go_tls.h's get_tls(r) macro.
func isBareTLS(op *ast.Operand) bool {
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
op.Addr.Base == "" && op.Addr.Index == ""
}
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
// statement nops out: the following (TLS*1) access folds to a direct
// segment-absolute load. The two-instruction systems keep the segment load,
// nine bytes with the R_TLSLE patch site at the disp32.
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
if size == 0 {
size = 8
}
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
reg, ok := dst.(Reg)
if !ok || reg.isVec() {
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
}
if link == nil || link.tlsOneInsn() {
return nil, nil, nil, nil // noped out
}
seg := byte(0x64) // FS
if link.goos == "windows" {
seg = 0x65 // GS
}
e := &enc{}
i := &instr{
prefix: seg,
rexW: size == 8,
rexR: reg.idx >= 8,
opcode: []byte{0x8B},
modrm: 0x04 | (reg.idx&7)<<3,
sib: 0x25,
disp: le32(0),
tls: true,
}
if err := e.emit(i); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
}
return e.out, ps, nil, nil
}
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
// level, before operand conversion: the toolchain's shapes are FUNCDATA
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
// function body. The symbol reference must not run through the SB-operand
// path, which demands file-level resolution the statement never needs.
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
if len(s.Operands) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
}
a, b := s.Operands[0], s.Operands[1]
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil, nil
}
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
// target label, in the short (rel8) or long (rel32) form.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
// target label or from a numeric ±N(PC) instruction count, in the short
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
// numeric operand, negative when the operand is not one.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
if len(s.Operands) != 1 {
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
}
name, ok := labelName(s.Operands[0])
if !ok {
name, isLabel := labelName(s.Operands[0])
if !isLabel && numTarget < 0 {
return nil, fmt.Errorf("jump target must be a local label")
}
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
target, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
var target int
if isLabel {
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
t, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
}
target = t
} else {
target = numTarget
}
rel := int64(target - (pc + jumpSize(mnem, long)))
@@ -756,7 +1107,7 @@ func isSBCall(s *ast.Instr) bool {
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
if err != nil {
return nil, nil, err
}
@@ -797,6 +1148,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
return false
}
a := s.Operands[0].Addr
// ±N(PC) is the numeric relative form, the PC counts instructions from
// the branch: relative, not indirect.
if a.Base == "PC" || a.Index == "PC" {
return false
}
if a.Base != "" || a.Index != "" {
return true
}
@@ -813,7 +1169,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
@@ -830,8 +1186,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
var spReg = Reg{idx: 4, size: 8}
// operandFromAST converts a parsed operand into an encoder Operand, applying
// the frame translation to FP/SP pseudo-register operands.
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
// instruction's upper-case mnemonic, which the floating-point immediate gate
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
// accept one.
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
switch op.Kind {
case ast.OpImmediate:
if op.Imm.HasVal {
@@ -841,17 +1200,43 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Imm(v), nil
}
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
// $(-1.0) spelling (the constant-expression folder only folds
// integers, so that shape arrives with an empty Immediate and only
// the raw spelling carries the value). The toolchain rewrites it
// into a pooled-constant read on the SSE scalar paths and rejects
// it everywhere else.
if text, neg, ok := floatImmText(op); ok {
if !sseFloatImm[mnemUpper] {
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
}
return FloatImm{Text: text, Neg: neg}, nil
}
return nil, fmt.Errorf("non-integer immediate not supported")
case ast.OpAddr:
a := op.Addr
// A bracketed register range, [Z0-Z3]: the four-register source of
// the 4FMAPS/4VNNIW families. The EVEX quad-register emit path
// needs an encoder operand of its own, so the shape stays a named
// gap rather than an encoding.
// the 4FMAPS/4VNNIW families. The range must span four consecutive
// same-width vector registers, exactly what the toolchain's parser
// takes; the EVEX quad-register emit path reads the low end.
if a.Range != nil {
return nil, fmt.Errorf("register range %q needs quad-register encoder support", op.Raw)
lo, ok := ParseReg(a.Range.Lo)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
}
hi, ok := ParseReg(a.Range.Hi)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
}
if !lo.isVec() || lo.size != hi.size {
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
}
if hi.idx != lo.idx+3 {
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
}
return RegList{Lo: lo, Hi: hi}, nil
}
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
@@ -885,12 +1270,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
// Memory with a real base register: (base), off(base), (base)(index*scale).
if a.Base != "" {
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
// spellings. The segment override prefixes a disp32 absolute
// reference with no relocation.
if a.Base == "GS" || a.Base == "FS" {
seg := byte(0x64)
if a.Base == "GS" {
seg = 0x65
}
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
}
base, ok := ParseReg(a.Base)
if !ok {
return nil, fmt.Errorf("unknown base register %q", a.Base)
}
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
if a.Index != "" {
if a.Index == "TLS" {
// off(base)(TLS*1): the thread-local annotation. The
// one-instruction TLS form folds it to off(TLS), the
// segment-prefixed absolute whose disp32 carries an
// R_TLS_LE patch site; the base register disappears
// from the encoding, exactly as the toolchain's
// progedit rewrites the address.
seg := byte(0x64) // FS on linux, freebsd, plan9
if link != nil && link.goos == "windows" {
seg = 0x65 // GS
}
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
}
if a.Index == "GS" || a.Index == "FS" {
// 0(CX)(GS): the segment annotation rides the base
// access as the override prefix.
m.Seg = 0x64
if a.Index == "GS" {
m.Seg = 0x65
}
return m, nil
}
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
@@ -911,6 +1328,11 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
}
// A bare displacement with no base: the absolute address form,
// MOVL $0xf1, 0xf1. No segment and no relocation.
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
return SegAbs{Disp: a.Offset, Size: size}, nil
}
// Bare register.
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if r, ok := ParseReg(a.Sym.Name); ok {
@@ -921,3 +1343,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return nil, fmt.Errorf("unsupported operand")
}
// floatImmText recovers a floating-point immediate's magnitude and sign from
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
// folder cannot read it, and only the verbatim operand text still carries
// the value. Anything that is not a number a float parser accepts reports
// not-ok, so every other shape keeps its existing diagnostic.
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
if op.Imm.Float != "" {
return op.Imm.Float, op.Imm.Neg, true
}
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
return "", false, false
}
// joinRaw spaced the token texts; the compact spelling is what matters.
compact := strings.ReplaceAll(op.Raw, " ", "")
inner, ok := strings.CutPrefix(compact, "$(")
if !ok || !strings.HasSuffix(inner, ")") {
return "", false, false
}
inner = strings.TrimSuffix(inner, ")")
inner = strings.TrimPrefix(inner, "+")
if s, ok := strings.CutPrefix(inner, "-"); ok {
neg = true
inner = s
}
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
return "", false, false
}
if _, err := strconv.ParseFloat(inner, 64); err != nil {
return "", false, false
}
return inner, neg, true
}
+79 -2
View File
@@ -10,8 +10,8 @@ import (
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// firstText parses src and returns its first TEXT function.
@@ -370,6 +370,58 @@ end:
}
}
// TestAssembleNumericPCJumps pins the numeric ±N(PC) branch operands: N
// counts instruction statements, skipping labels, in both directions (the
// runtime's exit loops write JMP -3(PC)), N = 0 parks on the jump itself.
func TestAssembleNumericPCJumps(t *testing.T) {
fn := firstText(t, `
#include "textflag.h"
TEXT ·exit(SB), NOSPLIT, $0
MOVB $1, AL
lab:
MOVB $2, AL
MOVB $3, AL
JMP -3(PC)
MOVB $4, AL
park:
JMP 0(PC)
MOVB $5, AL
JMP 2(PC)
MOVB $6, AL
RET
`)
code, _, err := Assemble(fn)
if err != nil {
t.Fatalf("Assemble: %v", err)
}
// From the Go-assembled function:
// MOVB $1, AL b001
// MOVB $2, AL b002
// MOVB $3, AL b003
// JMP -3(PC) ebf8 (three instructions back, past lab:)
// MOVB $4, AL b004
// JMP 0(PC) ebfe (the park loop)
// MOVB $5, AL b005
// JMP 2(PC) eb02 (over MOVB $6 to the RET)
// MOVB $6, AL b006
// RET c3
want := []byte{
0xb0, 0x01,
0xb0, 0x02,
0xb0, 0x03,
0xeb, 0xf8,
0xb0, 0x04,
0xeb, 0xfe,
0xb0, 0x05,
0xeb, 0x02,
0xb0, 0x06,
0xc3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("numeric-PC mismatch:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
func TestAssemblePrefetch(t *testing.T) {
fn := firstText(t, `
#include "textflag.h"
@@ -568,3 +620,28 @@ TEXT ·framed(SB), $16-8
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
// TestAssembleRegRange pins the bracketed register range at the statement
// level: exactly four consecutive same-width vector registers assemble, the
// toolchain's rejected shapes all report an error.
func TestAssembleRegRange(t *testing.T) {
asm := func(t *testing.T, op string) ([]byte, error) {
t.Helper()
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse %s: %v", op, errs)
}
code, _, err := Assemble(f.Decls[0].(*ast.Text))
return code, err
}
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
if _, err := asm(t, op); err != nil {
t.Errorf("%s: %v", op, err)
}
}
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
if _, err := asm(t, op); err == nil {
t.Errorf("%s: assembled, want an error", op)
}
}
}
+1 -1
View File
@@ -8,7 +8,7 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// ulebIter reads ULEB128 values, the .debug_abbrev and line-header
+2 -2
View File
@@ -12,8 +12,8 @@ import (
"path/filepath"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// The object-file tests share one source: two exported functions, one
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestELFAARCH64Object checks the structure of the emitted AArch64 ELF64
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestELFLOONG64Object checks the structure of the emitted LoongArch ELF64
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
+3 -3
View File
@@ -20,9 +20,9 @@ func Encodable(mnemonic string) bool {
switch upper {
case "RET", "NOP", "CALL", "JMP",
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
// The literal-data pseudo-ops, the accepted-and-ignored END and the
// SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP":
// The literal-data pseudo-ops, the accepted-and-ignored END and
// bookkeeping statements, and the SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
return true
}
if _, ok := noOperandTable[upper]; ok {
+228 -2
View File
@@ -5,6 +5,8 @@ package asm
import (
"fmt"
"math"
"strconv"
"strings"
)
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
type enc struct {
out []byte
patches []encPatch // disp32 fields awaiting static-symbol resolution
// FloatPool collects the pooled constants the floating-point
// immediates reference, in first-use order.
floatPool []floatPoolEntry
floatPoolSeen map[string]bool
}
// floatPoolEntry is one pooled floating-point constant: the symbol name
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
type floatPoolEntry struct {
name string
data []byte
}
// addFloatPool records a pooled constant, deduplicated by symbol name.
func (e *enc) addFloatPool(name string, bits uint64, width int) {
if e.floatPoolSeen == nil {
e.floatPoolSeen = map[string]bool{}
}
if e.floatPoolSeen[name] {
return
}
e.floatPoolSeen[name] = true
data := make([]byte, width)
for i := range width {
data[i] = byte(bits >> (8 * i))
}
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
}
// floatPoolList returns the pooled constants in first-use order.
func (e *enc) floatPoolList() []floatPoolEntry {
return e.floatPool
}
// encPatch marks a 4-byte displacement field in enc.out that must receive the
@@ -29,6 +64,7 @@ type encPatch struct {
off int
name string
addend int64
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
}
func (e *enc) encode(mnem string, ops []Operand) error {
@@ -102,6 +138,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeEnd(ops)
case "ADJSP":
return e.encodeAdjsp(ops)
// The runtime's bookkeeping statements carry no text bytes: go tool asm
// records FUNCDATA and PCDATA in the program list only, so the encoded
// body shows nothing, on every architecture.
case "FUNCDATA", "PCDATA":
return e.encodeFuncdata(upper, ops)
}
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
@@ -141,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
}
// Legacy SSE packed binaries dispatch on the full name: the packed
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
// which the size split must not eat.
// which the size split must not eat. A floating-point immediate
// rewrites into a pooled-constant read on the scalar members.
if m, ok := sseBinTable[upper]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
if m, ok := sseBinTable[base]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
// The imm8-controlled legacy instructions, the lane extracts and inserts
@@ -222,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeCvtInt(base, ops, size)
case "FMOVD":
return e.encodeFmov(ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
case "MOVSD", "MOVSS":
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatMove(upper, f, ops)
}
return e.encodeSSEMove(sseMoveTable[base], ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
return e.encodeSSEMove(sseMoveTable[base], ops)
}
return fmt.Errorf("unsupported instruction %q", mnem)
@@ -285,6 +338,33 @@ func (e *enc) encodeData(mnem string, ops []Operand) error {
return nil
}
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
// for either (the entries live in the object's ancillary tables, not the
// function body), and the operand shapes it takes are exactly these: an
// integer count first, then a symbol reference for FUNCDATA and an integer
// value for PCDATA. The other architectures accept-and-ignore the same
// statements; amd64 now matches.
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
}
if _, ok := ops[0].(Imm); !ok {
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if _, ok := ops[1].(sbMem); !ok {
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if _, ok := ops[1].(Imm); !ok {
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil
}
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
// entirely: the AEND Prog is skipped when the program list is flushed, so
// the statements after an END still belong to the same function and the
@@ -319,6 +399,120 @@ func (e *enc) encodeAdjsp(ops []Operand) error {
return nil
}
// --- floating-point immediates ----------------------------------------------
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
// immediate, the set go tool asm rewrites into a pooled-constant read: the
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
// reject the immediate in the toolchain and are absent here on purpose.
var sseFloatImm = map[string]bool{
"MOVSD": true, "MOVSS": true,
"ADDSD": true, "ADDSS": true,
"SUBSD": true, "SUBSS": true,
"MULSD": true, "MULSS": true,
"DIVSD": true, "DIVSS": true,
"COMISD": true, "COMISS": true,
"UCOMISD": true, "UCOMISS": true,
}
// floatImmOperand reports whether the operand list opens with a
// floating-point immediate in the two-operand spelling (imm, dst).
func floatImmOperand(ops []Operand) (FloatImm, bool) {
if len(ops) != 2 {
return FloatImm{}, false
}
f, ok := ops[0].(FloatImm)
return f, ok
}
// floatPoolValue evaluates a floating-point immediate at the width its
// mnemonic encodes and names the pool constant the toolchain synthesises:
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
// rounding of the parsed value). The name carries the IEEE-754 bits; the
// section holds them little-endian.
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
v, err := strconv.ParseFloat(f.Text, 64)
if err != nil {
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
}
if f.Neg {
v = -v
}
if strings.HasSuffix(mnem, "D") {
bits = math.Float64bits(v)
return bits, fmt.Sprintf("$f64.%016x", bits), nil
}
bits = uint64(math.Float32bits(float32(v)))
return bits, fmt.Sprintf("$f32.%08x", bits), nil
}
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
// source. A positive zero needs no memory read: the toolchain emits
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
// file-level layout or the linker resolves.
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
if bits == 0 {
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
if err := setRM(i, dst, dst, 8); err != nil {
return err
}
return e.emit(i)
}
m := sseMoveTable[mnem]
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
// a floating-point immediate source: the constant is read from the pool into
// the instruction's r/m side (reg = destination), the rewrite go tool asm
// performs at the source level.
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
// spellings end in D, the single spellings in S.
func mwidth(mnem string) int {
if strings.HasSuffix(mnem, "D") {
return 8
}
return 4
}
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
func splitSize(upper string) (base string, size int) {
if upper == "" {
@@ -385,6 +579,7 @@ type instr struct {
disp []byte
imm []byte
sb *sbRef // static-symbol displacement in disp, awaiting resolution
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
}
// sbRef records that an instruction's displacement refers to a static symbol
@@ -427,6 +622,9 @@ func (e *enc) emit(i *instr) error {
if i.sb != nil {
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
}
if i.tls {
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
}
e.out = append(e.out, i.disp...)
e.out = append(e.out, i.imm...)
return nil
@@ -481,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
i.disp = le32(0)
i.sb = &sbRef{name: r.name, addend: r.addend}
return nil
case TLSMem:
// off(TLS): the segment-prefixed absolute access, mod=00 with the
// SIB escape's disp32 absolute form. The displacement is the TLS
// slot offset, patched by the linker's TLS relocation.
i.prefix = r.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(r.Disp)
i.tls = true
return nil
case SegAbs:
// 0x30(GS): the segment override with the SIB escape's disp32
// absolute form, no relocation.
setSegAbs(i, regField, r)
return nil
default:
return fmt.Errorf("invalid r/m operand %T", rm)
}
}
func setMem(i *instr, regField int, m Mem) error {
if m.Seg != 0 {
i.prefix = m.Seg
}
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
if err != nil {
return err
@@ -499,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
return nil
}
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
// override with the mod=00 SIB escape's disp32 absolute form and no
// relocation.
func setSegAbs(i *instr, regField int, m SegAbs) {
i.prefix = m.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(m.Disp)
}
// memComponents computes the ModR/M byte (with the given reg field), the SIB
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
+157
View File
@@ -9,6 +9,9 @@ import (
"testing"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// decode encodes an instruction and decodes it back, returning the decoded
@@ -1064,3 +1067,157 @@ func TestAdjsp(t *testing.T) {
t.Error("ADJSP AX assembled, want an error")
}
}
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
// byte for byte against go tool asm: the scalar moves and the scalar
// arithmetic read the constant from a synthesised read-only pool symbol
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
func TestFloatImmediateGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
// single spellings; negative zero keeps its sign bit and never takes the
// XORPS shortcut.
for _, c := range []struct {
mnem string
imm FloatImm
want string
}{
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
} {
_, name, err := floatPoolValue(c.mnem, c.imm)
if err != nil {
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
continue
}
if name != c.want {
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
}
}
// The shapes the toolchain's parser rejects: the packed and uniform
// forms, a non-vector destination, and the integer spellings.
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
// go tool asm emits no text bytes for either, on every architecture.
func TestBookkeepingGroundTruth(t *testing.T) {
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
} {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if len(code) != 0 {
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
}
}
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
// At the statement level the bookkeeping lines sit between real
// instructions and contribute nothing to the body, symbol reference
// included: the FUNCDATA operand never needs file-level resolution.
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
want := "90c3"
if got := hexCompact(img.Code); got != want {
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("FUNCDATA $1, X0 assembled, want an error")
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("PCDATA $1, X0 assembled, want an error")
}
// Encodable mirrors Encode for the names this work touched.
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
if !Encodable(mnem) {
t.Errorf("Encodable(%s) = false, want true", mnem)
}
}
}
// mustParse parses src or fails the test.
func mustParse(t *testing.T, src string) *ast.File {
t.Helper()
f, errs := parser.Parse("t_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
return f
}
+97 -2
View File
@@ -737,6 +737,91 @@ var evexTable = map[string]evexSpec{
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
}
// evexQuad describes one quad-register instruction: the opcode under
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
// registers the bracketed list and the destination take (512-bit ZMM for
// the packed forms, 128-bit XMM for the scalar ones).
type evexQuad struct {
opcode byte
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
}
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
// families) to their encoding. The operand shape is fixed: a single memory
// source in r/m, the bracketed register list whose LOW register travels the
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
// destination in reg. The vector length follows the destination (512-bit
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
// multiplier stays 16 for every member, the toolchain's own tuple choice.
var evexQuadTable = map[string]evexQuad{
"V4FMADDPS": {0x9A, 64},
"V4FMADDSS": {0x9B, 16},
"V4FNMADDPS": {0xAA, 64},
"V4FNMADDSS": {0xAB, 16},
"VP4DPWSSD": {0x52, 64},
"VP4DPWSSDS": {0x53, 64},
}
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
func isEvexQuad(upper string) bool {
_, ok := evexQuadTable[upper]
return ok
}
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
// The register list is the VVVV-side source: its low register fills the
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
// suffix applies.
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
if len(ops) != 3 && len(ops) != 4 {
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
}
mem, lst := ops[0], ops[1]
dst := ops[len(ops)-1]
mask := 0
if len(ops) == 4 {
k, ok := ops[2].(Reg)
if !ok || !k.mask {
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
}
if k.idx == 0 {
return fmt.Errorf("k0 is not a usable mask register")
}
mask = k.idx
}
list, ok := lst.(RegList)
if !ok {
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
}
if list.Lo.size != q.width {
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if dstReg.size != q.width {
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
}
if !memOperand(mem) {
return fmt.Errorf("%s: the source must be a memory operand", mnem)
}
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
// its fifth bit into the same field the list's low register occupies.
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
}
if sfx.zeroing && mask == 0 {
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
}
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
// The vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones), exactly as the oracle encodes it.
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
}
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
// depends on the source kind, a GPR source uses opReg, a memory source uses
// opMem with a disp8×N of n.
@@ -812,8 +897,10 @@ func isEvex(mnemUpper string) bool {
if _, ok := evexBcastTable[mnemUpper]; ok {
return true
}
_, ok := evexMoveTable[mnemUpper]
return ok
if _, ok := evexMoveTable[mnemUpper]; ok {
return true
}
return isEvexQuad(mnemUpper)
}
// evexRequired reports whether the operands force the EVEX encoding of a
@@ -995,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexRM(spec, ops, 0, sfx)
}
spec, inTable := evexTable[mnemUpper]
if q, ok := evexQuadTable[mnemUpper]; ok {
// The quad-register family carries no rounding, SAE or broadcast;
// only masking and zeroing apply.
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
}
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
}
if inTable {
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
+124
View File
@@ -810,3 +810,127 @@ func TestAvx512CorpusFamilies(t *testing.T) {
}
}
}
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
// source keeps r/m, the bracketed list's LOW register travels the inverted
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
// the vector length follows the destination (L'L=512 for the ZMM forms,
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
// member. The x86 decoder has no view of these forms, so no decode check
// runs.
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
sp := vreg(t, "RSP")
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a842411000000"},
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f22f4a9a842411000000"},
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f25f429a842411000000"},
{"V4FMADDPS Z8 dst", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
"62727f4a9a842411000000"},
{"V4FMADDPS disp8x16", "V4FMADDPS",
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a442404"},
{"V4FMADDPS unmasked", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f489a842411000000"},
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9bb007000000"},
{"V4FMADDSS (DI)", "V4FMADDSS",
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9b37"},
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
"62e22f0d9bb007000000"},
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
"62e25f059bb007000000"},
{"V4FMADDSS X30 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
"62627f0d9bb007000000"},
{"V4FMADDSS X3 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
"62f27f0d9b9807000000"},
{"V4FMADDSS disp8x16", "V4FMADDSS",
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
"62625f059b7001"},
{"V4FNMADDPS", "V4FNMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4aaa842411000000"},
{"V4FNMADDSS", "V4FNMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0dabb007000000"},
{"VP4DPWSSD", "VP4DPWSSD",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a52842411000000"},
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f4853842411000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
// the register class the list and the destination take is fixed per
// instruction, the source is memory only, the opmask slot is positional and
// the list's low register owns V'VVVV.
func TestEvexQuadRegisterErrors(t *testing.T) {
sp := vreg(t, "RSP")
list := func(lo, hi string) RegList {
return RegList{vreg(t, lo), vreg(t, hi)}
}
cases := []struct {
name string
mnem string
ops []Operand
}{
{"X list on the PS form", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"Z list on the SS form", "V4FMADDSS",
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
{"Y destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
{"register source", "V4FMADDPS",
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"non-mask third operand", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
{"k0 mask", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
{"K after the destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
{"zeroing without a mask", "V4FMADDPS.Z",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
{"SAE suffix", "V4FMADDPS.SAE",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"high index source", "VP4DPWSSD",
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"short operand list", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
}
for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
+1 -1
View File
@@ -12,7 +12,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// goobjView is a minimal parsed view of a GOOBJ payload, enough to check
+6 -6
View File
@@ -9,8 +9,8 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// The expected bytes are pinned from `go tool asm` output (Go 1.27, amd64,
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
func TestStackGuardGOObjInternalCall(t *testing.T) {
for _, tt := range []struct {
src string
assemble func(*ast.File) (*Image, error)
assemble func(*ast.File, ...AssembleOption) (*Image, error)
}{
{"g_amd64.s", AssembleFile},
{"g_arm64.s", AssembleFileARM64},
{"g_riscv64.s", AssembleFileRISCV},
{"g_loong64.s", AssembleFileLOONG64},
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
} {
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
+74 -21
View File
@@ -192,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
}
return e.emit(i)
case TLSMem:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
// offset with its R_TLSLE patch site.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case SegAbs:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, 0x30(GS): the segment-absolute load.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case Imm:
if dstIsReg {
v := int64(src)
@@ -232,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
i.imm = imm
return e.emit(i)
}
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
// destination slot is the absolute-address crash-store spelling,
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
// as an immediate, and the store's disp32 carries the address.
op := byte(0xC7)
if size == 1 {
op = 0xC6
}
if d, ok := dst.(Imm); ok {
i := newInstr(size, []byte{op})
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
immBytes, err := immediate(int64(src), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
i := newInstr(size, []byte{op})
if err := setRMDigit(i, 0, dst, size); err != nil {
return err
@@ -627,7 +664,16 @@ func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
func (e *enc) encodeImul(ops []Operand, size int) error {
switch len(ops) {
case 2:
// IMUL r, r/m: 0x0F 0xAF.
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
// Otherwise IMUL r, r/m: 0x0F 0xAF.
if imm, ok := ops[0].(Imm); ok {
dstReg, isReg := ops[1].(Reg)
if !isReg {
return fmt.Errorf("IMUL: destination must be a register")
}
return e.encodeImulImm(imm, dstReg, dstReg, size)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("IMUL: destination must be a register")
@@ -647,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
if !ok {
return fmt.Errorf("IMUL: immediate operand expected first")
}
// Plan 9 order: IMUL $imm, src, dst.
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
// r/m operand (setRM takes registers and memory alike).
return e.encodeImulImm(imm, ops[1], dstReg, size)
}
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
}
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
// when the value fits, 0x69 with a 32-bit immediate otherwise.
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
// --- PUSH / POP -------------------------------------------------------------
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
+3 -3
View File
@@ -16,7 +16,7 @@ import (
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel;
@@ -25,7 +25,7 @@ import (
func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx2_amd64.s"
if _, err := os.Stat(path); err != nil {
t.Skip("go-libraries repository not present next to gasm-devkit")
t.Skip("go-libraries repository not present next to gasm-sdk")
}
src, err := os.ReadFile(path)
if err != nil {
@@ -86,7 +86,7 @@ func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
func TestAssembleGoFlacAVX512Kernel(t *testing.T) {
path := "../../go-libraries/go-flac/avx512_amd64.s"
if _, err := os.Stat(path); err != nil {
t.Skip("go-libraries repository not present next to gasm-devkit")
t.Skip("go-libraries repository not present next to gasm-sdk")
}
src, err := os.ReadFile(path)
if err != nil {
+10 -8
View File
@@ -13,7 +13,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// The differential kernels for the DATA-path and front-end gaps are kept in
@@ -63,7 +63,10 @@ func toolAsmObject(t *testing.T, path, goarch string) []byte {
}
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
// toolchain object, keyed by the name the object records (pkg.name).
// toolchain object, keyed by the name the object records (pkg.name). Each
// function's span is its own symbol size: a toolchain object that follows
// the text with data symbols (the synthesised float-constant pool) would
// otherwise fold them into the last function's bytes.
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
t.Helper()
v := openGoobj(t, obj)
@@ -76,18 +79,13 @@ func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
preceding += len(v.blk(bi)) / symSize
}
total := preceding + len(nps)
out := make(map[string][]byte, len(nps))
for i, s := range nps {
if s.typ != kindSTEXT {
continue
}
start := le.Uint32(didx[4*(preceding+i):])
end := uint32(len(data))
if preceding+i+1 < total {
end = le.Uint32(didx[4*(preceding+i+1):])
}
out[s.name] = data[start:end]
out[s.name] = data[start : start+s.size]
}
return out
}
@@ -129,6 +127,10 @@ func TestDifferentialKernels(t *testing.T) {
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
} {
+1 -1
View File
@@ -12,7 +12,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestGOObjectLOONG64Structure checks the emitted loong64 object's blocks:
+109 -9
View File
@@ -5,10 +5,11 @@ package asm
import (
"fmt"
"math"
"sort"
"strconv"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// Image is an assembled file: the function bodies laid out in source order,
@@ -149,6 +150,18 @@ func (img *Image) Bytes() []byte {
return append(out, img.Data...)
}
// AssembleOption adjusts the file-level assembly context.
type AssembleOption func(*linkInfo)
// WithGOOS selects the target operating system for the forms that depend on
// it, the TLS access shape above all: linux and freebsd take the
// one-instruction form, windows and plan9 keep the two-instruction load.
func WithGOOS(goos string) AssembleOption {
return func(l *linkInfo) {
l.goos = goos
}
}
// AssembleFile assembles every TEXT function of a parsed file and lays out
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
// reference to a file-local static symbol becomes a RIP-relative load whose
@@ -156,7 +169,7 @@ func (img *Image) Bytes() []byte {
// GLOBL defines is recorded as an external relocation (Externals) with its
// displacement left zero, the object-file emitters resolve it at link
// time, while the raw image (Bytes) cannot represent it.
func AssembleFile(f *ast.File) (*Image, error) {
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
dataSyms, err := collectData(f)
if err != nil {
return nil, err
@@ -165,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
for _, d := range dataSyms {
known[d.name] = true
}
// TEXT symbols are file-level definitions too: a symbol immediate
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
for _, d := range f.Decls {
if t, ok := d.(*ast.Text); ok {
known[t.Name.Name] = true
}
}
link := &linkInfo{symbols: known, allowExternal: true}
for _, o := range opts {
o(link)
}
poolSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
textOff := map[string]int{}
@@ -179,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, patches, labels, steps, lines, err := assemble(t, link)
code, patches, labels, steps, lines, pool, err := assemble(t, link)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
// The pooled floating-point constants join the declared data as
// read-only symbols, deduplicated across the file (the toolchain
// synthesises the same symbols into its rodata).
for _, entry := range pool {
if poolSeen[entry.name] {
continue
}
poolSeen[entry.name] = true
dataSyms = append(dataSyms, dataSym{
name: entry.name,
buf: entry.data,
size: len(entry.data),
rodata: true,
dupok: true,
})
}
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
@@ -299,6 +342,10 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if err != nil {
return nil, err
}
// The pooled $i64 constants the wide MOV immediate loads refer to join
// the declared data as read-only symbols, deduplicated across the file
// (the toolchain synthesises the same symbols into its rodata).
litSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls {
@@ -306,10 +353,23 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, labels, relocs, lines, spadj, err := assembleRISCV(t)
code, labels, relocs, lines, spadj, lits, err := assembleRISCV(t)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
for _, lit := range lits {
if litSeen[lit.Name] {
continue
}
litSeen[lit.Name] = true
dataSyms = append(dataSyms, dataSym{
name: lit.Name,
buf: lit.Data,
size: len(lit.Data),
rodata: true,
dupok: true,
})
}
fl := FuncLayout{
Name: t.Name.Name,
Pkg: t.Name.Pkg,
@@ -561,11 +621,6 @@ func collectData(f *ast.File) ([]dataSym, error) {
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
}
w := dd.Width
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
off := dd.Name.Offset
buf := syms[i].buf
if off < 0 || off+int64(w) > int64(len(buf)) {
@@ -587,9 +642,54 @@ func collectData(f *ast.File) ([]dataSym, error) {
})
continue
}
// A string or rune value ("DATA s+0(SB)/20, $"text"") writes its
// bytes into the field and leaves the rest zero, the toolchain's
// WriteString: the declared width must hold every byte, and any
// width is legal.
if s := dd.Value.Imm.Str; s != "" && !dd.Value.Imm.HasVal {
text, err := strconv.Unquote(s)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid string value %s", dd.Name.Name, s)
}
if len(text) > w {
return nil, fmt.Errorf("DATA %q: string of %d bytes does not fit width %d", dd.Name.Name, len(text), w)
}
copy(buf[off:], text)
continue
}
// A floating-point value stores its IEEE-754 bits: /4 the float32
// rounding of the parsed double, /8 the full 64 bits, the
// toolchain's WriteFloat32 and WriteFloat64.
if f := dd.Value.Imm.Float; f != "" && !dd.Value.Imm.HasVal {
num, err := strconv.ParseFloat(f, 64)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid floating-point value %q", dd.Name.Name, f)
}
if dd.Value.Imm.Neg {
num = -num
}
var v uint64
switch w {
case 4:
v = uint64(math.Float32bits(float32(num)))
case 8:
v = math.Float64bits(num)
default:
return nil, fmt.Errorf("DATA %q: invalid width %d for a float (want 4 or 8)", dd.Name.Name, w)
}
for j := range w {
buf[off+int64(j)] = byte(v >> (8 * j))
}
continue
}
if !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
}
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
v := dd.Value.Imm.Val
if dd.Value.Imm.Neg {
v = -v
+74 -1
View File
@@ -5,13 +5,14 @@ package asm
import (
"encoding/binary"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestAssembleFileStaticData checks the whole-image layout; code, padding
@@ -418,3 +419,75 @@ func main() {
t.Fatalf("linked program failed: %v\n%s", err, out)
}
}
// TestCollectDataFloatAndStringValues covers the non-integer DATA values the
// runtime's math and asm files use: floating-point initialisers store their
// IEEE-754 bits (/4 the float32 rounding, /8 the full double) and string
// initialisers write their bytes zero-padded within the declared width.
func TestCollectDataFloatAndStringValues(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL vals<>(SB), RODATA, $44
DATA vals<>+0(SB)/8, $0.5
DATA vals<>+8(SB)/8, $-1.0
DATA vals<>+16(SB)/4, $1.5
DATA vals<>+20(SB)/16, $"call frame too "
DATA vals<>+36(SB)/4, $"hi"
`
f, errs := parser.Parse("fvals_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
d := byName["vals"]
if d.Size != 44 {
t.Fatalf("vals size = %d, want 44", d.Size)
}
buf := img.Data[d.Offset : d.Offset+44]
// 0.5 = 0x3FE0000000000000, -1.0 = 0xBFF0000000000000 (float64);
// 1.5 = 0x3FC00000 (float32).
for _, c := range []struct {
off int
want []byte
}{
{0, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xE0, 0x3F}},
{8, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xF0, 0xBF}},
{16, []byte{0x00, 0x00, 0xC0, 0x3F}},
{20, []byte("call frame too ")},
{36, []byte{'h', 'i', 0x00, 0x00}},
} {
if string(buf[c.off:c.off+len(c.want)]) != string(c.want) {
t.Errorf("vals+%d: got % x, want % x", c.off, buf[c.off:c.off+len(c.want)], c.want)
}
}
}
// TestCollectDataValueErrors pins the value-kind width rules: a float needs
// width 4 or 8, a string must fit its declared width, and a bad float
// literal is diagnosed rather than stored.
func TestCollectDataValueErrors(t *testing.T) {
cases := []string{
`GLOBL v<>(SB), RODATA, $4
DATA v<>+0(SB)/1, $0.5`,
`GLOBL v<>(SB), RODATA, $2
DATA v<>+0(SB)/2, $"toolarge"`,
}
for i, src := range cases {
full := "#include \"textflag.h\"\nTEXT ·Keep(SB), NOSPLIT, $0-8\n\tRET\n" + src
f, errs := parser.Parse(fmt.Sprintf("verr%d_amd64.s", i), full)
if len(errs) > 0 {
t.Fatalf("case %d parse: %v", i, errs)
}
if _, err := AssembleFile(f); err == nil {
t.Errorf("case %d: expected an error, got none", i)
}
}
}
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// assembleLOONG64 assembles a LoongArch (loong64) TEXT function body into
+2 -2
View File
@@ -8,8 +8,8 @@ import (
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// firstTextLOONG64 parses assembly source and returns the first TEXT body.
+1 -1
View File
@@ -6,7 +6,7 @@ package asm
import (
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// Loong64 frame mapping, matching the Go toolchain's loong64 backend.
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"bytes"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestLOONG64_sys exercises the no-operand system instructions and the
+1 -1
View File
@@ -6,7 +6,7 @@ package asm
import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestLOONG64RelocOffsetsIncludePrologue pins the function-relative
+46
View File
@@ -14,6 +14,51 @@ type Imm int64
func (Imm) isOperand() {}
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
// low register through the inverted 5-bit V'VVVV field; the three higher
// registers are implied by the instruction, so only the pair travels here.
type RegList struct {
Lo Reg
Hi Reg // implied by the encoding; Lo.idx+3 by construction
}
func (RegList) isOperand() {}
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
// encoding takes an XMM/memory source at that position rewrite it as a read
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
// own behaviour; every other instruction rejects it.
type FloatImm struct {
Text string // the numeric text as written, sign excluded
Neg bool // a leading minus
}
func (FloatImm) isOperand() {}
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
// (the linker fills the TLS slot offset).
type TLSMem struct {
Disp int64
Size int
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
}
func (TLSMem) isOperand() {}
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
// prefixes a disp32 absolute reference with no relocation. The base
// register spellings GS and FS produce it.
type SegAbs struct {
Disp int64
Size int
Seg byte // 0x64 FS, 0x65 GS
}
func (SegAbs) isOperand() {}
// Mem is a memory operand of the form disp(base)(index*scale).
type Mem struct {
Base Reg
@@ -23,6 +68,7 @@ type Mem struct {
Size int // operand width in bytes
HasBase bool
HasIndex bool
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
}
func (Mem) isOperand() {}
+287 -36
View File
@@ -6,21 +6,23 @@ package asm
import (
"errors"
"fmt"
"math/bits"
"slices"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// assembleRISCV assembles a RISC-V TEXT function body into machine code.
// It handles the full RV64IMAFDC instruction set including RVC compression.
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, []RiscvLiteral, error) {
fi := riscvComputeFrame(t)
prologue := riscvPrologue(fi)
guardLen, err := riscvGuardLen(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
lits := &riscvLiterals{}
var relocs []Reloc
var spadj []SpadjStep
@@ -74,9 +76,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
pc := len(prologue)
for i := range recs {
branchLike := isBranchLike(recs[i].instr.Mnemonic.Text) || riscvIsCondBranch(recs[i].instr.Mnemonic.Text)
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil) // no relocs in Pass 2
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil, lits) // no relocs in Pass 2
if err != nil && !(branchLike && riscvIsRangeError(err)) {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
}
if err != nil {
code = make([]byte, 4)
@@ -178,13 +180,20 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
}
}
if !changed {
// Capture the final pcs for the N(PC) branch forms: their target
// is the instruction N source slots away, resolved by index.
// Capture the final pcs for the N(PC) branch and jump forms: the
// target is the instruction N source slots away (N=0 the branch
// itself, N negative backwards), resolved by index against the
// final layout.
pcRelPcs = map[*ast.Instr]int{}
for i := range recs {
if _, ok := riscvPCRelOffset(recs[i].instr); ok {
pcRelPcs[recs[i].instr] = pcs[i]
n, ok := riscvPCRelOffset(recs[i].instr)
if !ok {
continue
}
if i+n < 0 || i+n >= len(recs) {
continue
}
pcRelPcs[recs[i].instr] = pcs[i+n]
}
break
}
@@ -198,7 +207,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
var out []byte
guardBytes, guardReloc, err := riscvGuard(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
if fi.needSplit {
out = append(out, guardBytes...)
@@ -220,11 +229,11 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
// The JMP a relaxation inserted: JAL X0 to the original target.
targetOff, ok := offsets[r.jmpTo]
if !ok {
return nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
}
offset := int32(targetOff - pc)
if err := riscvCheckJumpOffset(r.jmpTo, offset); err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
word := riscvJType(0, offset)
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
@@ -233,7 +242,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
// JMP, always the very next instruction (offset 4).
enc, rs1, rs2, ok := riscvInvertedBranchEnc(strings.ToUpper(r.instr.Mnemonic.Text), r.instr.Operands)
if !ok {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
}
word := riscvBType(enc, rs1, rs2, 4)
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
@@ -241,9 +250,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
code = r.code
default:
var err error
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs)
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs, lits)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
if c16, ok := tryCompressRVC(r.instr, fi); ok {
code = []byte{byte(c16), byte(c16 >> 8)}
@@ -270,7 +279,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
if fi.needSplit {
relocs = append(relocs, guardReloc)
}
return out, offsets, relocs, lines, spadj, nil
return out, offsets, relocs, lines, spadj, lits.list(), nil
}
// riscvImmAlias maps the R-type ALU mnemonics onto their I-type immediate
@@ -348,6 +357,11 @@ func riscvPadBytes(pad int) []byte {
func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
mnem := instr.Mnemonic.Text
ops := instr.Operands
mnem = riscvNormalisePseudo(mnem)
if mnem == "FUNCDATA" || mnem == "PCDATA" {
// The bookkeeping statements contribute no bytes.
return 0
}
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
if mnem == "RET" {
@@ -368,7 +382,25 @@ func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
}
// MOV $imm, rd → size depends on the immediate and RVC compression.
if isImmOperand(ops[0]) && ops[0].Imm.Sym == nil {
return riscvMovImmSize(regFromOperand(ops[1]), immFromOperand(ops[0]))
imm := riscvOperandImm64(ops[0])
if int64(int32(imm)) != imm {
return riscvMovImm64Size(regFromOperand(ops[1]), imm)
}
return riscvMovImmSize(regFromOperand(ops[1]), int32(imm))
}
// MOV $sym+off(FP|SP), rd → the frame-adjusted offset as an ADDI,
// compressed like riscvSPAddiBytes encodes it.
if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil &&
(ops[0].Imm.Sym.Pseudo == "FP" || ops[0].Imm.Sym.Pseudo == "SP") {
rd := regFromOperand(ops[1])
_, off := riscvResolvePseudo(ops[0].Imm.Sym, fi)
if rd > 0 && off == 0 {
return 2 // C.MV rd, SP
}
if isRVCIntReg(rd) && off > 0 && off < 1024 && off%4 == 0 {
return 2 // C.ADDI4SPN
}
return riscvItypeImmediateSize("ADDI", off)
}
// Frame-relative loads and stores: a frame offset beyond the signed
// 12-bit range materialises the address in X31 first.
@@ -664,9 +696,10 @@ func riscvCheckJumpOffset(target string, off int32) error {
}
// encodeRISCVInstr encodes a single RISC-V instruction.
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int, lits *riscvLiterals) ([]byte, error) {
mnem := instr.Mnemonic.Text
ops := instr.Operands
mnem = riscvNormalisePseudo(mnem)
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
var word uint32
@@ -677,6 +710,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// RET = epilogue (restore LR and close the frame when present) +
// uncompressed JALR X0, 0(X1) (the toolchain never compresses RET).
return riscvReturn(fi), nil
case "FUNCDATA":
// The assembler's bookkeeping statement, the expanded form of the
// GO_ARGS and NO_LOCAL_POINTERS macros: FUNCDATA $n, sym(SB)
// contributes no bytes, exactly as the toolchain's listing shows
// (the FUNCDATA entries and the instruction after them share a PC).
if len(ops) != 2 || !isImmOperand(ops[0]) {
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
}
return nil, nil
case "PCDATA":
// The other bookkeeping statement, the expanded form of
// GO_RESULTS_INITIALIZED: PCDATA $n, $m contributes no bytes too.
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
return nil, fmt.Errorf("PCDATA expects $n, $m")
}
return nil, nil
case "WORD":
// WORD $w lays down a raw 32-bit little-endian word.
if len(ops) != 1 {
@@ -739,6 +788,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
target = labelFromOperand(ops[0])
// JMP N(PC): the PC-relative slot form, resolved like the
// branches (the toolchain counts source instructions at a
// uniform 4 bytes, so JMP 0(PC) is a self-loop and JMP -3(PC)
// reaches twelve bytes back). It must be recognised before the
// indirect-register form, whose operand it resembles.
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
if err != nil {
return nil, err
}
offset := int32(off - pc)
if err := riscvCheckJumpOffset("", offset); err != nil {
return nil, err
}
word = riscvJType(0, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
// JMP (X5): an indirect branch, the toolchain's JALR X0, 0(X5).
if ops[0].Addr.Sym == nil && ops[0].Addr.Base != "" {
if ops[0].Addr.Offset != 0 || ops[0].Addr.Index != "" {
@@ -751,17 +816,6 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
if err != nil {
return nil, err
}
offset := int32(off - pc)
if err := riscvCheckJumpOffset("", offset); err != nil {
return nil, err
}
word = riscvJType(0, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
}
targetOff, ok := offsets[target]
if !ok {
@@ -810,7 +864,7 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// (MOVB/MOVH/MOVW and unsigned forms) select the access width, and
// MOVD/MOVF address the FP registers.
case "MOV", "MOVB", "MOVBU", "MOVH", "MOVHU", "MOVW", "MOVWU", "MOVF", "MOVD":
return encodeRISCVMov(instr, fi, relocs)
return encodeRISCVMov(instr, fi, relocs, lits)
// JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1).
case "JALR":
@@ -1316,7 +1370,7 @@ func isImmOperand(op *ast.Operand) bool {
// - MOV Rs, (Rd) register-relative store
// - MOV Rs, Rd register-to-register move (ADDI $0)
// - MOV $imm, Rd load immediate (ADDI or LUI+ADDIW)
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) {
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc, lits *riscvLiterals) ([]byte, error) {
ops := instr.Operands
if len(ops) != 2 {
return nil, fmt.Errorf("MOV expects 2 operands, got %d", len(ops))
@@ -1335,8 +1389,21 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
}
return encodeRISCVSBAddr(src.Imm.Sym, rd, relocs), nil
}
// MOV $sym+off(FP|SP), rd: the address of a frame slot as an
// immediate is the frame-adjusted offset against the hardware SP,
// the toolchain's ADDI $adj, SP, rd (argframe+0(FP) in the runtime's
// reflect trampolines is the spelling).
if src.Imm.Sym != nil && (src.Imm.Sym.Pseudo == "FP" || src.Imm.Sym.Pseudo == "SP") {
rd := regFromOperand(dst)
if rd < 0 {
return nil, fmt.Errorf("MOV $%s(%s): invalid destination register", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
}
_, off := riscvResolvePseudo(src.Imm.Sym, fi)
return riscvSPAddiBytes(rd, off), nil
}
// MOV $sym(FP/SP), rd, not supported: immediate symbol references
// other than SB cannot be encoded as a simple immediate.
// other than the frame pseudos cannot be encoded as a simple
// immediate.
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo != "" {
return nil, fmt.Errorf("MOV $%s(%s): unsupported immediate symbol reference (only SB is supported)", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
}
@@ -1344,11 +1411,14 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rd < 0 {
return nil, fmt.Errorf("MOV $imm: invalid destination register")
}
imm, err := riscvImm32FromOperand(src, false)
if err != nil {
return nil, err
imm := riscvOperandImm64(src)
if int64(int32(imm)) != imm {
// Beyond the signed 32-bit span the toolchain either builds the
// value from a shifted 32-bit part or loads it from the pooled
// $i64 constant it synthesises for the purpose.
return riscvLoadImm64(rd, imm, lits, relocs), nil
}
return encodeRISCVLoadImm(rd, imm), nil
return encodeRISCVLoadImm(rd, int32(imm)), nil
}
// Memory → register (load).
@@ -1557,6 +1627,186 @@ func splitRISCV32Imm(imm int32) (low, high int32) {
return low, high
}
// riscvNormalisePseudo rewrites the toolchain's UNDEF spelling onto EBREAK:
// the assembler accepts UNDEF where the hardware wants the trap instruction
// and emits ebreak (compressed to C.EBREAK under RVC), so every pass sees the
// canonical name.
func riscvNormalisePseudo(mnem string) string {
if strings.EqualFold(mnem, "UNDEF") {
return "EBREAK"
}
return mnem
}
// riscvOperandImm64 reads an immediate operand as a full signed 64-bit value,
// where immFromOperand would truncate to int32; the MOV immediate path uses
// it to classify the wide constants.
func riscvOperandImm64(op *ast.Operand) int64 {
if !op.Imm.HasVal {
return 0
}
v := op.Imm.Val
if op.Imm.Neg {
v = -v
}
return v
}
// riscvSplitShiftConst mirrors cmd/internal/obj/riscv's splitShiftConst: it
// looks for the signed 32-bit integer a constant can be rebuilt from with a
// left shift, a left-and-right shift pair (a run of ones), or a zero-extended
// 32-bit pattern. A constant that fits none of the shapes is materialised
// from the pooled $i64 data symbol instead.
func riscvSplitShiftConst(v int64) (imm int64, lsh int, rsh int, ok bool) {
// Rebuild from a signed 32-bit integer shifted left.
lsh = bits.TrailingZeros64(uint64(v))
c := v >> lsh
if int64(int32(c)) == c {
return c, lsh, 0, true
}
// Rebuild from a small negative constant: shift left into place, then
// shift the sign-extended ones run right.
rsh = bits.LeadingZeros64(uint64(v))
ones := bits.OnesCount64((uint64(v) >> lsh) >> 11)
if rsh+ones+lsh+11 == 64 {
c = (1<<11 | ((v >> lsh) & 0x7ff)) << 52 >> 52 // sign extend 12 bits
if lsh > 0 || c != -1 {
lsh += rsh
}
return c, lsh, rsh, true
}
// Rebuild from a zero-extended signed 32-bit integer.
if int64(uint32(c)) == c {
c = int64(int32(c))
lsh, rsh = 32, 32-lsh
return c, lsh, rsh, true
}
return 0, 0, 0, false
}
// riscvSPAddiBytes encodes ADDI rd, SP, imm for the frame-address immediates
// (the MOV $sym+off(FP|SP) form), using the compressed forms the toolchain
// picks under RVC: C.ADDI4SPN for a positive 4-byte multiple that fits, C.MV
// for the zero offset, the plain ADDI otherwise.
func riscvSPAddiBytes(rd int, imm int32) []byte {
if rd != 0 && imm == 0 {
return word16(rvcCR(0x8, uint32(rd), 2)) // C.MV rd, SP
}
if isRVCIntReg(rd) && imm > 0 && imm < 1024 && imm%4 == 0 {
return word16(rvcCIW(0x0, rvcReg3(rd), uint32(imm)))
}
return wordLE(riscvIType(riscvInstrTable["ADDI"], rd, 2, imm))
}
// riscvMovImm64Size returns the encoded byte length of MOV $imm, rd when the
// immediate sits outside the signed 32-bit span: the shifted-part sequences
// of riscvLoadImm64, or the 8-byte AUIPC+LD pool load.
func riscvMovImm64Size(rd int, imm int64) int {
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
if !ok {
return 8 // AUIPC + LD against the $i64 pool symbol
}
size := riscvMovImmSize(rd, int32(c))
if lsh > 0 {
size += riscvShiftImmSize(rd, true)
}
if rsh > 0 {
size += riscvShiftImmSize(rd, false)
}
return size
}
// riscvShiftImmSize returns the encoded size of one SLLI/SRLI expansion
// part: two bytes under RVC when the destination can carry a compressed
// shift (C.SLLI admits every register but X0, C.SRLI only X8 to X15), four
// otherwise.
func riscvShiftImmSize(rd int, left bool) int {
if rd != 0 && (left || isRVCIntReg(rd)) {
return 2
}
return 4
}
// riscvLoadImm64 encodes MOV $imm, rd for an immediate beyond the signed
// 32-bit span, mirroring the toolchain's instructionsForMOVConst: when a
// shifted 32-bit part rebuilds the value it emits that part (compressed like
// any written MOV) followed by the SLLI and SRLI shifts; otherwise it loads
// the constant from the pooled read-only $i64.<hex> symbol via AUIPC + LD
// and registers the literal so the data section carries its bytes.
func riscvLoadImm64(rd int, imm int64, lits *riscvLiterals, relocs *[]Reloc) []byte {
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
if !ok {
name := fmt.Sprintf("$i64.%016x", uint64(imm))
if lits != nil {
lits.add(name, riscvLiteralBytes(imm))
}
return encodeRISCVSBLoad(&ast.Symbol{Name: name}, rd, relocs)
}
out := encodeRISCVLoadImm(rd, int32(c))
if lsh > 0 {
out = append(out, riscvShiftImmBytes(rd, lsh, true)...)
}
if rsh > 0 {
out = append(out, riscvShiftImmBytes(rd, rsh, false)...)
}
return out
}
// riscvShiftImmBytes encodes one SLLI (left) or SRLI expansion part, using
// the compressed form the toolchain picks under RVC: C.SLLI admits every
// register but X0, C.SRLI only X8 to X15.
func riscvShiftImmBytes(rd, shamt int, left bool) []byte {
if rd != 0 && shamt >= 1 && shamt <= 63 && (left || isRVCIntReg(rd)) {
if left {
return word16(rvcSLLI(uint32(rd), uint32(shamt)&0x3F))
}
return word16(rvcCBShift(0x0, rvcReg3(rd), uint32(shamt)&0x3F))
}
enc := riscvEnc{0x13, 0x1, 0x00} // SLLI
imm := int32(shamt)
if !left {
enc = riscvEnc{0x13, 0x5, 0x00} // SRLI: funct6 000000, funct3 101
}
return wordLE(riscvIType(enc, rd, rd, imm))
}
// riscvLiteralBytes renders a 64-bit constant as the little-endian bytes the
// $i64 pool symbol holds.
func riscvLiteralBytes(v int64) []byte {
return []byte{byte(v), byte(v >> 8), byte(v >> 16), byte(v >> 24),
byte(v >> 32), byte(v >> 40), byte(v >> 48), byte(v >> 56)}
}
// RiscvLiteral is one pooled 64-bit constant: a MOV whose immediate sits
// beyond both the 32-bit span and the shift sequences loads its bits from a
// read-only data symbol named like the toolchain's $i64 pool.
type RiscvLiteral struct {
Name string
Data []byte
}
// riscvLiterals collects the pooled constants the MOV expansions refer to,
// deduplicated by name, in first-use order.
type riscvLiterals struct {
order []RiscvLiteral
seen map[string]bool
}
func (l *riscvLiterals) add(name string, data []byte) {
if l.seen == nil {
l.seen = map[string]bool{}
}
if !l.seen[name] {
l.seen[name] = true
l.order = append(l.order, RiscvLiteral{Name: name, Data: data})
}
}
func (l *riscvLiterals) list() []RiscvLiteral { return l.order }
// encodeRISCVItypeImmediate encodes an I-type arithmetic instruction, expanding
// large immediates for ADDI/ANDI/ORI/XORI into LUI+ADDIW+op (or two ADDIs for
// ADDI), matching the Go assembler.
@@ -1734,6 +1984,7 @@ func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
// RVC form. It returns the compressed instruction word and true on success.
func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
mnem := riscvCompressMnem(instr)
mnem = riscvNormalisePseudo(mnem)
ops := instr.Operands
// The immediate aliases fold onto their I-type mnemonics before
// compression: the toolchain compresses ADD $imm, rd as c.addi, exactly
+1 -1
View File
@@ -63,7 +63,7 @@ func riscvRegNum(name string) int {
return 24
case "X25", "S9":
return 25
case "X26", "S10":
case "X26", "S10", "CTXT":
return 26
case "X27", "S11", "g":
return 27
+154 -21
View File
@@ -10,8 +10,8 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// firstTextRISCV parses assembly source and returns the first TEXT function body.
@@ -33,7 +33,7 @@ func firstTextRISCV(t *testing.T, src string) *ast.Text {
// assembleRISCVHelper assembles one TEXT function and returns its code bytes.
func assembleRISCVHelper(t *testing.T, fn *ast.Text) []byte {
t.Helper()
code, _, _, _, _, err := assembleRISCV(fn)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -785,16 +785,140 @@ TEXT ·sys(SB), NOSPLIT, $0
}
}
func TestRISCV_MOV_sym_FP_error(t *testing.T) {
// MOV $sym(FP), rd should return an error (unsupported).
func TestRISCV_MOV_sym_FP(t *testing.T) {
// MOV $sym(FP), rd lowers to the frame-adjusted ADDI against SP: the
// toolchain's argframe spelling. A zero frame leaves the offset at the
// 8-byte link slot, compressed to C.ADDI4SPN.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·badfp(SB), NOSPLIT, $0
TEXT ·argfp(SB), NOSPLIT, $0
MOV $arg(FP), X10
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for MOV $arg(FP), got nil")
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// prologue (0: leaf, zero frame) + C.ADDI4SPN (2) + RET (4) = 6
want := []byte{0x28, 0x00, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_Bookkeeping(t *testing.T) {
// FUNCDATA and PCDATA contribute no bytes; UNDEF is the toolchain's
// ebreak, compressed to C.EBREAK under RVC.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·book(SB), NOSPLIT, $0-8
FUNCDATA $0, marks<>(SB)
PCDATA $1, $1
UNDEF
MOV $1, X10
MOV X10, ret+0(FP)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.EBREAK (2) + C.LI X10, 1 (2) + C.SWSP (2) + RET (4) = 10: the
// FUNCDATA and PCDATA statements contribute nothing.
want := []byte{0x02, 0x90, 0x05, 0x45, 0x2a, 0xe4, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_JMPPCRel(t *testing.T) {
// JMP N(PC): the displacement tracks the instruction N source slots
// away in the final layout (0 the jump itself, negative backwards).
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·slots(SB), NOSPLIT, $0-0
JMP 2(PC)
MOV $1, X11
MOV $2, X12
MOV X12, X11
JMP -3(PC)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// JMP 2(PC) lands on the C.MV six bytes ahead; JMP -3(PC) lands back on
// the first C.LI, six bytes behind.
want := []byte{
0x6f, 0x00, 0x60, 0x00, // JAL X0, 6
0x85, 0x45, // C.LI X11, 1
0x09, 0x46, // C.LI X12, 2
0xb2, 0x85, // C.MV X11, X12
0x6f, 0xf0, 0xbf, 0xff, // JAL X0, -6
0x67, 0x80, 0x00, 0x00, // RET
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVWideImm(t *testing.T) {
// Shift-sequence constants compress like the toolchain's expansion.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·wide(SB), NOSPLIT, $0-0
MOV $0x8000000000000000, X5
MOV $0x100000000, X5
MOV $0x000fffffffffffda, X5
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.LI -1, C.SLLI 63; C.LI 1, C.SLLI 32; C.LI -19, C.SLLI 13, SRLI 12.
want := []byte{
0xfd, 0x52, 0xfe, 0x12,
0x85, 0x42, 0x82, 0x12,
0xb5, 0x52, 0xb6, 0x02, 0x93, 0xd2, 0xc2, 0x00,
0x67, 0x80, 0x00, 0x00,
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVImmPool(t *testing.T) {
// A constant outside the shift shapes loads from the pooled $i64 data
// symbol via AUIPC+LD, named like the toolchain's pool.
src := `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0-8
MOV $0x0101010101010101, X16
MOV X16, ret+0(FP)
RET
`
f, errs := parser.Parse("pool_riscv64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
// AUIPC X16, 0 + LD X16, 0(X16): the relocation pair carries the symbol.
wantCode := []byte{0x17, 0x08, 0x00, 0x00, 0x03, 0x38, 0x08, 0x00}
if string(img.Code[0:8]) != string(wantCode) {
t.Errorf("pool load: got % x", img.Code[0:8])
}
var lit *DataSymbol
for i := range img.DataSyms {
if img.DataSyms[i].Name == "$i64.0101010101010101" {
lit = &img.DataSyms[i]
}
}
if lit == nil {
t.Fatalf("pool symbol missing: %v", img.DataSyms)
}
wantData := []byte{0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01}
if string(img.Data[lit.Offset:lit.Offset+8]) != string(wantData) {
t.Errorf("pool bytes: got % x", img.Data[lit.Offset:lit.Offset+8])
}
}
@@ -805,7 +929,7 @@ TEXT ·calltest(SB), NOSPLIT, $0
CALL ext(SB)
RET
`)
code, _, relocs, _, _, err := assembleRISCV(fn)
code, _, relocs, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -834,7 +958,7 @@ TEXT ·calllocal(SB), NOSPLIT, $0
sub:
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
_, _, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for CALL to local label, got nil")
}
@@ -868,7 +992,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
t.Helper()
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
instr := fn.Body[0].(*ast.Instr)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil, nil)
}
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
@@ -917,7 +1041,7 @@ func TestRISCVBranchFarBody(t *testing.T) {
}
sb.WriteString("done:\n\tRET\n")
fn := firstTextRISCV(t, sb.String())
out, _, _, _, _, err := assembleRISCV(fn)
out, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
@@ -943,7 +1067,7 @@ TEXT ·csrhi(SB), NOSPLIT, $0
CSRRW $4096, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected an out-of-range error for CSR $4096, got none")
}
fn = firstTextRISCV(t, `#include "textflag.h"
@@ -951,25 +1075,24 @@ TEXT ·csrmax(SB), NOSPLIT, $0
CSRRW $4095, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("CSR $4095 must assemble: %v", err)
}
}
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
// span are diagnosed instead of silently truncated to their low 32 bits (the
// toolchain materialises such constants via SLLI expansion, which this
// assembler does not implement).
// span are diagnosed instead of silently truncated to their low 32 bits for
// the I-type arithmetic; the MOV forms materialise the wide constant instead
// (shift sequence or pooled load), like the toolchain.
func TestRISCV_Imm64Rejected(t *testing.T) {
cases := []string{
"MOV $0x123456789, X10",
"ADDI $0x100000000, X10, X11",
"ANDI $-0x800000001, X10, X11",
"SUB $0x100000000, X10, X11",
}
for _, src := range cases {
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Errorf("%s: expected an out-of-range error, got none", src)
}
}
@@ -982,9 +1105,19 @@ TEXT ·edge(SB), NOSPLIT, $0
SUB $0x80000000, X12, X13
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("int32-span immediates must assemble: %v", err)
}
// Beyond the span the MOV forms materialise the constant like the
// toolchain instead of diagnosing it.
fn = firstTextRISCV(t, `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0
MOV $0x123456789, X10
RET
`)
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("MOV with a 64-bit immediate must assemble: %v", err)
}
}
// riscvWants decodes code as little-endian words and pins each one; the
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
)
// RISC-V frame mapping, matching the Go toolchain's riscv64 backend.
+1 -1
View File
@@ -6,7 +6,7 @@ package asm
import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestRISCVFrameSpadjAndLines checks that a framed function records its
+1 -1
View File
@@ -13,7 +13,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// TestGOObjectRISCVCallReloc checks that CALL sym(SB) emits a single JAL
+1 -1
View File
@@ -8,7 +8,7 @@
// to the arch package; the AST records syntax only.
package ast
import "sourcedock.dev/petrbalvin/gasm-devkit/token"
import "sourcedock.dev/petrbalvin/gasm-sdk/token"
// File is the parsed representation of one .s source file.
type File struct {
+1 -1
View File
@@ -6,7 +6,7 @@ package ast
import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
"sourcedock.dev/petrbalvin/gasm-sdk/token"
)
func pos(line, col int) token.Position { return token.Position{Line: line, Column: col} }
+1 -1
View File
@@ -17,7 +17,7 @@ import (
"regexp"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
)
// go_asm.h is the header the Go compiler writes for every package that
+57 -3
View File
@@ -5,6 +5,7 @@ package main
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
@@ -392,14 +393,67 @@ func TestRunCorpusAuditGOOS(t *testing.T) {
}
}
// TestRunCorpusAuditBuildConstraint covers the //go:build classification end
// to end: a generic-named file whose constraint admits one target is
// attempted there alone (cpu_x86.s on amd64), and a file whose constraint
// admits none of the four targets is never attempted (the msan and
// goexperiment trees).
func TestRunCorpusAuditBuildConstraint(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("x86.s", "//go:build 386 || amd64\n\nTEXT \xc2\xb7f(SB), NOSPLIT, $0\n\tRET\n")
write("racey.s", "//go:build race\n\nTEXT \xc2\xb7r(SB), NOSPLIT, $0\n\tRET\n")
write("plain.s", "TEXT \xc2\xb7p(SB), NOSPLIT, $0\n\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
tally := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
if stats.narrowed != 1 || stats.excluded != 1 || stats.generic != 1 {
t.Errorf("buckets = narrowed %d, excluded %d, generic %d; want 1, 1, 1", stats.narrowed, stats.excluded, stats.generic)
}
if a := tally("amd64"); a.attempted != 2 || a.assembled != 2 {
t.Errorf("amd64 = %d/%d, want 2/2 (x86.s and plain.s)", a.assembled, a.attempted)
}
for _, name := range []string{"arm64", "riscv64", "loong64"} {
if a := tally(name); a.attempted != 1 || a.assembled != 1 {
t.Errorf("%s = %d/%d, want 1/1 (plain.s only)", name, a.assembled, a.attempted)
}
}
if stats.full != 2 {
t.Errorf("full = %d, want 2 (x86.s over its one target, plain.s over all four)", stats.full)
}
}
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
// the runtime package, whose header the toolchain's own -asmhdr output was
// sampled from. Skipped in short mode: it type-checks the whole package.
// the runtime package of the ambient toolchain, whose header the toolchain's
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
// the whole package. The GOROOT comes from the go command itself, so the
// test follows whatever toolchain the host provides.
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
if testing.Short() {
t.Skip("type-checks the whole runtime package")
}
dir, err := generateGoAsmHeader("/usr/local/go/src/runtime", "", "amd64", t.TempDir())
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
t.Skipf("no Go toolchain: %v", err)
}
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
}
+170 -34
View File
@@ -5,6 +5,7 @@ package main
import (
"fmt"
"go/build/constraint"
"maps"
"os"
"os/exec"
@@ -15,9 +16,9 @@ import (
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// cmdAuditInstructions cross-checks a gasm encoder against the Go toolchain's
@@ -38,7 +39,7 @@ import (
// construction and are excluded from the diff; the other architectures list
// their conditional branches outright.
func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]", `
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
@@ -55,16 +56,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
the encodability backlog by frequency rather than by table order. With
-list the report also prints every failing file with its reason, per
architecture.
`)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
if err := fs.Parse(args); err != nil {
return err
}
if *corpus {
return cmdAuditCorpus(fs.Args(), dirs)
return cmdAuditCorpus(fs.Args(), dirs, *list)
}
archName := "amd64"
switch n := len(fs.Args()); {
@@ -403,20 +407,30 @@ type corpusTally struct {
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
fails []corpusFailure // every failure, in file order, for --list
}
func (t *corpusTally) fail(path, reason string) {
// corpusFailure is one failed attempt, recorded for the --list report.
type corpusFailure struct {
path string
reason string
detail string
}
func (t *corpusTally) fail(path string, err error) {
reason := corpusReason(err)
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
}
// cmdAuditCorpus implements audit-instructions --corpus. The include
// directories carry #include resolution over a corpus whose files refer to
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
// needs.
func cmdAuditCorpus(args []string, dirs includeDirs) error {
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
@@ -454,7 +468,7 @@ func cmdAuditCorpus(args []string, dirs includeDirs) error {
if err != nil {
return err
}
printCorpusStats(stats)
printCorpusStats(stats, list)
return nil
}
@@ -463,8 +477,10 @@ type corpusStats struct {
root string
files int
generic int // files attempted for all four architectures
narrowed int // files whose //go:build admits a proper subset of the four
excluded int // files whose //go:build admits none of the four: never compiled
otherPort int // files named for another Go port: never attempted
full int // files that assembled for every target architecture
full int // files that assembled for every applicable target architecture
targets []corpusTarget
tallies []*corpusTally
}
@@ -475,8 +491,8 @@ type corpusStats struct {
// set, even when gasm does not support the architecture.
var goPortSuffixes = []string{
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
"s390x", "wasm",
"mips64le", "mipsle", "mips64x", "mipsx", "ppc64", "ppc64le",
"ppc64x", "riscv", "riscv64", "s390x", "wasm",
}
// otherPortFile reports whether the file belongs to a build no supported
@@ -560,6 +576,57 @@ func otherGOOSFile(path string) bool {
return false
}
// buildConstraint returns the file's leading //go:build expression, or nil
// when the file carries none. The constraint governs the same header block
// go/build reads: blank lines and comments may precede it, and the first
// line that is neither ends the block. A constraint that does not parse
// narrows nothing, so the file stays in the attempted set: the audit must
// never exclude a file the toolchain would compile.
func buildConstraint(src string) constraint.Expr {
for line := range strings.SplitSeq(src, "\n") {
t := strings.TrimSpace(line)
switch {
case t == "":
continue
case strings.HasPrefix(t, "//"):
if constraint.IsGoBuild(t) {
e, err := constraint.Parse(t)
if err != nil {
return nil
}
return e
}
continue
default:
return nil
}
}
return nil
}
// unixOS is go/build's unixOS set: the GOOSes the unix build tag admits.
var unixOS = map[string]bool{
"aix": true, "android": true, "darwin": true, "dragonfly": true,
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
"linux": true, "netbsd": true, "openbsd": true, "solaris": true,
}
// constraintTags answers the build tags a plain `go build` sets for a
// target: the GOOS and GOARCH, gc, and unix on the unix-like GOOSes. No
// experiment, sanitiser or cgo tag is ever true: the audit models the
// default build, and no GOROOT assembly file's constraint hinges on cgo.
func constraintTags(goarch, goos string) func(string) bool {
return func(tag string) bool {
switch tag {
case goarch, goos, "gc":
return true
case "unix":
return unixOS[goos]
}
return false
}
}
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
@@ -577,8 +644,8 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
tallies[i] = &corpusTally{reasons: map[string]int{}, example: map[string]string{}}
}
// full is the north-star number: a file counts when every architecture
// its name allows assembles it.
full, generic, otherPort := 0, 0, 0
// its build admits assembles it.
full, generic, otherPort, narrowedCount, excluded := 0, 0, 0, 0, 0
// Header generation is created on first use, so a corpus with no
// go_asm.h includes never pays for a temp directory.
@@ -601,28 +668,78 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
// invisible to a file-name rule).
goos := goosFromFilename(path)
// The GOOS the header generation type-checks under follows the
// file's name when the name carries one; the ambient GOOS is the
// honest guess otherwise.
namedArch := arch.FromFilename(path)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
other := false
switch {
case namedArch != arch.Unknown:
for i, tg := range targets {
if tg.a == a {
if tg.a == namedArch {
wanted = append(wanted, i)
}
}
} else if otherPortFile(path) {
case otherPortFile(path):
// A file named for a Go port gasm does not support (arm,
// 386, s390x, ...) or for another GOOS is compiled by no
// supported-arch build, so it is neither generic nor a
// per-arch attempt: counting it as generic would make the
// headline unreachably low for reasons no supported target
// can fix.
other = true
otherPort++
} else {
generic++
default:
for i := range targets {
wanted = append(wanted, i)
}
}
// A //go:build constraint narrows the set of targets the file is
// assembled for, the way the go command compiles the file only for
// the targets the expression admits: cpu_x86.s belongs to the x86
// build alone, and a file whose constraint admits none of the four
// targets (the goexperiment.runtimesecret and msan trees) is
// compiled by no supported build. The tags mirror what a plain
// `go build` sets: the GOOS and GOARCH, gc, and unix on the
// unix-like GOOSes; no experiment, sanitiser or cgo tag is ever
// true. The GOOS is the file's own when the name carries one,
// else the ambient one.
goosForEval := goos
if goosForEval == "" {
goosForEval = runtime.GOOS
}
narrowed := false
if len(wanted) > 0 {
if ce := buildConstraint(src); ce != nil {
kept := make([]int, 0, len(wanted))
for _, i := range wanted {
tg := targets[i]
if ce.Eval(constraintTags(goarchName(tg.a), goosForEval)) {
kept = append(kept, i)
}
}
if len(kept) < len(wanted) {
narrowed = true
}
wanted = kept
}
}
switch {
case other:
// already tallied above
case len(wanted) == 0:
excluded++
case namedArch != arch.Unknown:
// a per-arch attempt over the constraint's subset
case narrowed:
narrowedCount++
default:
generic++
}
// A file that includes go_asm.h parses against a per-target header:
// the defines differ per architecture (internal/cpu's layout, for
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
@@ -645,21 +762,22 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: append(slices.Clone(dirs), hdrDir),
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
if len(errs) > 0 {
ok = false
t.fail(path, corpusReason(errs[0]))
t.fail(path, errs[0])
continue
}
if _, err := assembleFile(tg.a, f); err != nil {
if _, err := assembleFile(tg.a, f, goos); err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -670,21 +788,28 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
// The parse carries the target's platform predefines, so it
// cannot be shared across targets the way a header-free file's
// could: a #ifdef GOARCH_arm block must be live on arm64 and
// dead everywhere else.
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: dirs,
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
_, err = assembleFile(tg.a, f, goos)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -698,6 +823,8 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
root: root,
files: len(files),
generic: generic,
narrowed: narrowedCount,
excluded: excluded,
otherPort: otherPort,
full: full,
targets: targets,
@@ -706,14 +833,16 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
func printCorpusStats(s *corpusStats, list bool) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d narrowed by //go:build; %d excluded by //go:build; %d named for other Go ports, never attempted)\n",
s.root, s.files, s.generic, s.narrowed, s.excluded, s.otherPort)
// The rate is over the files a supported build would attempt: the
// other ports' files sit in the count for completeness but can never
// assemble, so counting them in the denominator would report the gap
// of architectures gasm deliberately does not target.
attemptable := max(s.files-s.otherPort, 1)
fmt.Printf(" assemble for every target architecture: %d of %d attemptable (%.1f%%)\n", s.full, attemptable, 100*float64(s.full)/float64(attemptable))
// other ports' files and the ones no supported target compiles sit in
// the count for completeness but can never assemble, so counting them
// in the denominator would report the gap of platforms gasm
// deliberately does not target.
attemptable := max(s.files-s.otherPort-s.excluded, 1)
fmt.Printf(" assemble for every applicable target: %d of %d attemptable (%.1f%%)\n", s.full, attemptable, 100*float64(s.full)/float64(attemptable))
for i, tg := range s.targets {
t := s.tallies[i]
fmt.Printf(" %s: %d/%d attempted\n", tg.name, t.assembled, t.attempted)
@@ -721,6 +850,13 @@ func printCorpusStats(s *corpusStats) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
if !list {
continue
}
for _, f := range t.fails {
fmt.Printf(" FAIL %s\n", f.path)
fmt.Printf(" %s: %s\n", f.reason, f.detail)
}
}
}
+61 -1
View File
@@ -4,9 +4,10 @@
package main
import (
"runtime"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
)
func TestDerivedFamily(t *testing.T) {
@@ -64,3 +65,62 @@ func TestGasmEncodable(t *testing.T) {
}
}
}
// TestBuildConstraint pins the //go:build reader: the constraint governs the
// leading comment block, the first non-comment line ends it (a tag below a
// #include governs nothing, exactly as go/build drops it), and a file
// without one admits every target.
func TestBuildConstraint(t *testing.T) {
admits := func(src, goarch, goos string) bool {
t.Helper()
e := buildConstraint(src)
if e == nil {
return true
}
return e.Eval(constraintTags(goarch, goos))
}
const ret = "TEXT \xc2\xb7f(SB), NOSPLIT, $0\n\tRET\n"
cases := []struct {
name string
src string
amd64, arm64 bool
}{
{"no constraint", ret, true, true},
{"x86 only", "//go:build 386 || amd64\n\n" + ret, true, false},
{"arm64 and linux", "//go:build arm64 && linux\n\n" + ret, false, true},
{"msan never", "//go:build msan\n\n" + ret, false, false},
{"experiment never", "//go:build goexperiment.runtimesecret\n\n" + ret, false, false},
{"below an include governs nothing", "#include \"textflag.h\"\n//go:build amd64\n" + ret, true, true},
{"unparsable narrows nothing", "//go:build (amd64\n" + ret, true, true},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := admits(c.src, "amd64", runtime.GOOS); got != c.amd64 {
t.Errorf("amd64 admission = %v, want %v", got, c.amd64)
}
if got := admits(c.src, "arm64", runtime.GOOS); got != c.arm64 {
t.Errorf("arm64 admission = %v, want %v", got, c.arm64)
}
})
}
}
// TestConstraintTags pins the tag set a plain `go build` sets: the GOOS and
// GOARCH, gc, unix on the unix-like GOOSes; nothing else is ever true.
func TestConstraintTags(t *testing.T) {
ok := constraintTags("amd64", "linux")
for _, tag := range []string{"amd64", "linux", "gc", "unix"} {
if !ok(tag) {
t.Errorf("tag %q = false, want true", tag)
}
}
for _, tag := range []string{"arm64", "freebsd", "darwin", "cgo", "race", "msan", "goexperiment.runtimesecret"} {
if ok(tag) {
t.Errorf("tag %q = true, want false", tag)
}
}
fb := constraintTags("arm64", "freebsd")
if !fb("unix") {
t.Error("unix on freebsd = false, want true")
}
}
+2 -2
View File
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build !linux
//go:build !(linux || (freebsd && (amd64 || arm64 || riscv64)))
package main
@@ -11,6 +11,6 @@ import (
)
func cmdDebug(args []string) int {
fmt.Fprintln(os.Stderr, "gasm debug: the interactive debugger requires Linux (ptrace)")
fmt.Fprintln(os.Stderr, "gasm debug: the interactive debugger requires Linux or FreeBSD (ptrace)")
return 1
}
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
package main
@@ -13,8 +13,8 @@ import (
"strings"
"time"
"sourcedock.dev/petrbalvin/gasm-devkit/debug"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/debug"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
func cmdDebug(args []string) int {
+4 -4
View File
@@ -9,9 +9,9 @@ import (
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// cmdDis disassembles machine code: either a raw binary (standard input with
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
if len(errs) > 0 {
return 1
}
img, err := assembleFile(target, f)
img, err := assembleFile(target, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 1
+43 -20
View File
@@ -27,15 +27,15 @@ import (
"sync"
"syscall"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/format"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/lint"
"sourcedock.dev/petrbalvin/gasm-devkit/lsp"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/format"
"sourcedock.dev/petrbalvin/gasm-sdk/lexer"
"sourcedock.dev/petrbalvin/gasm-sdk/lint"
"sourcedock.dev/petrbalvin/gasm-sdk/lsp"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// version reports the release the toolchain recorded for this build: the
@@ -574,7 +574,7 @@ naming the package.
defer cleanup()
dirs = append(dirs, hdrDir)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
@@ -582,7 +582,7 @@ naming the package.
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, goos)
if err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1
@@ -789,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
return 1
}
// platformPredefines mirrors the go command's assembler invocation, which
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
// exactly those names, so an assembler without them cannot see the platform
// definitions at all.
func platformPredefines(goarch, goos string) map[string]string {
return map[string]string{
"GOARCH_" + goarch: "1",
"GOOS_" + goos: "1",
}
}
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
// and nothing else.
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
goos := fileGoos
if goos == "" {
goos = runtime.GOOS
}
return platformPredefines(goarch, goos)
}
// assembleFile assembles a parsed file for the given architecture and returns the image.
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
switch targetArch {
case arch.AMD64:
return asm.AssembleFile(f)
return asm.AssembleFile(f, asm.WithGOOS(goos))
case arch.RISCV:
return asm.AssembleFileRISCV(f)
case arch.ARM64:
@@ -813,18 +836,18 @@ func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image,
if err != nil {
return nil, err
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return nil, fmt.Errorf("parse errors")
}
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
return assembleFile(target, f, "")
}
// printByteDiff shows the first few byte differences between two code blocks.
@@ -920,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
if len(errs) > 0 {
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
+2 -2
View File
@@ -14,8 +14,8 @@ import (
"syscall"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
)
const clean = "#include \"textflag.h\"\n" +
+2 -2
View File
@@ -11,8 +11,8 @@ import (
"os"
"strings"
gasmast "sourcedock.dev/petrbalvin/gasm-devkit/ast"
gasmparser "sourcedock.dev/petrbalvin/gasm-devkit/parser"
gasmast "sourcedock.dev/petrbalvin/gasm-sdk/ast"
gasmparser "sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
// cmdScaffold generates a differential test skeleton for every kernel in a
+1 -1
View File
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
package debug
+56
View File
@@ -0,0 +1,56 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
// memory and returns its text representation and length in bytes.
func (s *Session) Disassemble(addr uint64) (string, int, error) {
mem, err := s.ReadMemory(addr, 15)
if err != nil {
return "", 0, err
}
ins, err := disasm.Decode(arch.AMD64, mem, addr)
if err != nil {
return "", 0, err
}
return ins.Text, ins.Len, nil
}
// DisassembleN decodes up to n instructions starting at addr and returns
// them as a formatted string with addresses and byte offsets.
func (s *Session) DisassembleN(addr uint64, n int) string {
var result strings.Builder
pc := addr
for range n {
text, length, err := s.Disassemble(pc)
if err != nil {
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
break
}
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
if length == 0 {
length = 1
}
pc += uint64(length)
}
return result.String()
}
// isCallInsn reports whether disassembled text (x86asm.IntelSyntax) is a
// call. The first token must match exactly: a prefix test would also catch
// unrelated mnemonics.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
return strings.ToLower(m) == "call"
}
+60
View File
@@ -0,0 +1,60 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && arm64
package debug
import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
// memory and returns its text representation and length in bytes.
func (s *Session) Disassemble(addr uint64) (string, int, error) {
mem, err := s.ReadMemory(addr, 4)
if err != nil {
return "", 0, err
}
ins, err := disasm.Decode(arch.ARM64, mem, addr)
if err != nil {
return "", 0, err
}
return ins.Text, ins.Len, nil
}
// DisassembleN decodes up to n instructions starting at addr and returns
// them as a formatted string with addresses and byte offsets.
func (s *Session) DisassembleN(addr uint64, n int) string {
var result strings.Builder
pc := addr
for range n {
text, length, err := s.Disassemble(pc)
if err != nil {
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
break
}
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
if length == 0 {
length = 1
}
pc += uint64(length)
}
return result.String()
}
// isCallInsn reports whether disassembled text (arm64asm.GoSyntax) is a
// call. GoSyntax renders bl as CALL; the native mnemonic is accepted too.
// The first token must match exactly so branches never match.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
switch strings.ToLower(m) {
case "call", "bl":
return true
}
return false
}
+62
View File
@@ -0,0 +1,62 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && riscv64
package debug
import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
// memory and returns its text representation and length in bytes.
func (s *Session) Disassemble(addr uint64) (string, int, error) {
mem, err := s.ReadMemory(addr, 4)
if err != nil {
return "", 0, err
}
ins, err := disasm.Decode(arch.RISCV, mem, addr)
if err != nil {
return "", 0, err
}
return ins.Text, ins.Len, nil
}
// DisassembleN decodes up to n instructions starting at addr and returns
// them as a formatted string with addresses and byte offsets.
func (s *Session) DisassembleN(addr uint64, n int) string {
var result strings.Builder
pc := addr
for range n {
text, length, err := s.Disassemble(pc)
if err != nil {
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
break
}
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
if length == 0 {
length = 1
}
pc += uint64(length)
}
return result.String()
}
// isCallInsn reports whether disassembled text (riscv64asm.GoSyntax) is a
// call. GoSyntax renders jal and jalr calls as CALL; the native mnemonics
// are accepted too. The first token must match exactly: a prefix test on
// "bl" would catch branches on other architectures, and jalr as ret prints
// RET, which must not be stepped over.
func isCallInsn(text string) bool {
m, _, _ := strings.Cut(text, " ")
switch strings.ToLower(m) {
case "call", "jal", "jalr":
return true
}
return false
}
+2 -2
View File
@@ -9,8 +9,8 @@ import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
+2 -2
View File
@@ -9,8 +9,8 @@ import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
+2 -2
View File
@@ -9,8 +9,8 @@ import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
+2 -2
View File
@@ -9,8 +9,8 @@ import (
"fmt"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
)
// Disassemble decodes the instruction at the given address in the debuggee's
+136
View File
@@ -0,0 +1,136 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
import "fmt"
func printRegs(regs *Regs, codeBase, funcOff uint64) {
fmt.Printf(" RIP = %#016x (func+%#x)\n", regs.RIP, regs.RIP-codeBase-funcOff)
fmt.Printf(" RSP = %#016x RBP = %#016x\n", regs.RSP, regs.RBP)
fmt.Printf(" RAX = %#016x RBX = %#016x\n", regs.RAX, regs.RBX)
fmt.Printf(" RCX = %#016x RDX = %#016x\n", regs.RCX, regs.RDX)
fmt.Printf(" RSI = %#016x RDI = %#016x\n", regs.RSI, regs.RDI)
fmt.Printf(" R8 = %#016x R9 = %#016x\n", regs.R8, regs.R9)
fmt.Printf(" R10 = %#016x R11 = %#016x\n", regs.R10, regs.R11)
fmt.Printf(" R12 = %#016x R13 = %#016x\n", regs.R12, regs.R13)
fmt.Printf(" R14 = %#016x R15 = %#016x\n", regs.R14, regs.R15)
fmt.Printf(" RFLAGS = %#x [%s]\n", regs.RFLAGS, decodeRflags(regs.RFLAGS))
}
func printVectorRegs(v *VectorRegs) {
fmt.Println("\n Vector registers (YMM):")
for i := 0; i < 16; i += 2 {
fmt.Printf(" YMM%-2d = ", i)
printYMM(v.YMM[i][:])
fmt.Printf(" YMM%-2d = ", i+1)
printYMM(v.YMM[i+1][:])
fmt.Println()
}
}
func printYMM(b []byte) {
for j := 0; j < 32; j += 4 {
v := uint32(b[j]) | uint32(b[j+1])<<8 | uint32(b[j+2])<<16 | uint32(b[j+3])<<24
fmt.Printf("%08x ", v)
}
}
func decodeRflags(f uint64) string {
var flags string
if f&1 != 0 {
flags += "CF "
}
if f&(1<<2) != 0 {
flags += "PF "
}
if f&(1<<4) != 0 {
flags += "AF "
}
if f&(1<<6) != 0 {
flags += "ZF "
}
if f&(1<<7) != 0 {
flags += "SF "
}
if f&(1<<8) != 0 {
flags += "TF "
}
if f&(1<<9) != 0 {
flags += "IF "
}
if f&(1<<10) != 0 {
flags += "DF "
}
if f&(1<<11) != 0 {
flags += "OF "
}
if flags == "" {
return "none"
}
return flags[:len(flags)-1]
}
// SetReg modifies a register value in the debuggee.
func (s *Session) SetReg(name string, value uint64) error {
regs, err := s.GetRegs()
if err != nil {
return err
}
switch name {
case "rax", "eax", "ax", "al":
regs.RAX = value
case "rbx", "ebx", "bx", "bl":
regs.RBX = value
case "rcx", "ecx", "cx", "cl":
regs.RCX = value
case "rdx", "edx", "dx", "dl":
regs.RDX = value
case "rsi", "esi", "si":
regs.RSI = value
case "rdi", "edi", "di":
regs.RDI = value
case "rbp", "ebp", "bp":
regs.RBP = value
case "rsp", "esp", "sp":
regs.RSP = value
case "r8":
regs.R8 = value
case "r9":
regs.R9 = value
case "r10":
regs.R10 = value
case "r11":
regs.R11 = value
case "r12":
regs.R12 = value
case "r13":
regs.R13 = value
case "r14":
regs.R14 = value
case "r15":
regs.R15 = value
case "rip", "eip":
regs.RIP = value
default:
return fmt.Errorf("debug: unknown register %q", name)
}
return s.SetRegs(&regs)
}
// archReturnAddr reads the return address of the current frame (amd64
// ABI0 convention). A function that contains a CALL (or has a frame) is
// assembled with the prologue PUSHQ BP; MOVQ SP, BP, so mid-function the
// word at SP is the saved caller BP, a stack address, and the return
// address sits further up. Walk the stack from SP and take the first word
// that lies in an executable mapping: stack and data words never do, a
// return address always does. FreeBSD exposes no mapping list, so the
// walk degenerates to the raw entry convention, [SP] before any push.
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
return s.Peek(regs.RSP)
}
// archSPLabel returns the SP register name for display.
func archSPLabel() string { return "RSP" }
+127
View File
@@ -0,0 +1,127 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && arm64
package debug
import (
"encoding/binary"
"fmt"
)
func printRegs(regs *Regs, codeBase, funcOff uint64) {
fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff)
fmt.Printf(" SP = %#016x FP = %#016x\n", regs.SP, regs.X29)
fmt.Printf(" LR = %#016x\n", regs.X30)
fmt.Printf(" X0 = %#016x X1 = %#016x\n", regs.X0, regs.X1)
fmt.Printf(" X2 = %#016x X3 = %#016x\n", regs.X2, regs.X3)
fmt.Printf(" X4 = %#016x X5 = %#016x\n", regs.X4, regs.X5)
fmt.Printf(" X6 = %#016x X7 = %#016x\n", regs.X6, regs.X7)
fmt.Printf(" X8 = %#016x X9 = %#016x\n", regs.X8, regs.X9)
fmt.Printf(" X10 = %#016x X11 = %#016x\n", regs.X10, regs.X11)
fmt.Printf(" X12 = %#016x X13 = %#016x\n", regs.X12, regs.X13)
fmt.Printf(" X14 = %#016x X15 = %#016x\n", regs.X14, regs.X15)
fmt.Printf(" X16 = %#016x X17 = %#016x\n", regs.X16, regs.X17)
fmt.Printf(" X18 = %#016x X19 = %#016x\n", regs.X18, regs.X19)
fmt.Printf(" X20 = %#016x X21 = %#016x\n", regs.X20, regs.X21)
fmt.Printf(" X22 = %#016x X23 = %#016x\n", regs.X22, regs.X23)
fmt.Printf(" X24 = %#016x X25 = %#016x\n", regs.X24, regs.X25)
fmt.Printf(" X26 = %#016x X27 = %#016x\n", regs.X26, regs.X27)
fmt.Printf(" X28 = %#016x PSTATE = %#x\n", regs.X28, regs.PSTATE)
}
func printVectorRegs(v *VectorRegs) {
fmt.Println("\n Vector registers (V0-V31):")
for i := 0; i < 32; i += 2 {
fmt.Printf(" V%-2d = %016x%016x\n", i, binary.LittleEndian.Uint64(v.V[i][8:16]), binary.LittleEndian.Uint64(v.V[i][0:8]))
fmt.Printf(" V%-2d = %016x%016x\n", i+1, binary.LittleEndian.Uint64(v.V[i+1][8:16]), binary.LittleEndian.Uint64(v.V[i+1][0:8]))
}
}
// SetReg modifies a register value in the debuggee.
func (s *Session) SetReg(name string, value uint64) error {
regs, err := s.GetRegs()
if err != nil {
return err
}
switch name {
case "x0":
regs.X0 = value
case "x1":
regs.X1 = value
case "x2":
regs.X2 = value
case "x3":
regs.X3 = value
case "x4":
regs.X4 = value
case "x5":
regs.X5 = value
case "x6":
regs.X6 = value
case "x7":
regs.X7 = value
case "x8":
regs.X8 = value
case "x9":
regs.X9 = value
case "x10":
regs.X10 = value
case "x11":
regs.X11 = value
case "x12":
regs.X12 = value
case "x13":
regs.X13 = value
case "x14":
regs.X14 = value
case "x15":
regs.X15 = value
case "x16":
regs.X16 = value
case "x17":
regs.X17 = value
case "x18":
regs.X18 = value
case "x19":
regs.X19 = value
case "x20":
regs.X20 = value
case "x21":
regs.X21 = value
case "x22":
regs.X22 = value
case "x23":
regs.X23 = value
case "x24":
regs.X24 = value
case "x25":
regs.X25 = value
case "x26":
regs.X26 = value
case "x27":
regs.X27 = value
case "x28":
regs.X28 = value
case "x29", "fp":
regs.X29 = value
case "x30", "lr":
regs.X30 = value
case "sp":
regs.SP = value
case "pc":
regs.PC = value
default:
return fmt.Errorf("debug: unknown register %q", name)
}
return s.SetRegs(&regs)
}
// archReturnAddr reads the return address from LR (arm64 convention).
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
return regs.X30, nil
}
// archSPLabel returns the SP register name for display.
func archSPLabel() string { return "SP" }
+121
View File
@@ -0,0 +1,121 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && riscv64
package debug
import "fmt"
func printRegs(regs *Regs, codeBase, funcOff uint64) {
fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff)
fmt.Printf(" SP = %#016x FP = %#016x\n", regs.Sp, regs.S0)
fmt.Printf(" RA = %#016x\n", regs.Ra)
fmt.Printf(" A0 = %#016x A1 = %#016x\n", regs.A0, regs.A1)
fmt.Printf(" A2 = %#016x A3 = %#016x\n", regs.A2, regs.A3)
fmt.Printf(" A4 = %#016x A5 = %#016x\n", regs.A4, regs.A5)
fmt.Printf(" A6 = %#016x A7 = %#016x\n", regs.A6, regs.A7)
fmt.Printf(" T0 = %#016x T1 = %#016x\n", regs.T0, regs.T1)
fmt.Printf(" T2 = %#016x T3 = %#016x\n", regs.T2, regs.T3)
fmt.Printf(" T4 = %#016x T5 = %#016x\n", regs.T4, regs.T5)
fmt.Printf(" T6 = %#016x\n", regs.T6)
fmt.Printf(" S1 = %#016x S2 = %#016x\n", regs.S1, regs.S2)
fmt.Printf(" S3 = %#016x S4 = %#016x\n", regs.S3, regs.S4)
fmt.Printf(" S5 = %#016x S6 = %#016x\n", regs.S5, regs.S6)
fmt.Printf(" S7 = %#016x S8 = %#016x\n", regs.S7, regs.S8)
fmt.Printf(" S9 = %#016x S10 = %#016x\n", regs.S9, regs.S10)
fmt.Printf(" S11 = %#016x\n", regs.S11)
}
func printVectorRegs(v *VectorRegs) {
fmt.Println("\n FP registers (F0-F31):")
for i := 0; i < 32; i += 2 {
fmt.Printf(" F%-2d = %#018x F%-2d = %#018x\n", i, v.F[i], i+1, v.F[i+1])
}
fmt.Printf(" FCSR = %#x\n", v.FCSR)
}
// SetReg modifies a register value in the debuggee.
func (s *Session) SetReg(name string, value uint64) error {
regs, err := s.GetRegs()
if err != nil {
return err
}
switch name {
case "pc":
regs.PC = value
case "ra", "x1":
regs.Ra = value
case "sp", "x2":
regs.Sp = value
case "gp", "x3":
regs.Gp = value
case "tp", "x4":
regs.Tp = value
case "t0", "x5":
regs.T0 = value
case "t1", "x6":
regs.T1 = value
case "t2", "x7":
regs.T2 = value
case "s0", "fp", "x8":
regs.S0 = value
case "s1", "x9":
regs.S1 = value
case "a0", "x10":
regs.A0 = value
case "a1", "x11":
regs.A1 = value
case "a2", "x12":
regs.A2 = value
case "a3", "x13":
regs.A3 = value
case "a4", "x14":
regs.A4 = value
case "a5", "x15":
regs.A5 = value
case "a6", "x16":
regs.A6 = value
case "a7", "x17":
regs.A7 = value
case "s2", "x18":
regs.S2 = value
case "s3", "x19":
regs.S3 = value
case "s4", "x20":
regs.S4 = value
case "s5", "x21":
regs.S5 = value
case "s6", "x22":
regs.S6 = value
case "s7", "x23":
regs.S7 = value
case "s8", "x24":
regs.S8 = value
case "s9", "x25":
regs.S9 = value
case "s10", "x26":
regs.S10 = value
case "s11", "x27":
regs.S11 = value
case "t3", "x28":
regs.T3 = value
case "t4", "x29":
regs.T4 = value
case "t5", "x30":
regs.T5 = value
case "t6", "x31":
regs.T6 = value
default:
return fmt.Errorf("debug: unknown register %q", name)
}
return s.SetRegs(&regs)
}
// archReturnAddr reads the return address from RA (riscv64 convention).
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
return regs.Ra, nil
}
// archSPLabel returns the SP register name for display.
func archSPLabel() string { return "SP" }
+2 -2
View File
@@ -17,8 +17,8 @@ import (
"time"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// Integration tests beyond the basic entry breakpoint: hardware watchpoints,
+288
View File
@@ -0,0 +1,288 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && (amd64 || arm64 || riscv64)
package debug
import (
"fmt"
"os"
"os/exec"
"path/filepath"
"runtime"
"strings"
"syscall"
"time"
"golang.org/x/sys/unix"
)
// Session is a ptrace debugging session controlling one debuggee process.
// The FreeBSD implementation sits behind the same surface as the Linux one:
// PT_TRACE_ME from the debuggee, PT_CONTINUE/PT_STEP from the tracer, and
// tracee memory through PT_IO (FreeBSD has no /proc/pid/mem to fall back
// on, so PT_IO is the only supported route).
type Session struct {
pid int
cmd *exec.Cmd
stopped bool
exited bool
codeBase uint64 // base address of the JIT code in the debuggee
tmpDir string // scratch directory of the session, removed on Kill
wpSlots [16]bool // hardware watchpoint slots in use (DR0-DR3, arm64 dbw 0-15)
// lastSignal holds the signal of the most recent stop when that stop
// was a genuine signal-delivery-stop the caller must see (a fault such
// as SIGSEGV, SIGBUS, SIGFPE or SIGILL); 0 for breakpoint traps,
// single-steps, SIGSTOP and suppressed runtime signals.
lastSignal syscall.Signal
}
// Launch starts the debuggee subprocess (gasm debug --target ...) and
// attaches to it via ptrace.
func Launch(gasmBin, asmPath, funcName string, args []byte) (*Session, error) {
sess, _, err := LaunchWithBuffers(gasmBin, asmPath, funcName, args, "")
return sess, err
}
// LaunchWithBuffers is like Launch but also allocates buffers in the debuggee.
//
// It pins the calling goroutine to its OS thread and leaves it pinned: the
// debuggee's PT_TRACE_ME binds the tracer relation to the forking thread,
// and every ptrace request on the session must come from that same thread.
// All Session methods must therefore be called from the goroutine that
// launched the session (the REPL and coverage loops do exactly that).
func LaunchWithBuffers(gasmBin, asmPath, funcName string, args []byte, bufSpec string) (*Session, []uint64, error) {
runtime.LockOSThread() // ptrace requests must stay on the forking thread
self, err := os.Executable()
if err != nil {
return nil, nil, fmt.Errorf("debug: cannot find gasm binary: %w", err)
}
if gasmBin != "" {
self = gasmBin
}
tmpDir, err := os.MkdirTemp("", "gasm-debug-*")
if err != nil {
return nil, nil, fmt.Errorf("debug: tempdir: %w", err)
}
argsFile := filepath.Join(tmpDir, "args.bin")
if err := os.WriteFile(argsFile, args, 0o644); err != nil {
os.RemoveAll(tmpDir)
return nil, nil, fmt.Errorf("debug: write args: %w", err)
}
if bufSpec != "" {
if err := os.WriteFile(filepath.Join(tmpDir, "bufspec"), []byte(bufSpec), 0o644); err != nil {
os.RemoveAll(tmpDir)
return nil, nil, fmt.Errorf("debug: write bufspec: %w", err)
}
}
cmd := exec.Command(self, "debug", "--func", funcName, "--args", argsFile, asmPath)
cmd.Env = append(os.Environ(), "GASM_DEBUG_TARGET=1", "GASM_DEBUG_TMP="+tmpDir)
cmd.Stdout = nil
cmd.Stderr = os.Stderr
cmd.SysProcAttr = &syscall.SysProcAttr{}
if err := cmd.Start(); err != nil {
os.RemoveAll(tmpDir)
return nil, nil, fmt.Errorf("debug: start debuggee: %w", err)
}
s := &Session{pid: cmd.Process.Pid, cmd: cmd, tmpDir: tmpDir}
readyFile := filepath.Join(tmpDir, "ready")
for range 500 {
if _, err := os.Stat(readyFile); err == nil {
break
}
time.Sleep(5 * time.Millisecond)
}
// The debuggee parks itself with SIGSTOP once the JIT code is mapped.
// A Go tracee also reports SIGURG preemption as signal-delivery-stops,
// so the wait loops until a stop the debugger cares about instead of
// assuming the first event is the SIGSTOP.
if _, err := s.waitStopped(); err != nil {
cmd.Process.Kill()
os.RemoveAll(tmpDir)
return nil, nil, fmt.Errorf("debug: wait for debuggee: %w", err)
}
s.stopped = true
// The debuggee reports its JIT mapping in the codebase file; that is
// the supported path on FreeBSD, where no /proc/pid/maps exists to
// scan for the RWX region as a fallback.
if data, err := os.ReadFile(filepath.Join(tmpDir, "codebase")); err == nil {
fmt.Sscanf(string(data), "%d", &s.codeBase)
}
var bufAddrs []uint64
if bufSpec != "" {
addrFile := filepath.Join(tmpDir, "bufaddrs")
if data, err := os.ReadFile(addrFile); err == nil {
for line := range strings.SplitSeq(strings.TrimSpace(string(data)), "\n") {
var addr uint64
if _, err := fmt.Sscanf(line, "%d", &addr); err == nil {
bufAddrs = append(bufAddrs, addr)
}
}
}
}
return s, bufAddrs, nil
}
// waitStopped consumes ptrace-stop events until one the debugger cares
// about arrives: SIGTRAP (a breakpoint or a completed single-step), the
// debuggee's own SIGSTOP, or a genuine signal-delivery-stop. A Go tracee's
// runtime raises SIGURG for asynchronous preemption, and every signal on a
// traced thread surfaces as a signal-delivery-stop, so SIGURG is suppressed
// and the tracee resumed without it. Every other signal (SIGSEGV, SIGBUS,
// SIGFPE, SIGILL, ...) is returned to the caller: resuming with signal 0
// would restart the faulting instruction and fault forever, so a faulting
// kernel must surface as a stop the caller reports.
func (s *Session) waitStopped() (syscall.Signal, error) {
for {
var ws syscall.WaitStatus
if _, err := syscall.Wait4(s.pid, &ws, syscall.WUNTRACED, nil); err != nil {
return 0, err
}
if ws.Exited() {
s.exited = true
return 0, fmt.Errorf("debuggee exited with status %d", ws.ExitStatus())
}
if ws.Signaled() {
s.exited = true
return 0, fmt.Errorf("debuggee killed by signal %v", ws.Signal())
}
switch sig := ws.StopSignal(); sig {
case syscall.SIGTRAP, syscall.SIGSTOP:
s.stopped = true
s.lastSignal = 0
return sig, nil
case syscall.SIGURG:
// Go runtime asynchronous preemption: resume the tracee
// without delivering the signal.
s.lastSignal = 0
if err := unix.PtraceCont(s.pid, 0); err != nil {
return 0, fmt.Errorf("debug: PT_CONTINUE: %w", err)
}
default:
// A genuine signal-delivery-stop. Report it; the caller
// decides how to proceed.
s.stopped = true
s.lastSignal = sig
return sig, nil
}
}
}
// LastSignal returns the signal of the most recent stop when that stop was
// a genuine signal-delivery-stop (a fault such as SIGSEGV, SIGFPE, SIGILL
// or SIGBUS), and 0 for breakpoint traps, single-steps, SIGSTOP and
// suppressed runtime signals.
func (s *Session) LastSignal() syscall.Signal { return s.lastSignal }
// Peek reads a word (8 bytes) from the debuggee's memory at addr, through
// PT_IO with PIOD_READ_D.
func (s *Session) Peek(addr uint64) (uint64, error) {
var buf [8]byte
if _, err := unix.PtraceIO(unix.PIOD_READ_D, s.pid, uintptr(addr), buf[:], len(buf)); err != nil {
return 0, fmt.Errorf("debug: read mem %#x: %w", addr, err)
}
return uint64(buf[0]) | uint64(buf[1])<<8 | uint64(buf[2])<<16 | uint64(buf[3])<<24 |
uint64(buf[4])<<32 | uint64(buf[5])<<40 | uint64(buf[6])<<48 | uint64(buf[7])<<56, nil
}
// Poke writes a word (8 bytes) to the debuggee's memory at addr, through
// PT_IO with PIOD_WRITE_D.
func (s *Session) Poke(addr, val uint64) error {
buf := []byte{byte(val), byte(val >> 8), byte(val >> 16), byte(val >> 24),
byte(val >> 32), byte(val >> 40), byte(val >> 48), byte(val >> 56)}
if _, err := unix.PtraceIO(unix.PIOD_WRITE_D, s.pid, uintptr(addr), buf, len(buf)); err != nil {
return fmt.Errorf("debug: write mem %#x: %w", addr, err)
}
return nil
}
// ReadMemory reads len bytes from the debuggee's memory at addr in one
// PT_IO request, the shape the request is built for.
func (s *Session) ReadMemory(addr uint64, length int) ([]byte, error) {
out := make([]byte, length)
n, err := unix.PtraceIO(unix.PIOD_READ_D, s.pid, uintptr(addr), out, length)
return out[:n], err
}
// WriteMemory writes bytes to the debuggee's memory at addr in one PT_IO
// request.
func (s *Session) WriteMemory(addr uint64, data []byte) error {
_, err := unix.PtraceIO(unix.PIOD_WRITE_D, s.pid, uintptr(addr), data, len(data))
return err
}
// Step executes a single instruction in the debuggee.
func (s *Session) Step() error {
if s.exited {
return fmt.Errorf("debug: debuggee has exited")
}
if err := unix.PtraceSingleStep(s.pid); err != nil {
return fmt.Errorf("debug: PT_STEP: %w", err)
}
_, err := s.waitStopped()
return err
}
// Continue resumes execution until the next breakpoint or exit.
func (s *Session) Continue() error {
if s.exited {
return fmt.Errorf("debug: debuggee has exited")
}
if err := unix.PtraceCont(s.pid, 0); err != nil {
return fmt.Errorf("debug: PT_CONTINUE: %w", err)
}
_, err := s.waitStopped()
return err
}
// Exited returns true if the debuggee has terminated.
func (s *Session) Exited() bool { return s.exited }
// Pid returns the debuggee's process ID.
func (s *Session) Pid() int { return s.pid }
// CodeBase returns the base address of the JIT code in the debuggee.
func (s *Session) CodeBase() uint64 { return s.codeBase }
// Kill terminates the debuggee and removes the session's scratch
// directory, so a successful session leaves no gasm-debug-* debris behind.
func (s *Session) Kill() {
if !s.exited {
syscall.Kill(s.pid, syscall.SIGKILL)
syscall.Wait4(s.pid, nil, 0, nil)
s.exited = true
}
if s.cmd != nil && s.cmd.Process != nil {
s.cmd.Wait()
}
if s.tmpDir != "" {
os.RemoveAll(s.tmpDir)
s.tmpDir = ""
}
}
// execRange is one executable mapping of the debuggee.
type execRange struct {
lo, hi uint64
}
// execRanges is a stub on FreeBSD: there is no /proc/pid/maps to parse,
// and procfs(5) is not guaranteed to be mounted. The callers degrade
// gracefully: archReturnAddr falls back to the raw stack convention and
// the mapping scan is skipped.
func execRanges(pid int) []execRange { return nil }
// findRWXMapping is a stub on FreeBSD for the same reason: the codebase
// handshake file is the supported way the JIT region is located.
func findRWXMapping(pid int) uint64 { return 0 }
+160
View File
@@ -0,0 +1,160 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
import (
"encoding/binary"
"fmt"
"unsafe"
"golang.org/x/sys/unix"
)
// GetRegs reads the general-purpose registers of the stopped debuggee and
// converts the FreeBSD struct reg into the portable layout.
func (s *Session) GetRegs() (Regs, error) {
var ur unix.Reg
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
}
return Regs{
R15: uint64(ur.R15),
R14: uint64(ur.R14),
R13: uint64(ur.R13),
R12: uint64(ur.R12),
R11: uint64(ur.R11),
R10: uint64(ur.R10),
R9: uint64(ur.R9),
R8: uint64(ur.R8),
RDI: uint64(ur.Rdi),
RSI: uint64(ur.Rsi),
RBP: uint64(ur.Rbp),
RBX: uint64(ur.Rbx),
RDX: uint64(ur.Rdx),
RCX: uint64(ur.Rcx),
RAX: uint64(ur.Rax),
RIP: uint64(ur.Rip),
CS: uint64(ur.Cs),
RFLAGS: uint64(ur.Rflags),
RSP: uint64(ur.Rsp),
SS: uint64(ur.Ss),
FS: uint64(ur.Fs),
GS: uint64(ur.Gs),
DS: uint64(ur.Ds),
ES: uint64(ur.Es),
}, nil
}
// SetRegs writes the general-purpose registers of the stopped debuggee.
func (s *Session) SetRegs(regs *Regs) error {
// Read-modify-write keeps the fields FreeBSD owns (trapno, err) intact.
var ur unix.Reg
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
return fmt.Errorf("debug: PT_GETREGS: %w", err)
}
ur.R15 = int64(regs.R15)
ur.R14 = int64(regs.R14)
ur.R13 = int64(regs.R13)
ur.R12 = int64(regs.R12)
ur.R11 = int64(regs.R11)
ur.R10 = int64(regs.R10)
ur.R9 = int64(regs.R9)
ur.R8 = int64(regs.R8)
ur.Rdi = int64(regs.RDI)
ur.Rsi = int64(regs.RSI)
ur.Rbp = int64(regs.RBP)
ur.Rbx = int64(regs.RBX)
ur.Rdx = int64(regs.RDX)
ur.Rcx = int64(regs.RCX)
ur.Rax = int64(regs.RAX)
ur.Rip = int64(regs.RIP)
ur.Cs = int64(regs.CS)
ur.Rflags = int64(regs.RFLAGS)
ur.Rsp = int64(regs.RSP)
ur.Ss = int64(regs.SS)
return unix.PtraceSetRegs(s.pid, &ur)
}
// FPRegs holds the x87 FPU and SSE (XMM) register state, the FXSAVE image
// the FreeBSD struct fpreg mirrors: XMM0-15 at the same offsets.
type FPRegs struct {
XMM [16][16]byte // XMM0-15
}
// GetFPRegs retrieves the FPU/SSE register state via PT_GETFPREGS. The
// FreeBSD struct fpreg mirrors the FXSAVE image: the x87 environment and
// stack in Env/Acc, XMM0-15 in Xacc.
func (s *Session) GetFPRegs() (FPRegs, error) {
var fp FPRegs
var fr unix.FpReg
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
}
for i := range 16 {
copy(fp.XMM[i][:], fr.Xacc[i][:])
}
return fp, nil
}
// VectorRegs holds the YMM register state.
type VectorRegs struct {
YMM [16][32]byte // YMM0-15 (full 256-bit values)
}
// The XSAVE area the PT_GETXSTATE request returns follows the architectural
// layout (Intel SDM vol 1, "XSAVE"): the 512-byte legacy FXSAVE image (x87
// state in 0-159, XMM0-15 in 160-511), then the 64-byte xsave header whose
// first 8 bytes are xstate_bv, then one component per set feature bit, each
// 64-byte aligned. The YMM high halves are the first extended component,
// at offset 576; XFEATURE_STATE_BIT_AVX is bit 2 of xstate_bv.
const (
xsaveXMMOffset = 160
xsaveHeaderOffset = 512
xsaveBVOffset = xsaveHeaderOffset
ymmOffset = xsaveHeaderOffset + 64 // 576
ymmSize = 256 // 16 registers, 16 bytes each
xfeatureMaskYMM = 1 << 2
xstateMaxBuffer = 4096 // PT_GETXSTATE_INFO bounds the size far below this
)
// GetVectorRegs retrieves the YMM registers via PT_GETXSTATE. The low
// (XMM) halves always come from the legacy image; the high halves are
// copied only when xstate_bv reports the AVX state, and read as zero
// otherwise. When the request fails the FP image still provides correct
// XMM halves, so that is the fallback.
func (s *Session) GetVectorRegs() (VectorRegs, error) {
var v VectorRegs
buf := make([]byte, xstateMaxBuffer)
n, _, errno := unix.Syscall6(
unix.SYS_PTRACE,
uintptr(unix.PT_GETXSTATE),
uintptr(s.pid),
0,
uintptr(unsafe.Pointer(&buf[0])),
0, 0,
)
if errno != 0 {
fp, err := s.GetFPRegs()
if err != nil {
return v, err
}
for i := range 16 {
copy(v.YMM[i][:16], fp.XMM[i][:])
}
return v, nil
}
for i := range 16 {
copy(v.YMM[i][:16], buf[xsaveXMMOffset+16*i:xsaveXMMOffset+16*i+16])
}
if int(n) >= ymmOffset+ymmSize {
if binary.LittleEndian.Uint64(buf[xsaveBVOffset:xsaveBVOffset+8])&xfeatureMaskYMM != 0 {
for i := range 16 {
copy(v.YMM[i][16:], buf[ymmOffset+16*i:ymmOffset+16*i+16])
}
}
}
return v, nil
}
+114
View File
@@ -0,0 +1,114 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && arm64
package debug
import (
"fmt"
"golang.org/x/sys/unix"
)
// GetRegs reads the general-purpose registers of the stopped debuggee and
// converts the FreeBSD struct reg (x[30], lr, sp, elr, spsr) into the
// portable layout.
func (s *Session) GetRegs() (Regs, error) {
var ur unix.Reg
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
}
return Regs{
X0: ur.X[0],
X1: ur.X[1],
X2: ur.X[2],
X3: ur.X[3],
X4: ur.X[4],
X5: ur.X[5],
X6: ur.X[6],
X7: ur.X[7],
X8: ur.X[8],
X9: ur.X[9],
X10: ur.X[10],
X11: ur.X[11],
X12: ur.X[12],
X13: ur.X[13],
X14: ur.X[14],
X15: ur.X[15],
X16: ur.X[16],
X17: ur.X[17],
X18: ur.X[18],
X19: ur.X[19],
X20: ur.X[20],
X21: ur.X[21],
X22: ur.X[22],
X23: ur.X[23],
X24: ur.X[24],
X25: ur.X[25],
X26: ur.X[26],
X27: ur.X[27],
X28: ur.X[28],
X29: ur.X[29],
X30: ur.Lr,
SP: ur.Sp,
PC: ur.Elr,
PSTATE: uint64(ur.Spsr),
}, nil
}
// SetRegs writes the general-purpose registers of the stopped debuggee.
func (s *Session) SetRegs(regs *Regs) error {
var ur unix.Reg
ur.X = [30]uint64{
regs.X0, regs.X1, regs.X2, regs.X3, regs.X4, regs.X5, regs.X6,
regs.X7, regs.X8, regs.X9, regs.X10, regs.X11, regs.X12, regs.X13,
regs.X14, regs.X15, regs.X16, regs.X17, regs.X18, regs.X19, regs.X20,
regs.X21, regs.X22, regs.X23, regs.X24, regs.X25, regs.X26, regs.X27,
regs.X28, regs.X29,
}
ur.Lr = regs.X30
ur.Sp = regs.SP
ur.Elr = regs.PC
ur.Spsr = uint32(regs.PSTATE)
return unix.PtraceSetRegs(s.pid, &ur)
}
// FPRegs holds the arm64 FP/NEON register state: the 32 128-bit V
// registers, then FPSR and FPCR (the user_fpsimd shape).
type FPRegs struct {
V [32][16]byte // V0-V31 (128-bit NEON/FP registers)
FPSR uint32
FPCR uint32
}
// GetFPRegs retrieves the FP/NEON register state via PT_GETFPREGS. The
// FreeBSD struct fpreg holds the 32 128-bit V registers followed by FPSR
// and FPCR, the user_fpsimd shape.
func (s *Session) GetFPRegs() (FPRegs, error) {
var fp FPRegs
var fr unix.FpReg
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
}
for i := range 32 {
copy(fp.V[i][:], fr.Q[i][:])
}
return fp, nil
}
// VectorRegs holds the full SIMD register state.
type VectorRegs struct {
V [32][16]byte // V0-V31 (128-bit)
}
// GetVectorRegs retrieves the SIMD registers.
func (s *Session) GetVectorRegs() (VectorRegs, error) {
var v VectorRegs
fp, err := s.GetFPRegs()
if err != nil {
return v, err
}
copy(v.V[:][:], fp.V[:][:])
return v, nil
}
+115
View File
@@ -0,0 +1,115 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && riscv64
package debug
import (
"fmt"
"golang.org/x/sys/unix"
)
// GetRegs reads the general-purpose registers of the stopped debuggee and
// converts the FreeBSD struct reg into the portable layout. Sstatus rides
// the kernel's struct but the portable surface carries the GPRs and PC.
func (s *Session) GetRegs() (Regs, error) {
var ur unix.Reg
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
}
return Regs{
PC: ur.Sepc,
Ra: ur.Ra,
Sp: ur.Sp,
Gp: ur.Gp,
Tp: ur.Tp,
T0: ur.T[0],
T1: ur.T[1],
T2: ur.T[2],
S0: ur.S[0],
S1: ur.S[1],
A0: ur.A[0],
A1: ur.A[1],
A2: ur.A[2],
A3: ur.A[3],
A4: ur.A[4],
A5: ur.A[5],
A6: ur.A[6],
A7: ur.A[7],
S2: ur.S[2],
S3: ur.S[3],
S4: ur.S[4],
S5: ur.S[5],
S6: ur.S[6],
S7: ur.S[7],
S8: ur.S[8],
S9: ur.S[9],
S10: ur.S[10],
S11: ur.S[11],
T3: ur.T[3],
T4: ur.T[4],
T5: ur.T[5],
T6: ur.T[6],
}, nil
}
// SetRegs writes the general-purpose registers of the stopped debuggee.
// Read-modify-write keeps sstatus, which the kernel owns, intact.
func (s *Session) SetRegs(regs *Regs) error {
var ur unix.Reg
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
return fmt.Errorf("debug: PT_GETREGS: %w", err)
}
ur.Sepc = regs.PC
ur.Ra = regs.Ra
ur.Sp = regs.Sp
ur.Gp = regs.Gp
ur.Tp = regs.Tp
ur.T = [7]uint64{regs.T0, regs.T1, regs.T2, regs.T3, regs.T4, regs.T5, regs.T6}
ur.S = [12]uint64{regs.S0, regs.S1, regs.S2, regs.S3, regs.S4, regs.S5,
regs.S6, regs.S7, regs.S8, regs.S9, regs.S10, regs.S11}
ur.A = [8]uint64{regs.A0, regs.A1, regs.A2, regs.A3, regs.A4, regs.A5, regs.A6, regs.A7}
return unix.PtraceSetRegs(s.pid, &ur)
}
// FPRegs holds the RISC-V FP register state (32 64-bit FP registers plus
// fcsr).
type FPRegs struct {
F [32]uint64 // F0-F31 (64-bit FP registers)
FCSR uint32
}
// GetFPRegs retrieves the FP register state via PT_GETFPREGS. The FreeBSD
// struct fpreg carries each 64-bit FP register in a 128-bit slot (fp_x is
// the flat [64]-word area the x/sys type renders as [32][2]); the low word
// holds the register, and FCSR rides the tail.
func (s *Session) GetFPRegs() (FPRegs, error) {
var fp FPRegs
var fr unix.FpReg
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
}
for i := range 32 {
fp.F[i] = fr.X[i][0]
}
fp.FCSR = uint32(fr.Fcsr)
return fp, nil
}
// VectorRegs holds the FP register state shown by the regs command
// (riscv64 has 32 64-bit FP registers and fcsr).
type VectorRegs struct {
F [32]uint64
FCSR uint32
}
// GetVectorRegs retrieves the FP registers.
func (s *Session) GetVectorRegs() (VectorRegs, error) {
fp, err := s.GetFPRegs()
if err != nil {
return VectorRegs{}, err
}
return VectorRegs{F: fp.F, FCSR: fp.FCSR}, nil
}
+91
View File
@@ -0,0 +1,91 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
import (
"os/exec"
"path/filepath"
"runtime"
"testing"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// TestLaunchAndBreakpoint is the FreeBSD twin of the Linux integration
// test: it drives the whole launch, breakpoint, trap and register-rewind
// flow end to end. It needs a real FreeBSD kernel (ptrace does not work
// under emulation), so it only runs where it can.
func TestLaunchAndBreakpoint(t *testing.T) {
if runtime.GOARCH != "amd64" {
t.Skip("runs only on amd64 hosts")
}
// The tracer is the OS thread that forked the debuggee (PT_TRACE_ME
// binds the relation to that thread); every ptrace request must come
// from the same thread, so pin the test goroutine to one thread.
runtime.LockOSThread()
defer runtime.UnlockOSThread()
bin := filepath.Join(t.TempDir(), "gasm")
out, err := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm").CombinedOutput()
if err != nil {
t.Fatalf("build gasm: %v: %s", err, out)
}
const kernelPath = "../testdata/verify/basic_amd64.s"
k, err := verify.Load(kernelPath)
if err != nil {
t.Fatalf("Load: %v", err)
}
t.Cleanup(k.Close)
fl, err := k.Func("wideCopy")
if err != nil {
t.Fatalf("Func: %v", err)
}
sess, err := Launch(bin, kernelPath, "wideCopy", make([]byte, fl.Args))
if err != nil {
t.Fatalf("Launch: %v", err)
}
t.Cleanup(sess.Kill)
bm := NewBreakpoints(sess)
entry := sess.CodeBase() + uint64(fl.Offset)
if _, err := bm.Set(entry, "entry"); err != nil {
t.Fatalf("Set: %v", err)
}
// The INT3 must be visible in the debuggee's memory.
word, err := sess.Peek(entry)
if err != nil {
t.Fatalf("Peek: %v", err)
}
if b := word & 0xFF; b != 0xCC {
t.Fatalf("int3 not patched: first byte %#02x at %#x", b, entry)
}
// The debuggee raises a second SIGSTOP after the launch barrier (the
// child's RunTarget marks its entry), so like the REPL and the cover
// mode the test keeps resuming until the breakpoint trap arrives.
for range 10 {
if err := sess.Continue(); err != nil {
t.Fatalf("Continue: %v", err)
}
if sess.Exited() {
t.Fatal("debuggee exited instead of trapping on the breakpoint")
}
regs, err := sess.GetRegs()
if err != nil {
t.Fatalf("GetRegs: %v", err)
}
if bp := bm.HandleTrap(&regs); bp != nil {
if bp.Addr != entry {
t.Fatalf("trap at %#x, want %#x", bp.Addr, entry)
}
return // trap on the entry breakpoint: the whole flow works
}
}
t.Fatal("no breakpoint trap after 10 resumes")
}
+2 -2
View File
@@ -14,7 +14,7 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// buildGasm produces the gasm binary the debugger spawns as its debuggee.
@@ -24,7 +24,7 @@ func buildGasm(t *testing.T) string {
return p
}
bin := filepath.Join(t.TempDir(), "gasm")
cmd := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm")
cmd := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm")
out, err := cmd.CombinedOutput()
if err != nil {
t.Fatalf("build gasm: %v: %s", err, out)
+98
View File
@@ -0,0 +1,98 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
// Regs holds the full general-purpose register set of a traced process
// (the FreeBSD amd64 struct reg layout, sys/x86/include/reg.h). FreeBSD
// reports segment selectors (FS/GS/ES/DS), not the bases the Linux ptrace
// surface carries, and has no ORIG_RAX slot.
type Regs struct {
R15 uint64
R14 uint64
R13 uint64
R12 uint64
RBP uint64
RBX uint64
R11 uint64
R10 uint64
R9 uint64
R8 uint64
RAX uint64
RCX uint64
RDX uint64
RSI uint64
RDI uint64
RIP uint64
CS uint64
RFLAGS uint64
RSP uint64
SS uint64
FS uint64
GS uint64
DS uint64
ES uint64
}
// GetPC returns the program counter.
func (r *Regs) GetPC() uint64 { return r.RIP }
// SetPC sets the program counter.
func (r *Regs) SetPC(pc uint64) { r.RIP = pc }
// GetSP returns the stack pointer.
func (r *Regs) GetSP() uint64 { return r.RSP }
// RegValue returns the value of the named register, or false if unknown.
func (r *Regs) RegValue(name string) (uint64, bool) {
switch name {
case "rax", "eax", "ax", "al":
return r.RAX, true
case "rbx", "ebx", "bx", "bl":
return r.RBX, true
case "rcx", "ecx", "cx", "cl":
return r.RCX, true
case "rdx", "edx", "dx", "dl":
return r.RDX, true
case "rsi", "esi", "si":
return r.RSI, true
case "rdi", "edi", "di":
return r.RDI, true
case "rbp", "ebp", "bp":
return r.RBP, true
case "rsp", "esp", "sp":
return r.RSP, true
case "r8":
return r.R8, true
case "r9":
return r.R9, true
case "r10":
return r.R10, true
case "r11":
return r.R11, true
case "r12":
return r.R12, true
case "r13":
return r.R13, true
case "r14":
return r.R14, true
case "r15":
return r.R15, true
case "rip", "eip":
return r.RIP, true
default:
return 0, false
}
}
// breakpointInsn is the software breakpoint instruction.
var breakpointInsn = []byte{0xCC} // INT3
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
// a trap. INT3 leaves the hardware PC on the following instruction (Intel
// SDM vol 3, "Debug Exceptions") and the FreeBSD T_BPTFLT path delivers
// that frame unmodified (sys/amd64/amd64/trap.c), so the trap address is
// PC-1, the same correction the Linux side applies.
const breakpointPCAdjust = 1
+139
View File
@@ -0,0 +1,139 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && arm64
package debug
// Regs holds the full general-purpose register set of a traced process
// (the FreeBSD arm64 struct reg layout, sys/arm64/include/reg.h: x[30], lr,
// sp, elr, spsr).
type Regs struct {
X0 uint64
X1 uint64
X2 uint64
X3 uint64
X4 uint64
X5 uint64
X6 uint64
X7 uint64
X8 uint64
X9 uint64
X10 uint64
X11 uint64
X12 uint64
X13 uint64
X14 uint64
X15 uint64
X16 uint64
X17 uint64
X18 uint64
X19 uint64
X20 uint64
X21 uint64
X22 uint64
X23 uint64
X24 uint64
X25 uint64
X26 uint64
X27 uint64
X28 uint64
X29 uint64 // FP (frame pointer)
X30 uint64 // LR (link register)
SP uint64
PC uint64
PSTATE uint64
}
// GetPC returns the program counter.
func (r *Regs) GetPC() uint64 { return r.PC }
// SetPC sets the program counter.
func (r *Regs) SetPC(pc uint64) { r.PC = pc }
// GetSP returns the stack pointer.
func (r *Regs) GetSP() uint64 { return r.SP }
// RegValue returns the value of the named register, or false if unknown.
func (r *Regs) RegValue(name string) (uint64, bool) {
switch name {
case "x0":
return r.X0, true
case "x1":
return r.X1, true
case "x2":
return r.X2, true
case "x3":
return r.X3, true
case "x4":
return r.X4, true
case "x5":
return r.X5, true
case "x6":
return r.X6, true
case "x7":
return r.X7, true
case "x8":
return r.X8, true
case "x9":
return r.X9, true
case "x10":
return r.X10, true
case "x11":
return r.X11, true
case "x12":
return r.X12, true
case "x13":
return r.X13, true
case "x14":
return r.X14, true
case "x15":
return r.X15, true
case "x16":
return r.X16, true
case "x17":
return r.X17, true
case "x18":
return r.X18, true
case "x19":
return r.X19, true
case "x20":
return r.X20, true
case "x21":
return r.X21, true
case "x22":
return r.X22, true
case "x23":
return r.X23, true
case "x24":
return r.X24, true
case "x25":
return r.X25, true
case "x26":
return r.X26, true
case "x27":
return r.X27, true
case "x28":
return r.X28, true
case "x29", "fp":
return r.X29, true
case "x30", "lr":
return r.X30, true
case "sp":
return r.SP, true
case "pc":
return r.PC, true
default:
return 0, false
}
}
// breakpointInsn is the software breakpoint instruction (BRK #0).
var breakpointInsn = []byte{0x00, 0x00, 0x20, 0xD4} // BRK #0
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
// a trap: 0. The BRK synchronous exception leaves ELR_EL0 on the BRK
// itself (ARM DDI 0487), and the FreeBSD EXCP_BRKPT_EL0 handler delivers
// the frame's elr unmodified (sys/arm64/arm64/trap.c), so the trap address
// is the PC as reported.
const breakpointPCAdjust = 0
+134
View File
@@ -0,0 +1,134 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && riscv64
package debug
// Regs holds the full general-purpose register set of a traced process
// (the FreeBSD riscv64 struct reg layout: ra, sp, gp, tp, t0-t6, s0-s11,
// a0-a7, sepc, sstatus).
type Regs struct {
PC uint64 // sepc
Ra uint64 // x1 (return address)
Sp uint64 // x2
Gp uint64 // x3
Tp uint64 // x4
T0 uint64 // x5
T1 uint64 // x6
T2 uint64 // x7
S0 uint64 // x8 (frame pointer)
S1 uint64 // x9
A0 uint64 // x10
A1 uint64 // x11
A2 uint64 // x12
A3 uint64 // x13
A4 uint64 // x14
A5 uint64 // x15
A6 uint64 // x16
A7 uint64 // x17
S2 uint64 // x18
S3 uint64 // x19
S4 uint64 // x20
S5 uint64 // x21
S6 uint64 // x22
S7 uint64 // x23
S8 uint64 // x24
S9 uint64 // x25
S10 uint64 // x26
S11 uint64 // x27
T3 uint64 // x28
T4 uint64 // x29
T5 uint64 // x30
T6 uint64 // x31
}
// GetPC returns the program counter.
func (r *Regs) GetPC() uint64 { return r.PC }
// SetPC sets the program counter.
func (r *Regs) SetPC(pc uint64) { r.PC = pc }
// GetSP returns the stack pointer.
func (r *Regs) GetSP() uint64 { return r.Sp }
// RegValue returns the value of the named register, or false if unknown.
func (r *Regs) RegValue(name string) (uint64, bool) {
switch name {
case "pc":
return r.PC, true
case "ra", "x1":
return r.Ra, true
case "sp", "x2":
return r.Sp, true
case "gp", "x3":
return r.Gp, true
case "tp", "x4":
return r.Tp, true
case "t0", "x5":
return r.T0, true
case "t1", "x6":
return r.T1, true
case "t2", "x7":
return r.T2, true
case "s0", "fp", "x8":
return r.S0, true
case "s1", "x9":
return r.S1, true
case "a0", "x10":
return r.A0, true
case "a1", "x11":
return r.A1, true
case "a2", "x12":
return r.A2, true
case "a3", "x13":
return r.A3, true
case "a4", "x14":
return r.A4, true
case "a5", "x15":
return r.A5, true
case "a6", "x16":
return r.A6, true
case "a7", "x17":
return r.A7, true
case "s2", "x18":
return r.S2, true
case "s3", "x19":
return r.S3, true
case "s4", "x20":
return r.S4, true
case "s5", "x21":
return r.S5, true
case "s6", "x22":
return r.S6, true
case "s7", "x23":
return r.S7, true
case "s8", "x24":
return r.S8, true
case "s9", "x25":
return r.S9, true
case "s10", "x26":
return r.S10, true
case "s11", "x27":
return r.S11, true
case "t3", "x28":
return r.T3, true
case "t4", "x29":
return r.T4, true
case "t5", "x30":
return r.T5, true
case "t6", "x31":
return r.T6, true
default:
return 0, false
}
}
// breakpointInsn is the software breakpoint instruction (EBREAK).
var breakpointInsn = []byte{0x73, 0x00, 0x10, 0x00} // ebreak
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
// a trap: 0. The EBREAK synchronous exception leaves sepc on the ebreak
// itself (RISC-V privileged architecture), so the trap address is the PC as
// reported.
const breakpointPCAdjust = 0
+1 -1
View File
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
package debug
+73
View File
@@ -0,0 +1,73 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && (amd64 || arm64 || riscv64)
package debug
import (
"encoding/binary"
"syscall"
"unsafe"
"golang.org/x/sys/unix"
)
// FreeBSD TRAP_* si_code values (sys/signal.h). A breakpoint (INT3, BRK,
// EBREAK) arrives as TRAP_BRKPT on every supported architecture; TRAP_TRACE
// is shared by the completed single-step and the hardware watchpoint hit,
// so the watchpoint layer disambiguates from the debug registers.
const (
trapBRKPT = 1 // TRAP_BRKPT
trapTRACE = 2 // TRAP_TRACE
)
// StopReason describes why the debuggee stopped.
type StopReason int
const (
StopNone StopReason = iota
StopBreakpoint // software breakpoint hit
StopWatchpoint // hardware watchpoint triggered
StopSingleStep // single-step completed
StopSignal // stopped by a signal
StopExited // process exited
)
// StopInfo returns the reason the debuggee stopped and the faulting address
// (for watchpoints, the watched address that was accessed). FreeBSD has no
// PTRACE_GETSIGINFO; the stop's signal information comes from PT_LWPINFO,
// whose pl_siginfo carries the siginfo the kernel delivered. A ptrace stop
// with no signal behind it (a completed single-step, the initial attach)
// fills no siginfo at all.
func (s *Session) StopInfo() (StopReason, uint64) {
if s.exited {
return StopExited, 0
}
var info unix.PtraceLwpInfoStruct
if err := unix.PtraceLwpInfo(s.pid, &info); err != nil {
return StopNone, 0
}
// The siginfo layout is the FreeBSD siginfo_t: three leading ints
// (signo, errno, code), then the union, 8-byte aligned, whose _fault
// member puts the address at byte offset 16. The read is byte-wise
// because the blob's alignment is not guaranteed.
si := (*[64]byte)(unsafe.Pointer(&info.Siginfo))
signo := int32(binary.LittleEndian.Uint32(si[0:4]))
code := int32(binary.LittleEndian.Uint32(si[8:12]))
switch {
case signo == 0:
// A pure ptrace stop: single-step completion, attach, or the
// events the kernel resolves internally.
return StopSingleStep, 0
case signo != int32(syscall.SIGTRAP):
return StopSignal, uint64(code)
case code == trapBRKPT:
return StopBreakpoint, 0
case code == trapTRACE:
addr := binary.LittleEndian.Uint64(si[16:24])
return archStopTrace(s, addr)
default:
return StopSingleStep, 0
}
}
+195
View File
@@ -0,0 +1,195 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && (amd64 || arm64 || riscv64)
package debug
import (
"encoding/hex"
"fmt"
"os"
"runtime"
"strconv"
"strings"
"syscall"
"unsafe"
"golang.org/x/sys/unix"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// mapRWX maps code into a read-write-execute region.
func mapRWX(code []byte) ([]byte, error) {
const pageSize = 4096
size := (len(code) + pageSize - 1) &^ (pageSize - 1)
mem, err := syscall.Mmap(-1, 0, size,
syscall.PROT_READ|syscall.PROT_WRITE|syscall.PROT_EXEC,
syscall.MAP_PRIVATE|syscall.MAP_ANON)
if err != nil {
return nil, err
}
copy(mem, code)
return mem, nil
}
// setupBuffers allocates buffers in the debuggee's memory.
func setupBuffers(spec string, args []byte, tmpDir string) ([]byte, error) {
type bufSpec struct {
name string
size int
pattern string
}
var specs []bufSpec
for part := range strings.SplitSeq(spec, ",") {
fields := strings.SplitN(part, ":", 3)
if len(fields) != 3 {
continue
}
size, err := strconv.Atoi(fields[1])
if err != nil || size <= 0 {
continue
}
specs = append(specs, bufSpec{name: fields[0], size: size, pattern: fields[2]})
}
if len(specs) == 0 {
return args, nil
}
var bufAddrs []uint64
for _, s := range specs {
buf, err := syscall.Mmap(-1, 0, s.size,
syscall.PROT_READ|syscall.PROT_WRITE,
syscall.MAP_PRIVATE|syscall.MAP_ANON)
if err != nil {
return nil, fmt.Errorf("mmap buffer %s: %w", s.name, err)
}
fillBuffer(buf, s.pattern)
bufAddrs = append(bufAddrs, uint64(uintptr(unsafe.Pointer(&buf[0]))))
}
addrFile, err := os.Create(tmpDir + "/bufaddrs")
if err != nil {
return nil, err
}
for _, addr := range bufAddrs {
fmt.Fprintf(addrFile, "%d\n", addr)
}
addrFile.Close()
return args, nil
}
// fillBuffer fills a buffer with the specified pattern.
func fillBuffer(buf []byte, pattern string) {
switch pattern {
case "zero":
case "ones":
for i := range buf {
buf[i] = 0xFF
}
case "seq":
for i := range buf {
buf[i] = byte(i)
}
default:
if data, err := hex.DecodeString(pattern); err == nil && len(data) > 0 {
for i := range buf {
buf[i] = data[i%len(data)]
}
}
}
}
// RunTarget is the debuggee entry point (gasm debug --target).
func RunTarget(asmPath, funcName, argsFile, tmpDir string) error {
src, err := os.ReadFile(asmPath)
if err != nil {
return fmt.Errorf("debug target: %w", err)
}
file, errs := parser.Parse(asmPath, string(src))
if len(errs) > 0 {
return fmt.Errorf("debug target: parse: %v", errs[0])
}
img, err := asm.AssembleFile(file)
if err != nil {
return fmt.Errorf("debug target: assemble: %w", err)
}
var fl *asm.FuncLayout
for i := range img.Funcs {
if img.Funcs[i].Name == funcName {
fl = &img.Funcs[i]
break
}
}
if fl == nil {
return fmt.Errorf("debug target: function %q not found", funcName)
}
code := img.Bytes()
exec, err := mapRWX(code)
if err != nil {
return fmt.Errorf("debug target: mmap: %w", err)
}
codeBase := uintptr(unsafe.Pointer(&exec[0]))
if err := os.WriteFile(tmpDir+"/codebase", []byte(fmt.Sprintf("%d", codeBase)), 0o644); err != nil {
return fmt.Errorf("debug target: write codebase: %w", err)
}
meta := fmt.Sprintf("%d %d %d", fl.Offset, fl.Size, fl.Args)
os.WriteFile(tmpDir+"/funcmeta", []byte(meta), 0o644)
labelsFile, _ := os.Create(tmpDir + "/labels")
if labelsFile != nil {
for label, off := range fl.Labels {
fmt.Fprintf(labelsFile, "%s %d\n", label, off)
}
labelsFile.Close()
}
args, err := os.ReadFile(argsFile)
if err != nil {
return fmt.Errorf("debug target: read args: %w", err)
}
if len(args) < fl.Args {
padded := make([]byte, fl.Args)
copy(padded, args)
args = padded
}
bufSpecFile := tmpDir + "/bufspec"
if bufSpec, err := os.ReadFile(bufSpecFile); err == nil && len(bufSpec) > 0 {
args, err = setupBuffers(string(bufSpec), args, tmpDir)
if err != nil {
return fmt.Errorf("debug target: setup buffers: %w", err)
}
}
runtime.LockOSThread()
if _, _, errno := unix.RawSyscall(unix.SYS_PTRACE, uintptr(unix.PT_TRACE_ME), 0, 0); errno != 0 {
return fmt.Errorf("debug target: PT_TRACE_ME: %v", errno)
}
os.WriteFile(tmpDir+"/ready", []byte("ok"), 0o644)
syscall.Kill(syscall.Getpid(), syscall.SIGSTOP)
os.WriteFile(tmpDir+"/entry", []byte("ok"), 0o644)
syscall.Kill(syscall.Getpid(), syscall.SIGSTOP)
fnAddr := codeBase + uintptr(fl.Offset)
stackArgs := make([]byte, fl.Args)
copy(stackArgs, args)
if _, callErr := verify.Call(fnAddr, stackArgs); callErr != nil {
os.Exit(1)
}
// Success returns to the caller, which exits with status 0; the JIT
// code has already run to its own trampoline by the time Call returns.
return nil
}
+3 -3
View File
@@ -12,9 +12,9 @@ import (
"syscall"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// RunTarget is the debuggee entry point (gasm debug --target).
+3 -3
View File
@@ -12,9 +12,9 @@ import (
"syscall"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// RunTarget is the debuggee entry point (gasm debug --target).
+3 -3
View File
@@ -12,9 +12,9 @@ import (
"syscall"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// RunTarget is the debuggee entry point (gasm debug --target).
+3 -3
View File
@@ -12,9 +12,9 @@ import (
"syscall"
"unsafe"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
)
// RunTarget is the debuggee entry point (gasm debug --target).
+1 -1
View File
@@ -1,7 +1,7 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build linux
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
package debug
+190
View File
@@ -0,0 +1,190 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && amd64
package debug
import (
"fmt"
"unsafe"
"golang.org/x/sys/unix"
)
// Hardware watchpoint support via x86-64 debug registers (DR0-DR3, DR7),
// read and written as one blob through PT_GETDBREGS/PT_SETDBREGS. The
// FreeBSD struct dbreg is the raw DR file: dr[16], where DR0-DR3 are the
// address registers, DR6 the status and DR7 the control (sys/x86/include/
// reg.h; the DBREG_DRX accessor indexes the same array).
// dbreg mirrors FreeBSD's struct dbreg for PT_GETDBREGS/PT_SETDBREGS.
type dbreg struct {
Dr [16]uint64
}
// dbreg indices of the registers the watchpoint layer drives.
const (
drStatus = 6 // DR6: the trap status register
drControl = 7 // DR7: the debug control register
)
// WatchpointType selects what triggers the watchpoint.
type WatchpointType int
const (
WatchWrite WatchpointType = 1 // trigger on write
WatchRead WatchpointType = 3 // trigger on read or write
)
// maxWatchpoints reports the number of hardware watchpoint slots the
// architecture provides: four address registers, DR0-DR3.
func maxWatchpoints() int { return 4 }
// getDbRegs reads the debug register file of the stopped debuggee.
func (s *Session) getDbRegs() (*dbreg, error) {
var dr dbreg
if _, _, errno := unix.Syscall6(
unix.SYS_PTRACE,
uintptr(unix.PT_GETDBREGS),
uintptr(s.pid),
0,
uintptr(unsafe.Pointer(&dr)),
0, 0,
); errno != 0 {
return nil, errno
}
return &dr, nil
}
// setDbRegs writes the debug register file of the stopped debuggee.
func (s *Session) setDbRegs(dr *dbreg) error {
if _, _, errno := unix.Syscall6(
unix.SYS_PTRACE,
uintptr(unix.PT_SETDBREGS),
uintptr(s.pid),
0,
uintptr(unsafe.Pointer(dr)),
0, 0,
); errno != 0 {
return errno
}
return nil
}
// archStopTrace classifies a TRAP_TRACE stop. On amd64 the kernel
// delivers both the completed single-step and the debug-register hit
// through T_TRCTRAP with TRAP_TRACE (sys/amd64/amd64/trap.c), and DR6's
// B0-B3 bits name the watchpoint that fired.
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
dr, err := s.getDbRegs()
if err != nil {
return StopSingleStep, 0
}
if status := dr.Dr[drStatus]; status&0xF != 0 {
for slot := range 4 {
if status&(1<<slot) != 0 && dr.Dr[slot] != 0 {
return StopWatchpoint, dr.Dr[slot]
}
}
}
return StopSingleStep, 0
}
// FindFreeWatchpointSlot returns the index of the first free watchpoint slot
// (0-3), or -1 if all four hardware watchpoints are in use.
func (s *Session) FindFreeWatchpointSlot() int {
for i := range 4 {
if !s.wpSlots[i] {
return i
}
}
return -1
}
// IsWatchpointSlotUsed reports whether slot (0-3) currently holds a watchpoint.
func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot > 3 {
return false
}
return s.wpSlots[slot]
}
// SetWatchpoint installs a hardware watchpoint on the given address.
// DR7's encoding is architectural: a 2-bit local/global enable pair per
// slot at bit 2*slot, the R/W field at 16+4*slot and the length field at
// 18+4*slot (Intel SDM vol 3, "Debug Registers").
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
if slot < 0 || slot > 3 {
return fmt.Errorf("debug: watchpoint slot must be 0-3")
}
if s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
}
var lenBits uint64
switch size {
case 1:
lenBits = 0
case 2:
lenBits = 1
case 4:
lenBits = 3
case 8:
lenBits = 2
default:
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
}
dr, err := s.getDbRegs()
if err != nil {
return fmt.Errorf("debug: read debug registers: %w", err)
}
dr.Dr[slot] = addr
dr7 := dr.Dr[drControl]
enableBit := uint64(1) << (2 * slot)
rwBits := uint64(typ) << (16 + 4*slot)
lenField := lenBits << (18 + 4*slot)
mask := ^((uint64(1) << (2 * slot)) | (uint64(3) << (16 + 4*slot)) | (uint64(3) << (18 + 4*slot)))
dr.Dr[drControl] = (dr7 & mask) | enableBit | rwBits | lenField
if err := s.setDbRegs(dr); err != nil {
return fmt.Errorf("debug: set debug registers: %w", err)
}
s.wpSlots[slot] = true
return nil
}
// ClearWatchpoint removes a hardware watchpoint.
func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot > 3 {
return fmt.Errorf("debug: watchpoint slot must be 0-3")
}
if !s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
}
dr, err := s.getDbRegs()
if err != nil {
return err
}
dr.Dr[slot] = 0
dr.Dr[drControl] &^= uint64(1) << (2 * slot)
if err := s.setDbRegs(dr); err != nil {
return err
}
s.wpSlots[slot] = false
return nil
}
// ClearAllWatchpoints removes all hardware watchpoints.
func (s *Session) ClearAllWatchpoints() error {
for slot := range maxWatchpoints() {
if s.wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil {
return err
}
}
}
return nil
}
+205
View File
@@ -0,0 +1,205 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && arm64
package debug
import (
"fmt"
"unsafe"
"golang.org/x/sys/unix"
)
// Hardware watchpoint support via arm64 debug registers, read and written
// as one blob through PT_GETDBREGS/PT_SETDBREGS. The FreeBSD struct dbreg
// (sys/arm64/include/reg.h) opens with the debug-facility header and then
// carries 16 breakpoint and 16 watchpoint pairs of {address, control}.
// dbreg mirrors FreeBSD's struct dbreg for PT_GETDBREGS/PT_SETDBREGS.
type dbreg struct {
DbDebugVer uint8
DbNbkpts uint8
DbNwtpts uint8
_ [5]byte
DbBreakregs [16]struct {
Addr uint64
Ctrl uint32
_ uint32
}
DbWatchregs [16]struct {
Addr uint64
Ctrl uint32
_ uint32
}
}
// WatchpointType selects what triggers the watchpoint.
type WatchpointType int
const (
WatchWrite WatchpointType = 1 // trigger on write
WatchRead WatchpointType = 3 // trigger on read or write
)
// maxWatchpoints reports the number of hardware watchpoint slots the
// architecture provides: DBGWVR0-DBGWCR15.
func maxWatchpoints() int { return 16 }
// getDbRegs reads the debug register file of the stopped debuggee.
func (s *Session) getDbRegs() (*dbreg, error) {
var dr dbreg
if _, _, errno := unix.Syscall6(
unix.SYS_PTRACE,
uintptr(unix.PT_GETDBREGS),
uintptr(s.pid),
0,
uintptr(unsafe.Pointer(&dr)),
0, 0,
); errno != 0 {
return nil, errno
}
return &dr, nil
}
// setDbRegs writes the debug register file of the stopped debuggee.
func (s *Session) setDbRegs(dr *dbreg) error {
if _, _, errno := unix.Syscall6(
unix.SYS_PTRACE,
uintptr(unix.PT_SETDBREGS),
uintptr(s.pid),
0,
uintptr(unsafe.Pointer(dr)),
0, 0,
); errno != 0 {
return errno
}
return nil
}
// archStopTrace classifies a TRAP_TRACE stop. On arm64 the kernel
// delivers both the software single step and the watchpoint hit through
// EXCP_SOFTSTP_EL0/EXCP_WATCHPT_EL0 with TRAP_TRACE (sys/arm64/arm64/
// trap.c); the watchpoint address rides the FAR register, so a stop whose
// reported address falls inside an armed watchpoint's byte range is a
// watchpoint and everything else is a single step.
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
dr, err := s.getDbRegs()
if err != nil {
return StopSingleStep, 0
}
for slot := range 16 {
ctrl := uint64(dr.DbWatchregs[slot].Ctrl)
if ctrl&1 == 0 || dr.DbWatchregs[slot].Addr == 0 {
continue
}
if bas := (ctrl >> 5) & 0xFF; bas != 0 && siAddr >= dr.DbWatchregs[slot].Addr && siAddr < dr.DbWatchregs[slot].Addr+8 {
return StopWatchpoint, siAddr
}
}
return StopSingleStep, 0
}
// FindFreeWatchpointSlot returns the index of the first free watchpoint
// slot, or -1 if all of them are in use.
func (s *Session) FindFreeWatchpointSlot() int {
for i := range maxWatchpoints() {
if !s.wpSlots[i] {
return i
}
}
return -1
}
// IsWatchpointSlotUsed reports whether slot currently holds a watchpoint.
func (s *Session) IsWatchpointSlotUsed(slot int) bool {
if slot < 0 || slot >= maxWatchpoints() {
return false
}
return s.wpSlots[slot]
}
// SetWatchpoint installs a hardware watchpoint on the given address. The
// control word is the architectural DBGWCR (ARM DDI 0487): bit 0 enables,
// bits 3-4 select the access type (10 store, 11 load+store) and bits 5-12
// are the byte-address select, so the watch stays 8-byte aligned and names
// its watched bytes through BAS.
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
}
if s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
}
var bas uint64
switch size {
case 1:
bas = 0x01
case 2:
bas = 0x03
case 4:
bas = 0x0F
case 8:
bas = 0xFF
default:
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
}
dr, err := s.getDbRegs()
if err != nil {
return fmt.Errorf("debug: read debug registers: %w", err)
}
if uint8(slot) >= dr.DbNwtpts && dr.DbNwtpts != 0 {
return fmt.Errorf("debug: slot %d exceeds available watchpoints (%d)", slot, dr.DbNwtpts)
}
ctrl := uint64(1) // enable
switch typ {
case WatchWrite:
ctrl |= 2 << 3 // store only
case WatchRead:
ctrl |= 3 << 3 // load+store
}
ctrl |= bas << 5
dr.DbWatchregs[slot].Addr = addr
dr.DbWatchregs[slot].Ctrl = uint32(ctrl)
if err := s.setDbRegs(dr); err != nil {
return fmt.Errorf("debug: set debug registers: %w", err)
}
s.wpSlots[slot] = true
return nil
}
// ClearWatchpoint removes a hardware watchpoint.
func (s *Session) ClearWatchpoint(slot int) error {
if slot < 0 || slot >= maxWatchpoints() {
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
}
if !s.wpSlots[slot] {
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
}
dr, err := s.getDbRegs()
if err != nil {
return err
}
dr.DbWatchregs[slot].Addr = 0
dr.DbWatchregs[slot].Ctrl = 0
if err := s.setDbRegs(dr); err != nil {
return err
}
s.wpSlots[slot] = false
return nil
}
// ClearAllWatchpoints removes all hardware watchpoints.
func (s *Session) ClearAllWatchpoints() error {
for slot := range maxWatchpoints() {
if s.wpSlots[slot] {
if err := s.ClearWatchpoint(slot); err != nil {
return err
}
}
}
return nil
}
+50
View File
@@ -0,0 +1,50 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
//go:build freebsd && riscv64
package debug
import "fmt"
// The architecture has hardware watchpoint triggers, but FreeBSD exposes
// no PT_GETDBREGS request for riscv64, so there is no supported way to arm
// one: the watchpoint layer is honestly empty here.
// WatchpointType selects what triggers the watchpoint.
type WatchpointType int
const (
WatchWrite WatchpointType = 1 // trigger on write
WatchRead WatchpointType = 3 // trigger on read or write
)
// maxWatchpoints reports the number of hardware watchpoint slots the
// platform provides: FreeBSD exposes none for riscv64.
func maxWatchpoints() int { return 0 }
// archStopTrace classifies a TRAP_TRACE stop; with no watchpoint layer a
// trace stop is always a completed single step.
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
return StopSingleStep, 0
}
// FindFreeWatchpointSlot returns -1: no slots exist.
func (s *Session) FindFreeWatchpointSlot() int { return -1 }
// IsWatchpointSlotUsed reports whether slot currently holds a watchpoint.
func (s *Session) IsWatchpointSlotUsed(slot int) bool { return false }
// SetWatchpoint is unsupported: FreeBSD exposes no debug register request
// for riscv64.
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
return fmt.Errorf("debug: hardware watchpoints are not supported on freebsd/riscv64")
}
// ClearWatchpoint is unsupported for the same reason.
func (s *Session) ClearWatchpoint(slot int) error {
return fmt.Errorf("debug: hardware watchpoints are not supported on freebsd/riscv64")
}
// ClearAllWatchpoints is a no-op: no watchpoint can be armed.
func (s *Session) ClearAllWatchpoints() error { return nil }
+1 -1
View File
@@ -15,7 +15,7 @@ import (
"golang.org/x/arch/riscv64/riscv64asm"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
)
// Instruction is one decoded instruction: its text form, its length in bytes
+4 -4
View File
@@ -7,10 +7,10 @@ import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
)
func TestDecodeKnownBytes(t *testing.T) {
+30 -7
View File
@@ -1,8 +1,8 @@
# Architecture
How gasm-devkit is put together and why.
How gasm-sdk is put together and why.
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
Repository: [sourcedock.dev/petrbalvin/gasm-sdk](https://sourcedock.dev/petrbalvin/gasm-sdk)
## Overview
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
### Optional preprocessing
With `Options{Expand: true}` the parser runs a pre-parse pass
(`preproc.go`) that splices `#include` files (the source directory, then the
`-I` directories), expands object and parameterised `#define` macros,
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
folds constant expressions left in operands. The go command's platform
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
language server read the raw file. The command layer adds the go_asm.h
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
defines type-checked out of its Go files for the target architecture and
GOOS, with no compiler in the loop.
### `arch`
Register files are generated programmatically (the regular `R8`-`R15`,
@@ -232,11 +246,18 @@ document store, republishes diagnostics on every change, and provides:
pseudo-registers, labels, immediates, comments, directives, textflag macros);
- **navigation**: go-to-definition from a label reference to its definition,
find references, document highlights of every use of the symbol under the
cursor, rename, and workspace symbol search over the open documents;
cursor, rename, and workspace symbol search over the open documents and
the indexed workspace files: the `.s` files under the workspace root that
the editor has never opened, where an open buffer shadows its disk copy
and watched-file events plus a per-query freshness check keep the index
current;
- **assists**: document formatting through the `format` package, inlay hints
(the frame size after the TEXT argument area), signature help (the callee's
`// func` signature while the cursor is on a `CALL`), and code actions
offering quick fixes for the `missing-ret` and `unused-label` diagnostics.
offering quick fixes for the `missing-ret` and `unused-label` diagnostics,
the `missing-textflag-include` warning (the include after the last one in
the file) and the `abi-argsize` warning (the argument area set to the size
the `// func` signature implies).
- **document information**: pull diagnostics (`textDocument/diagnostic`),
#include document links (resolved against the document directory, then
`$GOROOT/pkg/include`) and folding ranges (one collapsible region per
@@ -558,7 +579,9 @@ AST, so neither depends on an encoding.
--ground-truth`, `go list -json -export` locates the archives of the packages
a GOOBJ object references, and `_gen` parses
`$GOROOT/src/cmd/internal/obj/<arch>/anames.go` to rebuild the tables.
- **Linux process interfaces** for the dynamic work: `mmap` and `mprotect` for
the JIT mapping, ptrace with `/proc/pid/mem` for the debugger. That is why
- **Linux and FreeBSD process interfaces** for the dynamic work: `mmap` and
`mprotect` for the JIT mapping, ptrace for the debugger — with tracee
memory through `/proc/pid/mem` on Linux and through `PT_IO` on FreeBSD,
and the tracee's stop reports read from `PT_LWPINFO` there. That is why
`verify` runs a JIT check only when the host architecture matches the
kernel's, and why `debug` is Linux-only.
kernel's, and why `debug` is bounded to those two kernels.
+13 -4
View File
@@ -287,7 +287,8 @@ Usage: gasm debug <file.s> --func <name>
The debugger re-executes the binary it is running as (`os.Executable()`) for the
traced child, so the child is the same `gasm`, whether it is installed on `$PATH`
or run with `go run ./cmd/gasm`; nothing has to be installed first. Requires
Linux (ptrace), and all four architectures are supported.
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64, arm64 and
riscv64 on FreeBSD.
REPL commands:
@@ -365,7 +366,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
## audit-instructions
```text
Usage: gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
```
Compare the gasm encoder for the given architecture (default amd64) against the
@@ -403,7 +404,9 @@ architecture; a file without one is attempted for all four, exactly as a
The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
drive the encodability backlog by frequency rather than by table order. With
`--list` the report additionally prints every failing file with its failure
reason, per architecture. A run
over GOROOT takes under a second.
```sh
@@ -453,7 +456,13 @@ semantic tokens, go-to-definition, find references, rename, document
formatting, inlay hints, code actions, signature help, document highlights,
workspace symbol search, #include document links, and folding ranges for
function bodies. Definition, references and rename work across every open
document.
document and the wider workspace on disk: the server indexes the `.s` files
under the workspace root that the editor has never opened, an open buffer
always shadows its disk copy, and watched-file events together with a
per-query freshness check keep the index current. The quick fixes add the
missing `#include "textflag.h"`, set the TEXT argument area to the size the
`// func` signature implies, add a missing `RET`, and remove an unused
label.
## version
+3 -3
View File
@@ -1,6 +1,6 @@
# Development Guide
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
Repository: [sourcedock.dev/petrbalvin/gasm-sdk](https://sourcedock.dev/petrbalvin/gasm-sdk)
## Prerequisites
@@ -19,8 +19,8 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
## Setup
```sh
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
cd gasm-devkit
git clone https://sourcedock.dev/petrbalvin/gasm-sdk.git
cd gasm-sdk
just build # compile bin/gasm, zero errors and zero warnings
just gates # build, fmt-check, vet, test, race: the definition of done
```
+657
View File
@@ -0,0 +1,657 @@
# The GOOBJ object file format
This document is a complete specification of GOOBJ, the object file format
that the Go toolchain's assembler, compiler and linker exchange, written for
implementers of independent producers and consumers. It documents the format
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
No comparable document exists upstream. The format is defined only by the
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
internal interface with no stability promise, and it can change in any
release. This specification was therefore produced by reverse engineering
that source and by parsing real objects produced by `go tool asm` and
`go tool compile`, byte for byte, against the layout described here. Within
gasm-sdk it is kept honest by the differential tests in `asm/goobj_test.go`
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
the toolchain's own products and feed gasm objects to `go build`.
Every numeric value in this document, every block index, structure size, flag
bit, type code and relocation number, was read from the Go 1.27.1 source at
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
`cmd/internal/objabi`, and exercised against assembled objects.
## Containers
The unit this document specifies is the **object**: one package's worth of
symbols, relocations and data. An object is never consumed naked. Two
wrappers exist in practice, and the linker dispatches on the first bytes of
the file.
**The bare object**, written by `go tool asm`:
```text
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
"!\n"
<GOOBJ blob>
```
The first line is the toolchain configuration string, produced by
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
version, an optional architecture qualifier such as `GOAMD64=v1`, and
`X:` followed by the enabled experiments, comma separated. The linker requires
this line to match its own configuration exactly and rejects the file
otherwise; the `-f` linker flag waives the check. Header lines may be
followed by export data delimited by `$$` markers; the header region always
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
immediately after that line.
**The package archive**, written by the compiler output pipeline and consumed
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
data in a `__.PKGDEF` member and one or more objects as further members, each
carrying the bare-object structure above. `go tool pack` creates and
inspects such archives.
| Consumer | Role |
|---|---|
| `cmd/asm` | writes objects from `.s` files |
| `cmd/compile` | writes objects from Go source |
| `cmd/link` | reads objects and archives, produces executables |
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
## Conventions
- All integers are **little endian**.
- There is **no alignment or padding** anywhere in the file; structures follow
one another byte by byte.
- Every offset stored in the file is **relative to the first byte of the GOOBJ
blob**, not to the start of the container.
- The blob opens with a 96 byte header that carries the byte offset of every
block. A block's length is the difference between its own offset and the
next block's, so the offset array is the only index the format needs.
### Layout overview
```mermaid
flowchart TB
A[Container header line and ! terminator] --> B[File header, 96 bytes]
B --> C[String table, implicit region]
C --> D[Autolib]
D --> E[PkgIndex]
E --> F[Files]
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
G --> H[RefFlags]
H --> I[Hash64 and Hash]
I --> J[RelocIndex, AuxIndex, DataIndex]
J --> K[Relocs]
K --> L[Aux]
L --> M[Data]
M --> N[RefNames]
N --> O[BlkEnd marks the end of the blob]
```
## The file header
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
offsets.
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
| 16 | 4 | Flags | Bit field, see below. |
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
Header flags:
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | ObjFlagShared | built with `-shared` |
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
| 4 | 16 | ObjFlagStd | standard library package |
### Block indices
The offset array is indexed by these constants, in file order:
| Index | Constant | Contents |
|---|---|---|
| 0 | BlkAutolib | imported packages |
| 1 | BlkPkgIndex | referenced packages, indexed |
| 2 | BlkFile | source file names |
| 3 | BlkSymdef | symbol definitions, package scope |
| 4 | BlkHashed64def | short hashed definitions |
| 5 | BlkHasheddef | hashed definitions |
| 6 | BlkNonpkgdef | non-package definitions |
| 7 | BlkNonpkgref | non-package references |
| 8 | BlkRefFlags | flags of referenced symbols |
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
| 10 | BlkHash | 16 byte hashes for hashed definitions |
| 11 | BlkRelocIndex | per symbol relocation start index |
| 12 | BlkAuxIndex | per symbol aux start index |
| 13 | BlkDataIndex | per symbol data offset |
| 14 | BlkReloc | relocations |
| 15 | BlkAux | aux symbol entries |
| 16 | BlkData | symbol payloads |
| 17 | BlkRefName | names of referenced symbols, for tools |
| 18 | BlkEnd | no contents; its offset is the end of the blob |
## The string table
There is no block index for strings. The table occupies the implicit region
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
string offset in the file points into that region. The writer de-duplicates:
each distinct string is stored once, in first-use order, and the empty string
is always the first entry, so its reference is length 0 and offset 96.
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
offset of the bytes. The bytes are stored raw, with no terminator.
## Symbol references and the package index
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
lives in:
| Value | Constant | SymIdx indexes |
|---|---|---|
| 0 | PkgIdxInvalid | never valid in a written file |
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
Assignment rules, as the toolchain performs them:
- Every definition a package exports to the linker by index lands in Symdefs
with PkgIdxSelf. The compiler puts its functions and data here; the
assembler puts only its file-local static symbols here, everything else by
name, see below.
- External package references take indices 1, 2, 3, in order of first
reference during assembly; the package names go into PkgIndex at those
indices, entry 0 is the empty package and is never referenced.
- References to the compiler's builtin functions become PkgIdxBuiltin with
SymIdx set to the builtin's index.
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
A producer must emit the definitions it made in NonPkgDefs and the pure
references in NonPkgRefs.
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
is referenced by name, PkgIdxNone, **except** file-local static symbols,
whose names carry `<>` and which are referenced by index. The compiler also
forces references by name for symbols marked `//go:linkname` and for any
symbol with the DUPOK attribute, which the linker de-duplicates by name.
## Symbol definition entries
The five definition and reference arrays (block indices 3 to 7) share one
element layout, 21 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Name | string reference |
| 8 | 2 | ABI | see table below |
| 10 | 1 | Type | symbol kind, see the kind table |
| 11 | 1 | Flag | bit field, see below |
| 12 | 1 | Flag2 | second bit field, see below |
| 13 | 4 | Siz | payload size in bytes, `uint32` |
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
The Name is a real string reference for hand-written symbols. The auxiliary
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
entries, have empty names: length 0, and their identity is only via the Aux
entries that point at them by index.
### The ABI field
| Value | Meaning |
|---|---|
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
| 1 | ABIInternal, the register ABI of compiler-generated functions |
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
### The Flag byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
| 1 | 2 | SymFlagLocal | file-local |
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
| 3 | 8 | SymFlagLeaf | leaf function |
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
amd64 the assembler itself marks any function whose frame is below
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
### The Flag2 byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
| 3 | 8 | SymFlagPkgInit | package initialisation function |
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
### The Type byte: symbol kinds
Values of `objabi.SymKind`, in numeric order:
| Value | Name | Meaning |
|---|---|---|
| 0 | Sxxx | invalid zero value |
| 1 | STEXT | executable code |
| 2 | STEXTFIPS | executable code, FIPS section |
| 3 | SRODATA | read only data |
| 4 | SRODATAFIPS | read only data, FIPS section |
| 5 | SNOPTRDATA | data without pointers |
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
| 7 | SDATA | data, may contain pointers |
| 8 | SDATAFIPS | data, FIPS section |
| 9 | SBSS | zero initialised data |
| 10 | SNOPTRBSS | zero initialised data without pointers |
| 11 | STLSBSS | thread local zero initialised data |
| 12 | SDWARFCUINFO | DWARF compile unit information |
| 13 | SDWARFCONST | DWARF constants |
| 14 | SDWARFFCN | DWARF function entry |
| 15 | SDWARFABSFCN | DWARF absolute function entry |
| 16 | SDWARFTYPE | DWARF type information |
| 17 | SDWARFVAR | DWARF variable information |
| 18 | SDWARFRANGE | DWARF range lists |
| 19 | SDWARFLOC | DWARF location lists |
| 20 | SDWARFLINES | DWARF line programs |
| 21 | SDWARFADDR | DWARF address table |
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
| 23 | SCOVERAGE_COUNTER | coverage counter |
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
## Referenced symbol flags (RefFlags)
Element size 10 bytes, one per referenced external indexed symbol that
carries a non-zero Flag2:
| Offset | Size | Field |
|---|---|---|
| 0 | 8 | Sym, a SymRef into another package |
| 8 | 1 | Flag, always 0 in current writers |
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
The linker uses these to preserve reachability of interface conversions
across package boundaries. Entries with no flags are omitted entirely.
## Hashes
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
a hash at all: the writer copies the **first 8 bytes of the symbol's
payload**. Only symbols whose content-hash section byte is 0 may use the
short form.
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
input, from `cmd/internal/obj/objfile.go`:
1. the payload size, little endian `uint64`;
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
symbols, otherwise 0;
3. for text symbols, the symbol name, which keeps distinct functions from
merging;
4. the payload with trailing zero bytes trimmed;
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
low type byte `uint8`, addend `int64`, followed by an encoding of the
target: tag byte 0 then the target's short hash, tag 1 then its full
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
for PkgIdxSelf and imported packages, no tag, then the package path
and the symbol index.
Two symbols with equal hashes are interchangeable at link time, which is what
makes content addressing work. A producer that computes these hashes wrongly
produces objects that link but de-duplicate wrongly; gasm verifies them by
byte comparison against `go tool asm`.
## The index arrays
Three arrays of `uint32`, one element per **defined** symbol plus one final
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
defined symbols, each array holds N + 1 entries, and the entry at N is the
total.
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
entry i + 1 minus entry i is its count.
- AuxIndex: the same construction over BlkAux.
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
the count is the difference of neighbours.
The toolchain writes relocations grouped per symbol in definition order, and
sorts each symbol's relocations by their Off field first. A producer that
skips the sort produces objects the linker still accepts, but that no longer
compare byte-for-byte with the toolchain's output.
## Relocations
Element size 23 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
| 4 | 1 | Siz | patch width in bytes |
| 5 | 2 | Type | relocation type, `uint16`, see the table |
| 7 | 8 | Add | addend, `int64` |
| 15 | 8 | Sym | target SymRef |
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
flavour the type prescribes is the linker's job; the object only records the
request. A size 0 relocation patches nothing and exists purely as a marker
for the linker's reachability analysis.
### Relocation types
Values of `objabi.RelocType`. The assembler and compiler emit the generic
ones plus their own architecture's family; the rest exist for other ports and
for the linker itself.
| Value | Name | Meaning |
|---|---|---|
| 1 | R_ADDR | absolute address |
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
| 6 | R_SIZE | size of the referenced symbol |
| 7 | R_CALL | direct call, PC relative |
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
| 9 | R_CALLARM64 | arm64: BL |
| 10 | R_CALLIND | indirect call marker |
| 11 | R_CALLPOWER | ppc64: call |
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
| 13 | R_CONST | constant value of the symbol |
| 14 | R_PCREL | PC relative displacement |
| 15 | R_TLS_LE | thread local, local exec offset |
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
| 17 | R_GOTOFF | offset from the GOT base |
| 18 | R_PLT0 | PLT sequence, first instruction |
| 19 | R_PLT1 | PLT sequence, second instruction |
| 20 | R_PLT2 | PLT sequence, third instruction |
| 21 | R_USEFIELD | field reachability marker |
| 22 | R_USETYPE | type reachability marker, no bytes patched |
| 23 | R_USEIFACE | interface conversion marker, size 0 |
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
| 27 | R_KEEP | keeps the target alive if the source survives |
| 28 | R_POWER_TOC | ppc64: TOC relative |
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
| 99 | R_WASMIMPORT | wasm: import module and name indices |
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
| 102 | R_INITORDER | orders inittask records, patches nothing |
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
type's value with bit 15 set. The linker strips the bit before dispatch.
## Aux symbol entries
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
auxiliary symbols to a definition; the arrays run per symbol in the order
given by AuxIndex.
| Value | Name | Attaches |
|---|---|---|
| 0 | AuxGotype | the Go type of a data symbol |
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
| 3 | AuxDwarfInfo | DWARF debug info for the function |
| 4 | AuxDwarfLoc | DWARF location lists |
| 5 | AuxDwarfRanges | DWARF range lists |
| 6 | AuxDwarfLines | DWARF line program |
| 7 | AuxPcsp | pc-value table: SP adjustments |
| 8 | AuxPcfile | pc-value table: source file indices |
| 9 | AuxPcline | pc-value table: line numbers |
| 10 | AuxPcinline | pc-value table: inlining tree positions |
| 11 | AuxPcdata | one pc-value table per live variable slot |
| 12 | AuxWasmImport | wasm import description |
| 13 | AuxWasmType | wasm export type description |
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
payload would be empty. A function assembled from `.s` source by Go 1.27.1
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
same set.
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
resolve to definitions in the package's compiled Go code when there is any.
## Symbol payloads (BlkData)
The payloads of all defined symbols, in definition order, concatenated with
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
its machine code, with the stack-split preamble and any morestack block
already included. A data symbol's payload is the bytes laid down by its DATA
directives, zero filled to its declared size. If a symbol was created from an
embedded file, the file's bytes follow the payload and count towards its
DataIndex extent; assembly producers never write this extension.
### The FuncInfo payload
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
little endian:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
| 4 | 4 | Locals | frame size in bytes |
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
| 12 | 4 | StartLine | source line of the TEXT declaration |
| 16 | 4 | NumFile | count of file indices that follow |
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
| then | 24 × NumInlTree | InlTree | nodes, see below |
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
`int32` parent PC.
The assembler derives FuncID from the symbol name through
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
specially gets that classification even when defined in assembly; an ordinary
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
### The pc-value tables
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
pc-value tables, each a sequence of value deltas and PC deltas:
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
raw delta divided by the architecture's minimum instruction length;
- the table ends with a final PC delta to the end of the function followed by
a zero byte.
The first value applies from function entry. The encoding is the one
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
final runtime pclntable carries.
### The DWARF payloads
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
symbols whose payloads are DWARF byte streams. The object format treats them
as opaque: the linker concatenates them into the final `.debug_*` sections
and resolves the relocations recorded inside them. The compiler produces
DWARF content per its own generation; gasm produces DWARF5 streams in
`asm/goobj_dwarf.go`.
## Builtins
Frequently referenced runtime functions are referenced by index rather than
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
ABI 0. Builtin names never enter the string table. The mapping only applies
while the object is not linked against shared libraries, and a linkname'd
symbol never counts as a builtin even when its name matches.
## Fingerprints
The 8 byte fingerprint identifies one build of a package. The compiler fills
it with a hash of the package's export data; the assembler leaves it zero.
The linker checks a package's fingerprint against the fingerprints its
importers recorded in their Autolib entries and rejects a mismatched build,
which is how stale objects are caught.
## What a producer must do
The checklist a third-party writer must satisfy for `go build` to accept its
objects, in one place:
1. Write the container exactly: the `go object` line matching the target
toolchain's configuration string, the `!\n` terminator, then the blob.
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
3. Deduplicate the string table, keep the empty string at offset 96, and
reference it everywhere a name appears.
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
5. Sort relocations by offset within each symbol.
6. Fill Siz with the true payload length, set Align for every
content-addressable symbol, and keep symbols under 2 GB.
7. Reference symbols by the package-index rules. An assembly producer
references everything outside the object by name, PkgIdxNone,
except its own file-local statics and the builtins; PkgIdxSelf is
reserved for definitions in this object. Assembly TEXT symbols
carry ABI 0.
8. Compute the content hashes exactly as the toolchain does, or emit no
hashed definitions at all.
## How gasm-sdk implements and verifies it
The writer lives in `asm/goobj.go`, which carries the shared container and the
amd64 relocation emission, with per-architecture relocation emitters in
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
resolution in `asm/goobj_resolve.go` and DWARF generation in
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
that `go build` consumes in place of the toolchain's own.
Verification is differential and continuous:
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
output for the same source, byte for byte;
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
objects and runs them;
- `gasm verify` keeps the machine code itself identical to the toolchain's,
which is the precondition for the object comparison to be meaningful.
## Versioning and drift
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
a toolchain release changes the format, it changes that string first, and the
linker refuses blobs whose magic it does not know. The watch points for a new
release are, in order: the magic, the block index list, the Aux type list,
the tail of the relocation table, the FuncInfo layout, and the builtin table
count. gasm's tests fail against any of these changes, which is the mechanism
that keeps this document and the writer current.
The authoritative sources, for the release this document covers:
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
document;
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
- `cmd/internal/objabi/reloctype.go`: relocation types;
- `cmd/internal/objabi/symkind.go`: symbol kinds;
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.

Some files were not shown because too many files have changed in this diff Show More