Compare commits
33
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4d01bb3ecf | ||
|
|
332c63e440 | ||
|
|
b306c210c6 | ||
|
|
b9015e1c2e | ||
|
|
26c5008136 | ||
|
|
74d6b90d69 | ||
|
|
7b11c62f53 | ||
|
|
8eed54b3da | ||
|
|
4be16dcdf5 | ||
|
|
ded9cabdf4 | ||
|
|
cf6bc6987e | ||
|
|
ff7b1452b1 | ||
|
|
517c1cea25 | ||
|
|
a3e3010e0f | ||
|
|
057c4eb545 | ||
|
|
f720381d43 | ||
|
|
2c9042d62c | ||
|
|
82ef289d3a | ||
|
|
7246b0e002 | ||
|
|
8cfd40aac8 | ||
|
|
5382c9a8e4 | ||
|
|
53de91b2df | ||
|
|
8a36af7c7d | ||
|
|
e9789ce3f4 | ||
|
|
837231c068 | ||
|
|
95025be1bc | ||
|
|
03a964bb2d | ||
|
|
123a16e346 | ||
|
|
9701812bee | ||
|
|
29ac03468e | ||
|
|
bfb7701db1 | ||
|
|
e8b6ff5d7c | ||
|
|
1456907000 |
@@ -342,7 +342,10 @@ jobs:
|
||||
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
|
||||
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
|
||||
q{-H}, q{Content-Type: application/octet-stream},
|
||||
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
|
||||
# The @ must not sit inside a qq{} string: there it starts an
|
||||
# array interpolation and the upload body collapses to empty,
|
||||
# which Gitea stores as a 201-created zero-byte attachment.
|
||||
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
|
||||
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
|
||||
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
|
||||
my $code = <$curl>;
|
||||
|
||||
@@ -56,6 +56,28 @@ jobs:
|
||||
- name: Build
|
||||
run: go build ./...
|
||||
|
||||
- name: FreeBSD build (amd64)
|
||||
# The debugger's ptrace surface and the JIT substrate are the two
|
||||
# FreeBSD-portable layers the tree carries; the forge has no FreeBSD
|
||||
# runner, so a push can only compile-gate them. Running the ptrace
|
||||
# suite needs real FreeBSD hardware.
|
||||
env:
|
||||
GOOS: freebsd
|
||||
GOARCH: amd64
|
||||
run: go build ./...
|
||||
|
||||
- name: FreeBSD build (arm64)
|
||||
env:
|
||||
GOOS: freebsd
|
||||
GOARCH: arm64
|
||||
run: go build ./...
|
||||
|
||||
- name: FreeBSD build (riscv64)
|
||||
env:
|
||||
GOOS: freebsd
|
||||
GOARCH: riscv64
|
||||
run: go build ./...
|
||||
|
||||
- name: Format
|
||||
run: |
|
||||
perl -e '
|
||||
|
||||
+168
-10
@@ -1,6 +1,6 @@
|
||||
# Changelog
|
||||
|
||||
All notable changes to gasm-devkit are documented here.
|
||||
All notable changes to gasm-sdk are documented here.
|
||||
|
||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
@@ -9,6 +9,87 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Added
|
||||
|
||||
- **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on
|
||||
amd64, arm64 and riscv64: the same interactive surface as on Linux —
|
||||
breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
|
||||
register file), single-stepping, register and memory access — behind the
|
||||
kernel's own ptrace requests, with tracee memory through `PT_IO` and
|
||||
stop reports through `PT_LWPINFO`. The JIT substrate maps executable
|
||||
memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD
|
||||
too. The pipeline compile-gates all three architectures; live
|
||||
validation awaits a FreeBSD machine.
|
||||
- **Workspace-wide navigation in the language server.** `gasm lsp` indexes
|
||||
the `.s` files under the workspace root beyond the documents the editor
|
||||
has open, so go-to-definition, find references and workspace symbol search
|
||||
reach files that were never opened. An open buffer always shadows its
|
||||
disk copy, and watched-file events together with a per-query freshness
|
||||
check keep the index current.
|
||||
- **Quick fixes for the textflag include and the argument area.** The
|
||||
`missing-textflag-include` warning offers to add the include after the
|
||||
last one in the file, and the `abi-argsize` warning offers to set the
|
||||
TEXT argument area to the size the `// func` signature implies, computed
|
||||
by the new `lint.ExpectedArgSize`.
|
||||
|
||||
### Changed
|
||||
|
||||
- **The corpus audit assembles like the build.** A file's `//go:build`
|
||||
constraint decides which target architectures attempt it: cpu_x86.s is
|
||||
an x86 build alone, and the msan and goexperiment.runtimesecret trees
|
||||
are compiled by no supported build, so they leave the measured set
|
||||
instead of failing it. The headline now reads "assemble for every
|
||||
applicable target": every real-code GOROOT assembly file, the tree
|
||||
without testdata, assembles for all four architectures (250 of 250,
|
||||
100 %); over the whole tree including testdata the measure is 271 of
|
||||
322 (84.2 %).
|
||||
- **The module moves to `sourcedock.dev/petrbalvin/gasm-sdk`.** The
|
||||
repository and the module rename together with the product, now the
|
||||
GAsm Software Development Kit. Fresh installs become
|
||||
`go install sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm@latest`, and
|
||||
installs pinned to the old `gasm-sdk` path stop resolving once the
|
||||
repository takes the new name: reinstall from the new path. The
|
||||
binary stays `gasm`.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **Rename edits land in their own documents.** A rename collected the
|
||||
ranges of every reference across the open documents but applied them all
|
||||
to the document that started it, so renaming a symbol used in a second
|
||||
file moved that file's text into the first. Each edit now applies to the
|
||||
document it was collected in.
|
||||
- **Negative numeric PC-relative jumps.** `JMP -3(PC)`, the shape the
|
||||
runtime's exit loops write (sys_linux_amd64.s, sys_netbsd_amd64.s),
|
||||
resolved to nothing: only the forward forms counted. A negative count
|
||||
now walks the same instruction statements backwards, labels excluded,
|
||||
byte-identical with the toolchain.
|
||||
- **The arm64 move-wide family reads its immediate as an unsigned
|
||||
pattern.** `MOVK $(40000<<48)` folds to a negative int64 and was
|
||||
rejected; the toolchain picks the 16-bit lane from the 64-bit bit
|
||||
pattern, so the encoder now does the same, and a zero immediate is
|
||||
rejected where the toolchain rejects it.
|
||||
|
||||
## [0.35.0] - 2026-09-22
|
||||
|
||||
### Added
|
||||
|
||||
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
|
||||
itself when an assembly file includes it: the Go files beside the source
|
||||
are type-checked for the target architecture and the constants and field
|
||||
offsets become assembler defines, so package-context files assemble with
|
||||
no compiler and no `go build` in the loop. `-GOOS` selects the
|
||||
type-checking GOOS for GOOS-specific files, and the corpus audit derives
|
||||
the GOOS from the file name.
|
||||
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
|
||||
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
|
||||
every architecture (amd64 carried them already), so standalone ELF
|
||||
objects link on all four targets.
|
||||
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
|
||||
prints every failing file with its failure reason, per architecture,
|
||||
instead of one representative file per reason.
|
||||
- **DATA with symbol values and relaxed symbol spellings.** DATA
|
||||
initialisers accept `$symbol(SB)` values, laid down as an absolute
|
||||
relocation at the data field (GOOBJ on all four architectures and ELF
|
||||
on all four as of this release), and U+2215 is accepted inside symbol
|
||||
package paths.
|
||||
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
|
||||
`gasm audit-instructions` now preprocess assembly the way the
|
||||
toolchain does: object and parameterised `#define` macros expand at
|
||||
@@ -19,7 +100,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
|
||||
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
|
||||
language server keep reading the raw file.
|
||||
- **The GOROOT instruction wave, part 1.** The encoder now covers the
|
||||
- **Encoder coverage: the instruction families GOROOT's real code
|
||||
uses.** The encoder now covers the
|
||||
instruction families GOROOT's real code uses that gasm lacked,
|
||||
byte-verified against `go tool asm`: on amd64 the carry ALU, the
|
||||
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
|
||||
@@ -35,14 +117,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
|
||||
riscv64 `VSETVLI` with an immediate length now canonicalises to
|
||||
`vsetivli` as the toolchain does.
|
||||
- **The corpus audit measures honestly.** Files named for Go ports gasm
|
||||
does not target (arm, 386, s390x, ...) are no longer attempted for the
|
||||
four supported architectures (no supported build compiles them), and
|
||||
the headline rate is reported over attemptable files: 136 of 433 on
|
||||
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
|
||||
127 that the previous release measured. The probe battery that
|
||||
decides encodability gained the operand shapes the new families use.
|
||||
-
|
||||
- **Encoder coverage: quad-register AVX-512 and floating-point
|
||||
immediates.** The encoder gains the
|
||||
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
|
||||
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
|
||||
floating-point immediates on the SSE scalar moves and arithmetic
|
||||
(the constant lands in a synthesised read-only pool, a positive zero
|
||||
collapses to XORPS exactly as the toolchain does), accept-and-ignore
|
||||
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
|
||||
operands for the legacy SSE moves, and the pooled 64-bit immediate
|
||||
materialisation on riscv64. The parser carries bracketed register
|
||||
ranges, index-only VSIB memory operands and bare trailing immediates;
|
||||
macro substitution reaches parameters used with element suffixes
|
||||
(`A.S4`), and `;` separates statements in plain files.
|
||||
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
|
||||
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
|
||||
roles the ABI fixes, addressing, operand order with every special form,
|
||||
constants and materialisation, alignment, fences and the relocations
|
||||
each target emits. An instruction inventory appendix per architecture
|
||||
is generated from the toolchain's own tables by `just gen`, and the
|
||||
regenerated tables recognise 147 more mnemonics than the previous
|
||||
release carried (arm64 107, riscv64 31, loong64 9).
|
||||
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
|
||||
opens the complete language reference with its common core: the lexicon,
|
||||
statement structure and constant expressions, the operand grammar with
|
||||
the pseudo-registers and symbol naming, the directives and the function
|
||||
flag vocabulary, preprocessing with `#define` and `#include`, and the
|
||||
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
|
||||
runtime contract). Every claim is verified against `go tool asm` of
|
||||
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
|
||||
generated instruction appendices follow.
|
||||
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
|
||||
documents the Go object file format in full: both containers, the 96
|
||||
byte header and all 19 blocks, every structure with its byte
|
||||
offsets, symbol kinds and flag bits, all 106 relocation types with
|
||||
the weak variants, aux symbols, the FuncInfo payload, the pc-value
|
||||
table encoding, the content hashes and the builtin table, all
|
||||
verified byte for byte against objects produced by Go 1.27.1's own
|
||||
tools.
|
||||
|
||||
### Changed
|
||||
|
||||
- **The corpus audit measures like a build.** Files named for a Go port
|
||||
gasm does not target (arm, 386, s390x, ...) are never attempted, because
|
||||
no supported build compiles them; the GOOS comes from the file name; and
|
||||
each target's go_asm.h is generated on the fly. The headline is reported
|
||||
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
|
||||
for every target architecture and 295 of 303 on real code (97.4 %),
|
||||
against 108 of 627 over all files (17.2 %) that the previous release
|
||||
measured.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
|
||||
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
|
||||
own instruction counting and fold jump-to-jump chains exactly as its
|
||||
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
|
||||
assemble to the toolchain's RIP-relative LEA with an R_PCREL
|
||||
relocation; negated constant expressions in operands (`ADJSP
|
||||
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
|
||||
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
|
||||
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
|
||||
one-instruction form (the bare `MOVQ TLS, r` load nops out and
|
||||
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
|
||||
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
|
||||
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
|
||||
the zero-immediate store (`MOVD $0, mem` through the zero register,
|
||||
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
|
||||
aligns on amd64, padding with the toolchain's greedy
|
||||
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
|
||||
AX` and the store direction) and the absolute crash-store
|
||||
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
|
||||
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
|
||||
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
|
||||
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
|
||||
GOROOT corpus measure moves to 291 of 353 files assembling for every
|
||||
target architecture (82.4 %), 97.4 % of the real-code corpus, from
|
||||
70.8 % and 82.2 %.
|
||||
- **Tool corrections across the pipeline.** The formatter keeps square
|
||||
brackets in SIMD operands, statement separators and canonical macro
|
||||
bodies; the linter drops false positives on shift counts, SETcc
|
||||
spellings and ABIInternal references; the lexer treats a trailing
|
||||
carriage return as a line end so comment text stays idempotent; and
|
||||
arm64 rejects bare BTI with a diagnostic while accepting the full
|
||||
family.
|
||||
|
||||
## [0.34.0] - 2026-09-20
|
||||
|
||||
|
||||
+4
-4
@@ -1,6 +1,6 @@
|
||||
# Contributing
|
||||
|
||||
Contributions to **gasm-devkit** are governed by the Contributor terms
|
||||
Contributions to **gasm-sdk** are governed by the Contributor terms
|
||||
below; submitting one means you accept them.
|
||||
|
||||
## Contributor terms
|
||||
@@ -29,8 +29,8 @@ compiler (gcc), because `just gates` includes `just race` and the race
|
||||
detector needs cgo.
|
||||
|
||||
```sh
|
||||
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
|
||||
cd gasm-devkit
|
||||
git clone https://sourcedock.dev/petrbalvin/gasm-sdk.git
|
||||
cd gasm-sdk
|
||||
just build
|
||||
just gates
|
||||
```
|
||||
@@ -126,7 +126,7 @@ tag, where it would double the time and the memory a shared runner cannot spare.
|
||||
|
||||
## Reporting bugs
|
||||
|
||||
Open an issue at `https://sourcedock.dev/petrbalvin/gasm-devkit/issues` with the
|
||||
Open an issue at `https://sourcedock.dev/petrbalvin/gasm-sdk/issues` with the
|
||||
version, the operating system and architecture, the exact command, the full output,
|
||||
and the expected against the actual behaviour.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Plan 9 assembly tooling, inside and outside Go
|
||||
# GAsm: Software Development Kit for Plan 9 Assembly
|
||||
|
||||
> **Warning: this is an experiment.** gasm-devkit is under active
|
||||
> **Warning: this is an experiment.** gasm-sdk is under active
|
||||
> development and is not stable. The version is 0.x.x: commands, flags,
|
||||
> output formats and behaviour can change without warning at any time.
|
||||
> A 1.0.0 release is light years away. Nothing in this document is a
|
||||
@@ -13,7 +13,7 @@
|
||||
there is no formatter, no linter and no debugger for `.s` files, and no
|
||||
assembler that works without a Go installation. Developers write
|
||||
assembly blind, validate it by benchmark, and debug it by print
|
||||
statement. gasm-devkit is the missing toolkit: a single, self-contained
|
||||
statement. gasm-sdk is the missing toolkit: a single, self-contained
|
||||
binary, `gasm`, that serves both purposes.
|
||||
|
||||
- **Help develop Plan 9 assembly.** Formatting, linting, disassembly,
|
||||
@@ -58,7 +58,7 @@ Plan 9 (Go): MOVQ AX, total-16(SP)
|
||||
|
||||
The same lines, but only one of them tells you what the number is for.
|
||||
The syntax is uppercase, regular and boring, which is the highest
|
||||
compliment a language for machine code can earn. gasm-devkit exists
|
||||
compliment a language for machine code can earn. gasm-sdk exists
|
||||
to give that syntax the tooling it deserves.
|
||||
|
||||
## Features
|
||||
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
|
||||
GOOBJ format, which needs the installed toolchain and which `go build`
|
||||
consumes in place of the toolchain's output. Framed functions get the
|
||||
stack-split guard and the morestack block, byte-identical to the
|
||||
toolchain's, so split functions link too.
|
||||
toolchain's, so split functions link too. The assembler preprocesses
|
||||
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
|
||||
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
|
||||
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
|
||||
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
|
||||
offsets after assembling, or disassembles raw bytes from a file or stdin.
|
||||
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
|
||||
@@ -91,13 +94,16 @@ to give that syntax the tooling it deserves.
|
||||
- **Debugger.** `gasm debug` is a source-level ptrace debugger with
|
||||
breakpoints (optionally conditional), hardware watchpoints, register and
|
||||
memory inspection, and headless script runs that report instruction and
|
||||
label coverage.
|
||||
label coverage; it runs on Linux (all four architectures) and FreeBSD
|
||||
(amd64, arm64, riscv64).
|
||||
- **Language server.** `gasm lsp` serves completion, hover, document symbols,
|
||||
push and pull diagnostics, semantic-token highlighting, go-to-definition,
|
||||
find references, rename, formatting, inlay hints, code actions, signature
|
||||
help, document highlights, workspace symbol search, #include document
|
||||
links and folding ranges over stdio; definition, references and rename
|
||||
work across every open document.
|
||||
work across every open document and the indexed workspace files beyond
|
||||
them, and the quick fixes add a missing textflag.h include and set the
|
||||
argument area from the // func signature.
|
||||
- **Comparators and audits.** `gasm diff` compares the machine code of two
|
||||
assembly files byte-for-byte, `gasm profile` shows basic-block structure,
|
||||
`gasm audit-instructions` diffs the encoder against the installed toolchain,
|
||||
@@ -110,9 +116,9 @@ Four architectures, the four that matter in practice:
|
||||
| Architecture | GOARCH | File suffix | Instructions recognised |
|
||||
|--------------|-------------|--------------|---------------------------------------------|
|
||||
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
|
||||
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
|
||||
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
|
||||
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
|
||||
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
|
||||
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
|
||||
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
|
||||
|
||||
"Common opcodes" are the instructions shared by every architecture (`RET`,
|
||||
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
|
||||
@@ -124,10 +130,13 @@ can emit today is narrower, and a recognised but unencodable instruction is
|
||||
reported as an explicit error, never as a wrong byte.
|
||||
|
||||
The same measurement runs over GOROOT's whole assembly corpus:
|
||||
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
|
||||
(31.4 %) assembling for every target architecture today (files named for
|
||||
other Go ports are counted but never attempted), with the top failure
|
||||
reasons per architecture; the number moves with every release.
|
||||
`gasm audit-instructions --corpus` reports every real-code GOROOT assembly
|
||||
file (the tree without testdata) assembling for every target its build
|
||||
admits: 250 of 250, 100 %. Over the whole tree including testdata the
|
||||
measure is 271 of 322 attemptable (84.2 %); files named for other Go ports
|
||||
are counted but never attempted, and `//go:build` constraints decide which
|
||||
targets attempt a file at all, exactly as the build does. The number moves
|
||||
with every release.
|
||||
|
||||
### Validation status
|
||||
|
||||
@@ -142,7 +151,7 @@ actually been executed.
|
||||
|---|---|---|
|
||||
| Encoding: byte-for-byte against `go tool asm` | native hardware | native hardware (the toolchain cross-assembles any GOARCH on any host) |
|
||||
| Execution: JIT calls, ABI checks, differential fuzzing | native hardware | qemu-user emulation |
|
||||
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more |
|
||||
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64, riscv64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
|
||||
|
||||
Consequences, stated plainly. An emulator is a model of a CPU, not the
|
||||
CPU: instruction semantics are implemented in software and can differ
|
||||
@@ -157,6 +166,39 @@ been compiled and read, never executed. Its architecture-neutral units
|
||||
run under `go test ./...`, which the race workflow and a manual run
|
||||
perform; the default `just test` gate does not sweep `./debug/...`.
|
||||
|
||||
## The documentation goal
|
||||
|
||||
The toolkit is the primary goal. The secondary one is documentation: a
|
||||
specification of the Plan 9 assembly language and of the GOOBJ object
|
||||
format that is 100 % complete, detailed enough to implement against,
|
||||
and written to a professional standard. These are the two subjects this
|
||||
project works with every day, and they are the two for which no usable
|
||||
documentation exists.
|
||||
|
||||
Go documents the language on a single page, "A Quick Guide to Go's
|
||||
Assembler", which carries no section for loong64, one of the four
|
||||
architectures gasm supports, and covers a fraction of what each
|
||||
assembler accepts. What exists beyond it lives as comments inside the
|
||||
toolchain's internal source: per-architecture reference manuals for
|
||||
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
|
||||
maintainers rather than for an outside reader, and none at all for
|
||||
amd64. GOOBJ fares worst of all. The format that `go build` consumes
|
||||
has no specification anywhere: it is described by a comment in an
|
||||
internal package, it is not a stable interface, and it can change with
|
||||
any toolchain release.
|
||||
|
||||
The gap is therefore filled the only way it can be filled: by reverse
|
||||
engineering the toolchain itself, the same work the encoders already
|
||||
perform. Most of the documentation can come from nowhere else, and it
|
||||
is written as that knowledge is produced during development. It is
|
||||
verified the way the code is verified: an encoding documented here is
|
||||
one that differential tests against `go tool asm` confirm
|
||||
byte-for-byte, and a format field documented here is one the linker
|
||||
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
|
||||
specifies the object file format completely, and
|
||||
[docs/asm/README.md](docs/asm/README.md) opens the language reference
|
||||
with its common core. The per-architecture pages follow.
|
||||
|
||||
## Direction
|
||||
|
||||
The plan, in the order it is being worked:
|
||||
@@ -179,9 +221,11 @@ The plan, in the order it is being worked:
|
||||
toolchain itself does not support; through ELF, Plan 9 assembly becomes
|
||||
usable outside Go entirely.
|
||||
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
|
||||
architectures and is where the binary builds. FreeBSD follows: the
|
||||
JIT's executable-memory mapping and the ptrace debugger layer are the
|
||||
two pieces of porting work. Other unix systems may follow those two.
|
||||
architectures and is where the binary builds. FreeBSD follows on amd64,
|
||||
arm64 and riscv64: the JIT's executable-memory mapping and the ptrace
|
||||
debugger layer are ported (the debugger's live validation awaits a
|
||||
FreeBSD machine, as the validation status states). Other unix systems
|
||||
may follow those two.
|
||||
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
|
||||
No others are planned.
|
||||
|
||||
@@ -189,11 +233,11 @@ The plan, in the order it is being worked:
|
||||
|
||||
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
|
||||
linux/loong64 are on the
|
||||
[releases page](https://sourcedock.dev/petrbalvin/gasm-devkit/releases).
|
||||
[releases page](https://sourcedock.dev/petrbalvin/gasm-sdk/releases).
|
||||
From source (Go 1.27.1):
|
||||
|
||||
```sh
|
||||
go install sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm@latest
|
||||
go install sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm@latest
|
||||
```
|
||||
|
||||
Or from a repository checkout:
|
||||
@@ -282,6 +326,8 @@ recipe.
|
||||
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
|
||||
them
|
||||
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
|
||||
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
|
||||
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
|
||||
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
|
||||
- [CHANGELOG.md](CHANGELOG.md): release history
|
||||
|
||||
|
||||
+1
-1
@@ -7,7 +7,7 @@ releases do not receive them.
|
||||
|
||||
| Version | Supported |
|
||||
|---|---|
|
||||
| 0.34.0 | yes |
|
||||
| 0.35.0 | yes |
|
||||
| older releases | no |
|
||||
|
||||
## Reporting a vulnerability
|
||||
|
||||
+105
-7
@@ -5,9 +5,13 @@
|
||||
// toolchain's own assembler source. Go's Plan 9 assembler defines the exact,
|
||||
// complete set of mnemonics it accepts for each architecture in
|
||||
// $GOROOT/src/cmd/internal/obj/<arch>/anames.go; this tool extracts those
|
||||
// names so gasm-devkit supports every instruction the real assembler does,
|
||||
// names so gasm-sdk supports every instruction the real assembler does,
|
||||
// with no hand-maintained (and therefore inevitably incomplete) lists.
|
||||
//
|
||||
// The same data feeds the generated instruction appendices of the assembly
|
||||
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
|
||||
// cannot drift from the tables it documents.
|
||||
//
|
||||
// Usage (via the justfile):
|
||||
//
|
||||
// just gen
|
||||
@@ -26,9 +30,12 @@ import (
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
)
|
||||
|
||||
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
|
||||
// archDirs maps a gasm-sdk architecture name to its obj sub-directory.
|
||||
var archDirs = []struct {
|
||||
arch string
|
||||
sub string
|
||||
@@ -39,11 +46,30 @@ var archDirs = []struct {
|
||||
{"loong64", "loong64"},
|
||||
}
|
||||
|
||||
// docPages maps an architecture to its generated appendix in the language
|
||||
// reference. The amd64 page carries a per-mnemonic encodability column,
|
||||
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
|
||||
// other targets have no single cheap predicate, so their pages carry the
|
||||
// inventory and point at the live measurement instead.
|
||||
var docPages = []struct {
|
||||
arch arch.Arch
|
||||
title string
|
||||
file string
|
||||
anames string
|
||||
encodable bool
|
||||
}{
|
||||
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
|
||||
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
|
||||
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
|
||||
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
|
||||
}
|
||||
|
||||
func main() {
|
||||
goroot := strings.TrimSpace(runGoEnvGOROOT())
|
||||
if goroot == "" {
|
||||
fatal("could not determine GOROOT")
|
||||
}
|
||||
version := strings.TrimSpace(runGoEnv("GOVERSION"))
|
||||
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
|
||||
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
|
||||
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
|
||||
@@ -57,16 +83,24 @@ func main() {
|
||||
}
|
||||
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
|
||||
|
||||
names := map[string][]string{}
|
||||
for _, a := range archDirs {
|
||||
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
|
||||
names, err := extractInstrs(path)
|
||||
names[a.arch], err = extractInstrs(path)
|
||||
if err != nil {
|
||||
fatal("extract %s: %v", a.arch, err)
|
||||
}
|
||||
if err := writeGen(a.arch, a.sub, names); err != nil {
|
||||
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
|
||||
fatal("write %s: %v", a.arch, err)
|
||||
}
|
||||
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
|
||||
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
|
||||
}
|
||||
|
||||
for _, p := range docPages {
|
||||
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
|
||||
fatal("write %s: %v", p.file, err)
|
||||
}
|
||||
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -85,7 +119,7 @@ func filterCommon(names []string) []string {
|
||||
// writeCommon emits arch/common_gen.go.
|
||||
func writeCommon(names []string) error {
|
||||
var b strings.Builder
|
||||
b.WriteString("// Code generated by gasm-devkit _gen; DO NOT EDIT.\n")
|
||||
b.WriteString("// Code generated by gasm-sdk _gen; DO NOT EDIT.\n")
|
||||
b.WriteString("// Source: cmd/internal/obj/util.go from the Go toolchain.\n")
|
||||
b.WriteString("//\n")
|
||||
b.WriteString("// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)\n")
|
||||
@@ -156,7 +190,7 @@ func stringLit(elt ast.Expr) string {
|
||||
// writeGen emits arch/<arch>_gen.go.
|
||||
func writeGen(arch, sub string, names []string) error {
|
||||
var b strings.Builder
|
||||
b.WriteString("// Code generated by gasm-devkit _gen; DO NOT EDIT.\n")
|
||||
b.WriteString("// Code generated by gasm-sdk _gen; DO NOT EDIT.\n")
|
||||
b.WriteString("// Source: cmd/internal/obj/" + sub + "/anames.go from the Go toolchain.\n")
|
||||
b.WriteString("//\n")
|
||||
b.WriteString("// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)\n")
|
||||
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
|
||||
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
|
||||
}
|
||||
|
||||
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
|
||||
// the language reference for one architecture: every mnemonic the toolchain
|
||||
// accepts, with the curated summary where the architecture table carries one
|
||||
// and, on amd64, a per-mnemonic encodability column.
|
||||
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
|
||||
table := arch.ForArch(a)
|
||||
instrs := table.Instructions()
|
||||
|
||||
var b strings.Builder
|
||||
b.WriteString("# " + title + ": instruction inventory\n\n")
|
||||
b.WriteString("Generated by gasm-sdk's `_gen` from the Go toolchain's instruction table\n")
|
||||
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
|
||||
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
|
||||
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
|
||||
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
|
||||
|
||||
encodableCount := 0
|
||||
if encodable {
|
||||
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
|
||||
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
|
||||
b.WriteString("`gasm audit-instructions`.\n\n")
|
||||
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
|
||||
b.WriteString("|---|---|---|\n")
|
||||
for _, in := range instrs {
|
||||
ok := asm.Encodable(in.Name)
|
||||
if ok {
|
||||
encodableCount++
|
||||
}
|
||||
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
|
||||
}
|
||||
b.WriteString("\n")
|
||||
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
|
||||
} else {
|
||||
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
|
||||
b.WriteString("encodability is decided per operand shape, and the live measured\n")
|
||||
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
|
||||
b.WriteString("| Mnemonic | Notes |\n")
|
||||
b.WriteString("|---|---|\n")
|
||||
for _, in := range instrs {
|
||||
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
|
||||
}
|
||||
b.WriteString("\n")
|
||||
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
|
||||
}
|
||||
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
|
||||
}
|
||||
|
||||
// yesNo renders a boolean as the word the appendix tables use.
|
||||
func yesNo(v bool) string {
|
||||
if v {
|
||||
return "yes"
|
||||
}
|
||||
return "no"
|
||||
}
|
||||
|
||||
func runGoEnvGOROOT() string {
|
||||
out, err := exec.Command("go", "env", "GOROOT").Output()
|
||||
if err != nil {
|
||||
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
|
||||
return string(out)
|
||||
}
|
||||
|
||||
// runGoEnv runs `go env` for a single variable.
|
||||
func runGoEnv(name string) string {
|
||||
out, err := exec.Command("go", "env", name).Output()
|
||||
if err != nil {
|
||||
return ""
|
||||
}
|
||||
return string(out)
|
||||
}
|
||||
|
||||
func fatal(format string, args ...any) {
|
||||
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
|
||||
os.Exit(1)
|
||||
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
// Code generated by gasm-devkit _gen; DO NOT EDIT.
|
||||
// Code generated by gasm-sdk _gen; DO NOT EDIT.
|
||||
// Source: cmd/internal/obj/x86/anames.go from the Go toolchain.
|
||||
//
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
|
||||
+108
-1
@@ -1,4 +1,4 @@
|
||||
// Code generated by gasm-devkit _gen; DO NOT EDIT.
|
||||
// Code generated by gasm-sdk _gen; DO NOT EDIT.
|
||||
// Source: cmd/internal/obj/arm64/anames.go from the Go toolchain.
|
||||
//
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
|
||||
"REVW",
|
||||
"ROR",
|
||||
"RORW",
|
||||
"RPRFM",
|
||||
"SB",
|
||||
"SBC",
|
||||
"SBCS",
|
||||
"SBCSW",
|
||||
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
|
||||
"UXTH",
|
||||
"UXTHW",
|
||||
"UXTW",
|
||||
"VABS",
|
||||
"VADD",
|
||||
"VADDP",
|
||||
"VADDV",
|
||||
"VAND",
|
||||
"VBCAX",
|
||||
"VBIC",
|
||||
"VBIF",
|
||||
"VBIT",
|
||||
"VBSL",
|
||||
"VCLS",
|
||||
"VCLZ",
|
||||
"VCMEQ",
|
||||
"VCMGE",
|
||||
"VCMGT",
|
||||
"VCMHI",
|
||||
"VCMHS",
|
||||
"VCMLE",
|
||||
"VCMLT",
|
||||
"VCMTST",
|
||||
"VCNT",
|
||||
"VDUP",
|
||||
"VEOR",
|
||||
"VEOR3",
|
||||
"VEXT",
|
||||
"VFABS",
|
||||
"VFADD",
|
||||
"VFADDP",
|
||||
"VFCMEQ",
|
||||
"VFCMGE",
|
||||
"VFCMGT",
|
||||
"VFCMLE",
|
||||
"VFCMLT",
|
||||
"VFCVTL",
|
||||
"VFCVTL2",
|
||||
"VFCVTN",
|
||||
"VFCVTN2",
|
||||
"VFCVTZS",
|
||||
"VFCVTZU",
|
||||
"VFDIV",
|
||||
"VFMAX",
|
||||
"VFMAXNM",
|
||||
"VFMAXNMP",
|
||||
"VFMAXNMV",
|
||||
"VFMAXP",
|
||||
"VFMAXV",
|
||||
"VFMIN",
|
||||
"VFMINNM",
|
||||
"VFMINNMP",
|
||||
"VFMINNMV",
|
||||
"VFMINP",
|
||||
"VFMINV",
|
||||
"VFMLA",
|
||||
"VFMLS",
|
||||
"VFMUL",
|
||||
"VFNEG",
|
||||
"VFRINTM",
|
||||
"VFRINTN",
|
||||
"VFRINTP",
|
||||
"VFRINTZ",
|
||||
"VFSQRT",
|
||||
"VFSUB",
|
||||
"VLD1",
|
||||
"VLD1R",
|
||||
"VLD2",
|
||||
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
|
||||
"VLD3R",
|
||||
"VLD4",
|
||||
"VLD4R",
|
||||
"VMLA",
|
||||
"VMLS",
|
||||
"VMOV",
|
||||
"VMOVD",
|
||||
"VMOVI",
|
||||
"VMOVQ",
|
||||
"VMOVS",
|
||||
"VMUL",
|
||||
"VNEG",
|
||||
"VNOT",
|
||||
"VORN",
|
||||
"VORR",
|
||||
"VPMULL",
|
||||
"VPMULL2",
|
||||
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
|
||||
"VREV16",
|
||||
"VREV32",
|
||||
"VREV64",
|
||||
"VSCVTF",
|
||||
"VSHADD",
|
||||
"VSHL",
|
||||
"VSHRN",
|
||||
"VSHRN2",
|
||||
"VSLI",
|
||||
"VSMAX",
|
||||
"VSMAXP",
|
||||
"VSMAXV",
|
||||
"VSMIN",
|
||||
"VSMINP",
|
||||
"VSMINV",
|
||||
"VSMLAL",
|
||||
"VSMLAL2",
|
||||
"VSMLSL",
|
||||
"VSMLSL2",
|
||||
"VSMULL",
|
||||
"VSMULL2",
|
||||
"VSQABS",
|
||||
"VSQADD",
|
||||
"VSQNEG",
|
||||
"VSQSHL",
|
||||
"VSQSUB",
|
||||
"VSQXTN",
|
||||
"VSQXTN2",
|
||||
"VSQXTUN",
|
||||
"VSQXTUN2",
|
||||
"VSRHADD",
|
||||
"VSRI",
|
||||
"VSRSHR",
|
||||
"VSSHL",
|
||||
"VSSHLL",
|
||||
"VSSHLL2",
|
||||
"VSSHR",
|
||||
"VST1",
|
||||
"VST2",
|
||||
"VST3",
|
||||
"VST4",
|
||||
"VSUB",
|
||||
"VSXTL",
|
||||
"VSXTL2",
|
||||
"VTBL",
|
||||
"VTBX",
|
||||
"VTRN1",
|
||||
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
|
||||
"VUADDLV",
|
||||
"VUADDW",
|
||||
"VUADDW2",
|
||||
"VUCVTF",
|
||||
"VUHADD",
|
||||
"VUMAX",
|
||||
"VUMAXP",
|
||||
"VUMAXV",
|
||||
"VUMIN",
|
||||
"VUMINP",
|
||||
"VUMINV",
|
||||
"VUMLAL",
|
||||
"VUMLAL2",
|
||||
"VUMLSL",
|
||||
"VUMLSL2",
|
||||
"VUMULL",
|
||||
"VUMULL2",
|
||||
"VUQADD",
|
||||
"VUQSHL",
|
||||
"VUQSUB",
|
||||
"VUQXTN",
|
||||
"VUQXTN2",
|
||||
"VURHADD",
|
||||
"VUSHL",
|
||||
"VUSHLL",
|
||||
"VUSHLL2",
|
||||
"VUSHR",
|
||||
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
|
||||
"VUZP1",
|
||||
"VUZP2",
|
||||
"VXAR",
|
||||
"VXTN",
|
||||
"VXTN2",
|
||||
"VZIP1",
|
||||
"VZIP2",
|
||||
"WFE",
|
||||
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
// Code generated by gasm-devkit _gen; DO NOT EDIT.
|
||||
// Code generated by gasm-sdk _gen; DO NOT EDIT.
|
||||
// Source: cmd/internal/obj/util.go from the Go toolchain.
|
||||
//
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
|
||||
+10
-1
@@ -1,4 +1,4 @@
|
||||
// Code generated by gasm-devkit _gen; DO NOT EDIT.
|
||||
// Code generated by gasm-sdk _gen; DO NOT EDIT.
|
||||
// Source: cmd/internal/obj/loong64/anames.go from the Go toolchain.
|
||||
//
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
|
||||
"FNMADDF",
|
||||
"FNMSUBD",
|
||||
"FNMSUBF",
|
||||
"FRINTD",
|
||||
"FRINTF",
|
||||
"FSCALEBD",
|
||||
"FSCALEBF",
|
||||
"FSEL",
|
||||
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
|
||||
"FTINTWF",
|
||||
"JIRL",
|
||||
"LL",
|
||||
"LLACQV",
|
||||
"LLACQW",
|
||||
"LLV",
|
||||
"LLW",
|
||||
"LU12IW",
|
||||
"LU32ID",
|
||||
"LU52ID",
|
||||
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
|
||||
"ROTR",
|
||||
"ROTRV",
|
||||
"SC",
|
||||
"SCQ",
|
||||
"SCRELV",
|
||||
"SCRELW",
|
||||
"SCV",
|
||||
"SCW",
|
||||
"SGT",
|
||||
"SGTU",
|
||||
"SLL",
|
||||
|
||||
+32
-1
@@ -1,4 +1,4 @@
|
||||
// Code generated by gasm-devkit _gen; DO NOT EDIT.
|
||||
// Code generated by gasm-sdk _gen; DO NOT EDIT.
|
||||
// Source: cmd/internal/obj/riscv/anames.go from the Go toolchain.
|
||||
//
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
|
||||
"CLD",
|
||||
"CLDSP",
|
||||
"CLI",
|
||||
"CLMUL",
|
||||
"CLMULH",
|
||||
"CLMULR",
|
||||
"CLUI",
|
||||
"CLW",
|
||||
"CLWSP",
|
||||
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
|
||||
"CSDSP",
|
||||
"CSLLI",
|
||||
"CSRAI",
|
||||
"CSRC",
|
||||
"CSRCI",
|
||||
"CSRLI",
|
||||
"CSRR",
|
||||
"CSRRC",
|
||||
"CSRRCI",
|
||||
"CSRRS",
|
||||
"CSRRSI",
|
||||
"CSRRW",
|
||||
"CSRRWI",
|
||||
"CSRS",
|
||||
"CSRSI",
|
||||
"CSRW",
|
||||
"CSRWI",
|
||||
"CSUB",
|
||||
"CSUBW",
|
||||
"CSW",
|
||||
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
|
||||
"ORCB",
|
||||
"ORI",
|
||||
"ORN",
|
||||
"PAUSE",
|
||||
"RDCYCLE",
|
||||
"RDINSTRET",
|
||||
"RDTIME",
|
||||
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
|
||||
"VADDVI",
|
||||
"VADDVV",
|
||||
"VADDVX",
|
||||
"VANDNVV",
|
||||
"VANDNVX",
|
||||
"VANDVI",
|
||||
"VANDVV",
|
||||
"VANDVX",
|
||||
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
|
||||
"VASUBUVX",
|
||||
"VASUBVV",
|
||||
"VASUBVX",
|
||||
"VBREV8V",
|
||||
"VBREVV",
|
||||
"VCLMULHVV",
|
||||
"VCLMULHVX",
|
||||
"VCLMULVV",
|
||||
"VCLMULVX",
|
||||
"VCLZV",
|
||||
"VCOMPRESSVM",
|
||||
"VCPOPM",
|
||||
"VCPOPV",
|
||||
"VCTZV",
|
||||
"VDIVUVV",
|
||||
"VDIVUVX",
|
||||
"VDIVVV",
|
||||
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
|
||||
"VREMUVX",
|
||||
"VREMVV",
|
||||
"VREMVX",
|
||||
"VREV8V",
|
||||
"VRGATHEREI16VV",
|
||||
"VRGATHERVI",
|
||||
"VRGATHERVV",
|
||||
"VRGATHERVX",
|
||||
"VROLVV",
|
||||
"VROLVX",
|
||||
"VRORVI",
|
||||
"VRORVV",
|
||||
"VRORVX",
|
||||
"VRSUBVI",
|
||||
"VRSUBVX",
|
||||
"VS1RV",
|
||||
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
|
||||
"VWMULVX",
|
||||
"VWREDSUMUVS",
|
||||
"VWREDSUMVS",
|
||||
"VWSLLVI",
|
||||
"VWSLLVV",
|
||||
"VWSLLVX",
|
||||
"VWSUBUVV",
|
||||
"VWSUBUVX",
|
||||
"VWSUBUWV",
|
||||
|
||||
@@ -11,7 +11,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestGOObjectAARCH64Structure checks the basic structure of the emitted
|
||||
|
||||
+38
-15
@@ -9,7 +9,7 @@ import (
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// assembleARM64 assembles an AArch64 (arm64) TEXT function body into machine
|
||||
@@ -637,6 +637,20 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
|
||||
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
|
||||
}
|
||||
|
||||
// The bare spelling BL R9 is the same indirect branch: the parser reads
|
||||
// a bare identifier as a symbol, and one named for a register is an
|
||||
// indirect branch through it, which the toolchain accepts alongside the
|
||||
// parenthesised form (BL (R3) and BL R3 both encode BLR R3).
|
||||
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" && op.Addr.Base == "" && op.Addr.Index == "" {
|
||||
if rn := arm64RegNum(op.Addr.Sym.Name); rn >= 0 {
|
||||
opc := uint32(0) // BR
|
||||
if link {
|
||||
opc = 1 // BLR
|
||||
}
|
||||
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
|
||||
}
|
||||
}
|
||||
|
||||
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
|
||||
// relocation (R_CALLARM64 either way).
|
||||
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
|
||||
@@ -1454,6 +1468,15 @@ func encodeARM64Mov(instr *ast.Instr, mnem string, wb string, fi arm64FrameInfo,
|
||||
}
|
||||
return encodeARM64SBAddr(src.Imm.Sym, rd, relocs), nil
|
||||
}
|
||||
// Immediate → memory: only storing zero is encodable (the ZR
|
||||
// register); the toolchain rejects any other immediate-to-memory
|
||||
// combination ("illegal combination").
|
||||
if isMemOperand(dst) {
|
||||
if arm64Imm64(src) != 0 {
|
||||
return nil, fmt.Errorf("%s: illegal combination: an immediate store must be zero", mnem)
|
||||
}
|
||||
return encodeARM64MemOp(mnem, dst, 31, false, fi, "")
|
||||
}
|
||||
rd := arm64RegNum(operandRegName(dst))
|
||||
if rd < 0 {
|
||||
return nil, fmt.Errorf("%s $imm: invalid destination register", mnem)
|
||||
@@ -3046,29 +3069,29 @@ func encodeARM64MoveWide(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte
|
||||
// base, so MOVZ and MOVN come along for free.
|
||||
opc := baseOp >> 29 & 3
|
||||
sf := baseOp >> 31 & 1
|
||||
v := arm64Imm64(ops[0])
|
||||
if v < 0 {
|
||||
return nil, fmt.Errorf("%s: negative immediate %d", mnem, v)
|
||||
// The toolchain's optab case 33, shared by the whole family in both
|
||||
// widths: the immediate is one unsigned 64-bit pattern (a high-lane
|
||||
// constant such as $(40000<<48) arrives negative through int64
|
||||
// folding), it must occupy exactly one 16-bit lane, zero is rejected,
|
||||
// and the W forms cannot reach the top half.
|
||||
u := uint64(arm64Imm64(ops[0]))
|
||||
if u == 0 {
|
||||
return nil, fmt.Errorf("%s: zero immediate cannot be handled", mnem)
|
||||
}
|
||||
hw := -1
|
||||
for i := range 4 {
|
||||
if v>>(uint(i)*16)&0xFFFF != 0 {
|
||||
hw = i
|
||||
for lane := range 4 {
|
||||
if u&^(uint64(0xFFFF)<<(lane*16)) == 0 {
|
||||
hw = lane
|
||||
break
|
||||
}
|
||||
}
|
||||
if hw < 0 {
|
||||
hw = 0 // zero: every chunk is zero, hw = 0 carries it
|
||||
}
|
||||
for i := hw + 1; i < 4; i++ {
|
||||
if v>>(uint(i)*16)&0xFFFF != 0 {
|
||||
return nil, fmt.Errorf("%s: immediate %d does not fit one 16-bit chunk", mnem, v)
|
||||
}
|
||||
return nil, fmt.Errorf("%s: immediate %#x does not fit one 16-bit chunk", mnem, u)
|
||||
}
|
||||
if sf == 0 && hw > 1 {
|
||||
return nil, fmt.Errorf("%s: immediate %d out of range for the 32-bit form", mnem, v)
|
||||
return nil, fmt.Errorf("%s: immediate %#x out of range for the 32-bit form", mnem, u)
|
||||
}
|
||||
return a64wordLE(a64MoveWide(sf, opc, uint32(hw), uint32(v>>uint(hw*16)&0xFFFF), uint32(rd))), nil
|
||||
return a64wordLE(a64MoveWide(sf, opc, uint32(hw), uint32(u>>uint(hw*16)&0xFFFF), uint32(rd))), nil
|
||||
}
|
||||
|
||||
// ---- Bitfield/EXTR encoding ----
|
||||
|
||||
+1
-1
@@ -33,7 +33,7 @@ import (
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// arm64RegNum returns the 5-bit register number for an AArch64 register name:
|
||||
|
||||
@@ -7,8 +7,8 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
func TestArm64LDRSTREncoding(t *testing.T) {
|
||||
@@ -1051,6 +1051,41 @@ func TestArm64MOVK(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64MOVKHighLane pins the shifted high-lane immediate the arm64 test
|
||||
// kernels write: $(40000<<48) folds to a negative int64, and the toolchain
|
||||
// reads the value as an unsigned 64-bit pattern when it picks the lane.
|
||||
func TestArm64MOVKHighLane(t *testing.T) {
|
||||
got := arm64Words(t, "\tMOVK $(40000<<48), R0\n\tMOVK $0x9c40000000000000, R1\n")
|
||||
want := []uint32{
|
||||
0xf2f38800, // MOVK $(40000<<48), R0 (go tool asm: f2f38800)
|
||||
0xf2f38801, // MOVK hw=3
|
||||
0xd65f03c0,
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64MoveWideZeroImmediate pins the toolchain's rejection of a zero
|
||||
// immediate in the move-wide family (optab case 33: "zero shifts cannot be
|
||||
// handled"): every lane is zero, so no hw field can carry it.
|
||||
func TestArm64MoveWideZeroImmediate(t *testing.T) {
|
||||
for _, mnem := range []string{"MOVK", "MOVZ", "MOVN"} {
|
||||
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\t"+mnem+" $0, R0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("%s: parse: %v", mnem, errs)
|
||||
}
|
||||
if _, err := AssembleFileARM64(f); err == nil {
|
||||
t.Errorf("%s $0: expected error, got nil", mnem)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64LoadImm64 tests 64-bit immediate loading.
|
||||
func TestArm64LoadImm64(t *testing.T) {
|
||||
src := `#include "textflag.h"
|
||||
|
||||
+1
-1
@@ -53,7 +53,7 @@ package asm
|
||||
import (
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// arm64FrameInfo holds the frame layout derived from a TEXT directive.
|
||||
|
||||
@@ -7,7 +7,7 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// parseArm64File is a helper assembling one arm64 source file.
|
||||
|
||||
+500
-44
@@ -5,9 +5,10 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// Assemble encodes the body of a TEXT function into x86-64 machine code,
|
||||
@@ -28,7 +29,7 @@ import (
|
||||
// emitted: the bytes match go tool asm only for NOSPLIT functions or
|
||||
// zero-frame leaves, where the toolchain emits no guard either.
|
||||
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
|
||||
code, _, labels, _, _, err := assemble(t, nil)
|
||||
code, _, labels, _, _, _, err := assemble(t, nil)
|
||||
return code, labels, err
|
||||
}
|
||||
|
||||
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
|
||||
// rejects SB operands outright (single-function assembly cannot resolve
|
||||
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
|
||||
// file defines is recorded as an external relocation instead of failing
|
||||
// the object-file emitters resolve it at link time.
|
||||
// the object-file emitters resolve it at link time. goos selects the TLS
|
||||
// access form: the empty default behaves as linux.
|
||||
type linkInfo struct {
|
||||
symbols map[string]bool
|
||||
allowExternal bool
|
||||
goos string
|
||||
}
|
||||
|
||||
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
|
||||
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
|
||||
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
|
||||
// keep the two-instruction form; shared linux does too, which gasm's raw
|
||||
// path does not model and therefore does not select.
|
||||
func (l *linkInfo) tlsOneInsn() bool {
|
||||
switch l.goos {
|
||||
case "", "linux", "freebsd":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// sbPatch is a function-relative static-symbol relocation: the disp32 field
|
||||
@@ -66,9 +82,9 @@ type spadjStep struct {
|
||||
// assemble encodes a TEXT body, returning the machine code, the static-symbol
|
||||
// patch sites (for the file-level layout to resolve), the label table and the
|
||||
// stack-adjustment boundaries.
|
||||
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
|
||||
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
|
||||
if err := checkAdjspBalance(t); err != nil {
|
||||
return nil, nil, nil, nil, nil, err
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
fi := computeFrame(t)
|
||||
chain := jumpChain(t)
|
||||
@@ -85,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
// outgrows the short form.
|
||||
long := make([]bool, len(t.Body))
|
||||
sizes := make([]int, len(t.Body))
|
||||
numTargets := make([]int, len(t.Body))
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
offsets := map[string]int{}
|
||||
pcs := make([]int, len(t.Body))
|
||||
var guardJBlong, guardJBElong, moreJMPlong bool
|
||||
poolSeen := map[string]bool{}
|
||||
var poolList []floatPoolEntry
|
||||
for {
|
||||
guard := fi.guardLen(guardJBlong, guardJBElong)
|
||||
pos := guard + len(fi.prologue)
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
idxAtPc := map[int]int{}
|
||||
for i, stmt := range t.Body {
|
||||
switch s := stmt.(type) {
|
||||
case *ast.Label:
|
||||
offsets[s.Name.Text] = pos
|
||||
case *ast.Instr:
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
|
||||
// The alignment pseudo-statement: its size is the
|
||||
// padding to the next boundary at this very position,
|
||||
// filled with NOPs at emission.
|
||||
pad, err := pcAlignPad(pcAlignValue(s), pos)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
|
||||
}
|
||||
sizes[i] = pad
|
||||
pcs[i] = pos
|
||||
pos += pad
|
||||
continue
|
||||
}
|
||||
sz, err := instrSize(s, fi, long[i], link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
}
|
||||
sizes[i] = sz
|
||||
pcs[i] = pos
|
||||
idxAtPc[pos] = i
|
||||
pos += sz
|
||||
}
|
||||
}
|
||||
bodyLen := pos - (guard + len(fi.prologue))
|
||||
// Expand any short jump whose displacement no longer fits rel8.
|
||||
changed := false
|
||||
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
|
||||
// emission pass reads the same table after the loop converges. A
|
||||
// target that is itself an unconditional local JMP is chased to the
|
||||
// ultimate target: the toolchain's brloop pass collapses branch-to-
|
||||
// branch chains before it encodes, so matching its bytes requires
|
||||
// the same redirection.
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if len(s.Operands) == 1 {
|
||||
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
|
||||
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
|
||||
numTargets[i] = target
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for i := range numTargets {
|
||||
if numTargets[i] < 0 {
|
||||
continue
|
||||
}
|
||||
tgt := numTargets[i]
|
||||
for hop := 0; hop < len(t.Body); hop++ {
|
||||
idx, ok := idxAtPc[tgt]
|
||||
if !ok {
|
||||
break
|
||||
}
|
||||
in, ok := t.Body[idx].(*ast.Instr)
|
||||
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
|
||||
break
|
||||
}
|
||||
if name, isLabel := labelName(in.Operands[0]); isLabel {
|
||||
tgt = offsets[resolve(name)]
|
||||
continue
|
||||
}
|
||||
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
|
||||
next, okT := pcJumpTarget(t, idx, n, pcs)
|
||||
if !okT {
|
||||
break
|
||||
}
|
||||
tgt = next
|
||||
continue
|
||||
}
|
||||
break // JMP through a register or memory: the chain ends
|
||||
}
|
||||
numTargets[i] = tgt
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if numTargets[i] >= 0 && !long[i] {
|
||||
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
|
||||
if !fits8(rel) {
|
||||
long[i] = true
|
||||
changed = true
|
||||
}
|
||||
}
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
@@ -229,12 +334,18 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
spadjStep{pos + epi, 0},
|
||||
)
|
||||
}
|
||||
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
|
||||
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
}
|
||||
for _, entry := range pool {
|
||||
if !poolSeen[entry.name] {
|
||||
poolSeen[entry.name] = true
|
||||
poolList = append(poolList, entry)
|
||||
}
|
||||
}
|
||||
if len(code) != sizes[i] {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
|
||||
}
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
|
||||
for k := range ps {
|
||||
@@ -272,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
pos += len(suffix)
|
||||
}
|
||||
_ = pos
|
||||
return out, patches, offsets, steps, lines, nil
|
||||
return out, patches, offsets, steps, lines, poolList, nil
|
||||
}
|
||||
|
||||
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
|
||||
@@ -419,6 +530,99 @@ func computeFrame(t *ast.Text) frameInfo {
|
||||
return fi
|
||||
}
|
||||
|
||||
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
|
||||
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
|
||||
// the second instruction boundary after the branch.
|
||||
func pcJumpOffset(op *ast.Operand) (int, bool) {
|
||||
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
|
||||
return 0, false
|
||||
}
|
||||
return int(op.Addr.Offset), true
|
||||
}
|
||||
|
||||
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
|
||||
// instruction statements after the jump itself (N = 0 is the jump's own
|
||||
// address, the classic park loop), and the target is the start of the Nth
|
||||
// one. A negative N counts the same way backwards, before the jump: the
|
||||
// exit loops write JMP -3(PC) to land three instructions earlier. Labels
|
||||
// count not, in either direction. It reports false when the count runs
|
||||
// past the end of the function, or before its first instruction.
|
||||
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
|
||||
if n == 0 {
|
||||
return pcs[j], true
|
||||
}
|
||||
if n < 0 {
|
||||
seen := 0
|
||||
for k := j - 1; k >= 0; k-- {
|
||||
if _, ok := t.Body[k].(*ast.Instr); !ok {
|
||||
continue
|
||||
}
|
||||
seen--
|
||||
if seen == n {
|
||||
return pcs[k], true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
seen := 0
|
||||
for k := j + 1; k < len(t.Body); k++ {
|
||||
if _, ok := t.Body[k].(*ast.Instr); !ok {
|
||||
continue
|
||||
}
|
||||
seen++
|
||||
if seen == n {
|
||||
return pcs[k], true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
|
||||
// toolchain's asm6.go nop table); longer padding repeats the largest that
|
||||
// fits, greedy from the end.
|
||||
var x86Nops = [][]byte{
|
||||
{0x90},
|
||||
{0x66, 0x90},
|
||||
{0x0F, 0x1F, 0x00},
|
||||
{0x0F, 0x1F, 0x40, 0x00},
|
||||
{0x0F, 0x1F, 0x44, 0x00, 0x00},
|
||||
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
|
||||
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
|
||||
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
|
||||
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
|
||||
}
|
||||
|
||||
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
|
||||
// the toolchain's fillnop.
|
||||
func fillNOPs(p []byte) {
|
||||
for len(p) > 0 {
|
||||
m := min(len(p), len(x86Nops))
|
||||
copy(p[:m], x86Nops[m-1])
|
||||
p = p[m:]
|
||||
}
|
||||
}
|
||||
|
||||
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
|
||||
// alignment must be a power of two in [8, 2048] and the padding runs to the
|
||||
// next boundary (zero when the position is already aligned).
|
||||
func pcAlignPad(align, pos int) (int, error) {
|
||||
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
|
||||
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
|
||||
}
|
||||
if lob := pos & (align - 1); lob != 0 {
|
||||
return align - lob, nil
|
||||
}
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
// pcAlignValue reads a PCALIGN statement's alignment operand.
|
||||
func pcAlignValue(s *ast.Instr) int {
|
||||
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
|
||||
return int(s.Operands[0].Imm.Val)
|
||||
}
|
||||
return 0 // rejected by pcAlignPad's range check
|
||||
}
|
||||
|
||||
// hasCall reports whether the function body contains a CALL instruction.
|
||||
func hasCall(t *ast.Text) bool {
|
||||
for _, stmt := range t.Body {
|
||||
@@ -593,7 +797,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
|
||||
}
|
||||
return jumpSize(mnem, long), nil
|
||||
}
|
||||
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
|
||||
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
@@ -628,9 +832,21 @@ func jumpSize(mnem string, long bool) int {
|
||||
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
|
||||
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
|
||||
// jump label through the jump-to-jump chain before the offset lookup.
|
||||
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
mnem := strings.ToUpper(s.Mnemonic.Text)
|
||||
|
||||
if mnem == "PCALIGN" {
|
||||
// The layout pass already accounted the padding; emit the same
|
||||
// amount of NOP bytes for the statement's own position.
|
||||
pad, err := pcAlignPad(pcAlignValue(s), pc)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
out := make([]byte, pad)
|
||||
fillNOPs(out)
|
||||
return out, nil, nil, nil
|
||||
}
|
||||
|
||||
var prefix []byte
|
||||
if mnem == "RET" && fi.useFP {
|
||||
prefix = fi.epilogue
|
||||
@@ -638,6 +854,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
|
||||
var code []byte
|
||||
var ps []sbPatch
|
||||
var pool []floatPoolEntry
|
||||
var err error
|
||||
if isJumpMnemonic(mnem) {
|
||||
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
|
||||
@@ -646,7 +863,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
// or the linker.
|
||||
code, ps, err = encodeSBCall(s, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
for i := range ps {
|
||||
ps[i].kind = RelCall
|
||||
@@ -656,23 +873,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
ps[i].off += body
|
||||
ps[i].after = body + len(code)
|
||||
}
|
||||
return append(prefix, code...), ps, nil
|
||||
return append(prefix, code...), ps, nil, nil
|
||||
}
|
||||
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
|
||||
// JMP/CALL through a register or memory: no relocation and no
|
||||
// label to resolve, the operand fully determines the bytes.
|
||||
code, err = encodeIndirectJump(s, mnem)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
return append(prefix, code...), nil, nil
|
||||
return append(prefix, code...), nil, nil, nil
|
||||
}
|
||||
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
|
||||
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
|
||||
} else {
|
||||
code, ps, err = encodeNormal(s, fi, link)
|
||||
code, ps, pool, err = encodeNormal(s, fi, link)
|
||||
}
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
// Anchor the patch fields at function-relative positions: off indexes the
|
||||
// disp32 field, after is the address just past the instruction.
|
||||
@@ -681,49 +898,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
ps[i].off += body
|
||||
ps[i].after = body + len(code)
|
||||
}
|
||||
return append(prefix, code...), ps, nil
|
||||
return append(prefix, code...), ps, pool, nil
|
||||
}
|
||||
|
||||
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
|
||||
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
|
||||
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
|
||||
code, err := encodeBookkeeping(mnemUpper, s)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
return code, nil, nil, nil
|
||||
}
|
||||
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
|
||||
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
|
||||
// never as a 64-bit absolute immediate (verified against go tool asm).
|
||||
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
|
||||
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
|
||||
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
|
||||
s.Operands[0].Imm.Sym.Pseudo == "SB" {
|
||||
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
|
||||
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
e := &enc{}
|
||||
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
|
||||
}
|
||||
return e.out, ps, nil, nil
|
||||
}
|
||||
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
|
||||
// it out on the one-instruction TLS systems (linux and freebsd, not
|
||||
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
|
||||
// the macro GOROOT's go_tls.h defines, expands to exactly this
|
||||
// statement, and the toolchain's pairing pass removes it whenever the
|
||||
// following instruction's (TLS*1) index folds.
|
||||
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
|
||||
return encodeTLSBaseLoad(s, fi, link)
|
||||
}
|
||||
_, size := splitSize(mnemUpper)
|
||||
if size == 0 {
|
||||
size = 8
|
||||
}
|
||||
ops := make([]Operand, len(s.Operands))
|
||||
for i, op := range s.Operands {
|
||||
o, err := operandFromAST(op, size, fi, link)
|
||||
o, err := operandFromAST(mnemUpper, op, size, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ops[i] = o
|
||||
}
|
||||
e := &enc{}
|
||||
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
|
||||
if p.tls {
|
||||
ps[i].kind = RelTLSLE
|
||||
}
|
||||
}
|
||||
return e.out, ps, nil
|
||||
return e.out, ps, e.floatPoolList(), nil
|
||||
}
|
||||
|
||||
// isBareTLS reports whether the operand is the bare TLS pseudo-register
|
||||
// load source, the expansion of go_tls.h's get_tls(r) macro.
|
||||
func isBareTLS(op *ast.Operand) bool {
|
||||
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
|
||||
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
|
||||
op.Addr.Base == "" && op.Addr.Index == ""
|
||||
}
|
||||
|
||||
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
|
||||
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
|
||||
// statement nops out: the following (TLS*1) access folds to a direct
|
||||
// segment-absolute load. The two-instruction systems keep the segment load,
|
||||
// nine bytes with the R_TLSLE patch site at the disp32.
|
||||
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
|
||||
if size == 0 {
|
||||
size = 8
|
||||
}
|
||||
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
reg, ok := dst.(Reg)
|
||||
if !ok || reg.isVec() {
|
||||
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
|
||||
}
|
||||
if link == nil || link.tlsOneInsn() {
|
||||
return nil, nil, nil, nil // noped out
|
||||
}
|
||||
seg := byte(0x64) // FS
|
||||
if link.goos == "windows" {
|
||||
seg = 0x65 // GS
|
||||
}
|
||||
e := &enc{}
|
||||
i := &instr{
|
||||
prefix: seg,
|
||||
rexW: size == 8,
|
||||
rexR: reg.idx >= 8,
|
||||
opcode: []byte{0x8B},
|
||||
modrm: 0x04 | (reg.idx&7)<<3,
|
||||
sib: 0x25,
|
||||
disp: le32(0),
|
||||
tls: true,
|
||||
}
|
||||
if err := e.emit(i); err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
|
||||
}
|
||||
return e.out, ps, nil, nil
|
||||
}
|
||||
|
||||
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
|
||||
// level, before operand conversion: the toolchain's shapes are FUNCDATA
|
||||
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
|
||||
// function body. The symbol reference must not run through the SB-operand
|
||||
// path, which demands file-level resolution the statement never needs.
|
||||
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
|
||||
if len(s.Operands) != 2 {
|
||||
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
|
||||
}
|
||||
a, b := s.Operands[0], s.Operands[1]
|
||||
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
|
||||
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
|
||||
}
|
||||
switch upper {
|
||||
case "FUNCDATA":
|
||||
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
|
||||
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
|
||||
}
|
||||
case "PCDATA":
|
||||
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
|
||||
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
|
||||
}
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
|
||||
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
|
||||
// target label, in the short (rel8) or long (rel32) form.
|
||||
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
|
||||
// target label or from a numeric ±N(PC) instruction count, in the short
|
||||
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
|
||||
// numeric operand, negative when the operand is not one.
|
||||
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
|
||||
if len(s.Operands) != 1 {
|
||||
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
|
||||
}
|
||||
name, ok := labelName(s.Operands[0])
|
||||
if !ok {
|
||||
name, isLabel := labelName(s.Operands[0])
|
||||
if !isLabel && numTarget < 0 {
|
||||
return nil, fmt.Errorf("jump target must be a local label")
|
||||
}
|
||||
if resolve != nil && mnem != "CALL" {
|
||||
name = resolve(name)
|
||||
}
|
||||
target, ok := offsets[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q", name)
|
||||
var target int
|
||||
if isLabel {
|
||||
if resolve != nil && mnem != "CALL" {
|
||||
name = resolve(name)
|
||||
}
|
||||
t, ok := offsets[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q", name)
|
||||
}
|
||||
target = t
|
||||
} else {
|
||||
target = numTarget
|
||||
}
|
||||
rel := int64(target - (pc + jumpSize(mnem, long)))
|
||||
|
||||
@@ -756,7 +1107,7 @@ func isSBCall(s *ast.Instr) bool {
|
||||
|
||||
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
|
||||
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
|
||||
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
}
|
||||
@@ -797,6 +1148,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
|
||||
return false
|
||||
}
|
||||
a := s.Operands[0].Addr
|
||||
// ±N(PC) is the numeric relative form, the PC counts instructions from
|
||||
// the branch: relative, not indirect.
|
||||
if a.Base == "PC" || a.Index == "PC" {
|
||||
return false
|
||||
}
|
||||
if a.Base != "" || a.Index != "" {
|
||||
return true
|
||||
}
|
||||
@@ -813,7 +1169,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
|
||||
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
|
||||
ops := make([]Operand, len(s.Operands))
|
||||
for i, op := range s.Operands {
|
||||
o, err := operandFromAST(op, 8, frameInfo{}, nil)
|
||||
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -830,8 +1186,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
|
||||
var spReg = Reg{idx: 4, size: 8}
|
||||
|
||||
// operandFromAST converts a parsed operand into an encoder Operand, applying
|
||||
// the frame translation to FP/SP pseudo-register operands.
|
||||
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
|
||||
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
|
||||
// instruction's upper-case mnemonic, which the floating-point immediate gate
|
||||
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
|
||||
// accept one.
|
||||
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
|
||||
switch op.Kind {
|
||||
case ast.OpImmediate:
|
||||
if op.Imm.HasVal {
|
||||
@@ -841,17 +1200,43 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return Imm(v), nil
|
||||
}
|
||||
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
|
||||
// $(-1.0) spelling (the constant-expression folder only folds
|
||||
// integers, so that shape arrives with an empty Immediate and only
|
||||
// the raw spelling carries the value). The toolchain rewrites it
|
||||
// into a pooled-constant read on the SSE scalar paths and rejects
|
||||
// it everywhere else.
|
||||
if text, neg, ok := floatImmText(op); ok {
|
||||
if !sseFloatImm[mnemUpper] {
|
||||
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
|
||||
}
|
||||
return FloatImm{Text: text, Neg: neg}, nil
|
||||
}
|
||||
return nil, fmt.Errorf("non-integer immediate not supported")
|
||||
|
||||
case ast.OpAddr:
|
||||
a := op.Addr
|
||||
|
||||
// A bracketed register range, [Z0-Z3]: the four-register source of
|
||||
// the 4FMAPS/4VNNIW families. The EVEX quad-register emit path
|
||||
// needs an encoder operand of its own, so the shape stays a named
|
||||
// gap rather than an encoding.
|
||||
// the 4FMAPS/4VNNIW families. The range must span four consecutive
|
||||
// same-width vector registers, exactly what the toolchain's parser
|
||||
// takes; the EVEX quad-register emit path reads the low end.
|
||||
if a.Range != nil {
|
||||
return nil, fmt.Errorf("register range %q needs quad-register encoder support", op.Raw)
|
||||
lo, ok := ParseReg(a.Range.Lo)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
|
||||
}
|
||||
hi, ok := ParseReg(a.Range.Hi)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
|
||||
}
|
||||
if !lo.isVec() || lo.size != hi.size {
|
||||
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
|
||||
}
|
||||
if hi.idx != lo.idx+3 {
|
||||
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
|
||||
}
|
||||
return RegList{Lo: lo, Hi: hi}, nil
|
||||
}
|
||||
|
||||
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
|
||||
@@ -885,12 +1270,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
|
||||
// Memory with a real base register: (base), off(base), (base)(index*scale).
|
||||
if a.Base != "" {
|
||||
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
|
||||
// spellings. The segment override prefixes a disp32 absolute
|
||||
// reference with no relocation.
|
||||
if a.Base == "GS" || a.Base == "FS" {
|
||||
seg := byte(0x64)
|
||||
if a.Base == "GS" {
|
||||
seg = 0x65
|
||||
}
|
||||
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
|
||||
}
|
||||
base, ok := ParseReg(a.Base)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown base register %q", a.Base)
|
||||
}
|
||||
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
|
||||
if a.Index != "" {
|
||||
if a.Index == "TLS" {
|
||||
// off(base)(TLS*1): the thread-local annotation. The
|
||||
// one-instruction TLS form folds it to off(TLS), the
|
||||
// segment-prefixed absolute whose disp32 carries an
|
||||
// R_TLS_LE patch site; the base register disappears
|
||||
// from the encoding, exactly as the toolchain's
|
||||
// progedit rewrites the address.
|
||||
seg := byte(0x64) // FS on linux, freebsd, plan9
|
||||
if link != nil && link.goos == "windows" {
|
||||
seg = 0x65 // GS
|
||||
}
|
||||
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
|
||||
}
|
||||
if a.Index == "GS" || a.Index == "FS" {
|
||||
// 0(CX)(GS): the segment annotation rides the base
|
||||
// access as the override prefix.
|
||||
m.Seg = 0x64
|
||||
if a.Index == "GS" {
|
||||
m.Seg = 0x65
|
||||
}
|
||||
return m, nil
|
||||
}
|
||||
idx, ok := ParseReg(a.Index)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown index register %q", a.Index)
|
||||
@@ -911,6 +1328,11 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
|
||||
}
|
||||
// A bare displacement with no base: the absolute address form,
|
||||
// MOVL $0xf1, 0xf1. No segment and no relocation.
|
||||
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
|
||||
return SegAbs{Disp: a.Offset, Size: size}, nil
|
||||
}
|
||||
// Bare register.
|
||||
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
|
||||
if r, ok := ParseReg(a.Sym.Name); ok {
|
||||
@@ -921,3 +1343,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return nil, fmt.Errorf("unsupported operand")
|
||||
}
|
||||
|
||||
// floatImmText recovers a floating-point immediate's magnitude and sign from
|
||||
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
|
||||
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
|
||||
// folder cannot read it, and only the verbatim operand text still carries
|
||||
// the value. Anything that is not a number a float parser accepts reports
|
||||
// not-ok, so every other shape keeps its existing diagnostic.
|
||||
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
|
||||
if op.Imm.Float != "" {
|
||||
return op.Imm.Float, op.Imm.Neg, true
|
||||
}
|
||||
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
|
||||
return "", false, false
|
||||
}
|
||||
// joinRaw spaced the token texts; the compact spelling is what matters.
|
||||
compact := strings.ReplaceAll(op.Raw, " ", "")
|
||||
inner, ok := strings.CutPrefix(compact, "$(")
|
||||
if !ok || !strings.HasSuffix(inner, ")") {
|
||||
return "", false, false
|
||||
}
|
||||
inner = strings.TrimSuffix(inner, ")")
|
||||
inner = strings.TrimPrefix(inner, "+")
|
||||
if s, ok := strings.CutPrefix(inner, "-"); ok {
|
||||
neg = true
|
||||
inner = s
|
||||
}
|
||||
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
|
||||
return "", false, false
|
||||
}
|
||||
if _, err := strconv.ParseFloat(inner, 64); err != nil {
|
||||
return "", false, false
|
||||
}
|
||||
return inner, neg, true
|
||||
}
|
||||
|
||||
+79
-2
@@ -10,8 +10,8 @@ import (
|
||||
|
||||
"golang.org/x/arch/x86/x86asm"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// firstText parses src and returns its first TEXT function.
|
||||
@@ -370,6 +370,58 @@ end:
|
||||
}
|
||||
}
|
||||
|
||||
// TestAssembleNumericPCJumps pins the numeric ±N(PC) branch operands: N
|
||||
// counts instruction statements, skipping labels, in both directions (the
|
||||
// runtime's exit loops write JMP -3(PC)), N = 0 parks on the jump itself.
|
||||
func TestAssembleNumericPCJumps(t *testing.T) {
|
||||
fn := firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·exit(SB), NOSPLIT, $0
|
||||
MOVB $1, AL
|
||||
lab:
|
||||
MOVB $2, AL
|
||||
MOVB $3, AL
|
||||
JMP -3(PC)
|
||||
MOVB $4, AL
|
||||
park:
|
||||
JMP 0(PC)
|
||||
MOVB $5, AL
|
||||
JMP 2(PC)
|
||||
MOVB $6, AL
|
||||
RET
|
||||
`)
|
||||
code, _, err := Assemble(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("Assemble: %v", err)
|
||||
}
|
||||
// From the Go-assembled function:
|
||||
// MOVB $1, AL b001
|
||||
// MOVB $2, AL b002
|
||||
// MOVB $3, AL b003
|
||||
// JMP -3(PC) ebf8 (three instructions back, past lab:)
|
||||
// MOVB $4, AL b004
|
||||
// JMP 0(PC) ebfe (the park loop)
|
||||
// MOVB $5, AL b005
|
||||
// JMP 2(PC) eb02 (over MOVB $6 to the RET)
|
||||
// MOVB $6, AL b006
|
||||
// RET c3
|
||||
want := []byte{
|
||||
0xb0, 0x01,
|
||||
0xb0, 0x02,
|
||||
0xb0, 0x03,
|
||||
0xeb, 0xf8,
|
||||
0xb0, 0x04,
|
||||
0xeb, 0xfe,
|
||||
0xb0, 0x05,
|
||||
0xeb, 0x02,
|
||||
0xb0, 0x06,
|
||||
0xc3,
|
||||
}
|
||||
if hexBytes(code) != hexBytes(want) {
|
||||
t.Errorf("numeric-PC mismatch:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
|
||||
}
|
||||
}
|
||||
|
||||
func TestAssemblePrefetch(t *testing.T) {
|
||||
fn := firstText(t, `
|
||||
#include "textflag.h"
|
||||
@@ -568,3 +620,28 @@ TEXT ·framed(SB), $16-8
|
||||
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
|
||||
}
|
||||
}
|
||||
|
||||
// TestAssembleRegRange pins the bracketed register range at the statement
|
||||
// level: exactly four consecutive same-width vector registers assemble, the
|
||||
// toolchain's rejected shapes all report an error.
|
||||
func TestAssembleRegRange(t *testing.T) {
|
||||
asm := func(t *testing.T, op string) ([]byte, error) {
|
||||
t.Helper()
|
||||
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse %s: %v", op, errs)
|
||||
}
|
||||
code, _, err := Assemble(f.Decls[0].(*ast.Text))
|
||||
return code, err
|
||||
}
|
||||
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
|
||||
if _, err := asm(t, op); err != nil {
|
||||
t.Errorf("%s: %v", op, err)
|
||||
}
|
||||
}
|
||||
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
|
||||
if _, err := asm(t, op); err == nil {
|
||||
t.Errorf("%s: assembled, want an error", op)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -8,7 +8,7 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// ulebIter reads ULEB128 values, the .debug_abbrev and line-header
|
||||
|
||||
+2
-2
@@ -12,8 +12,8 @@ import (
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// The object-file tests share one source: two exported functions, one
|
||||
|
||||
@@ -9,7 +9,7 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestELFAARCH64Object checks the structure of the emitted AArch64 ELF64
|
||||
|
||||
@@ -9,7 +9,7 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestELFLOONG64Object checks the structure of the emitted LoongArch ELF64
|
||||
|
||||
@@ -9,7 +9,7 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
|
||||
|
||||
+3
-3
@@ -20,9 +20,9 @@ func Encodable(mnemonic string) bool {
|
||||
switch upper {
|
||||
case "RET", "NOP", "CALL", "JMP",
|
||||
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
|
||||
// The literal-data pseudo-ops, the accepted-and-ignored END and the
|
||||
// SP adjust.
|
||||
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP":
|
||||
// The literal-data pseudo-ops, the accepted-and-ignored END and
|
||||
// bookkeeping statements, and the SP adjust.
|
||||
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
|
||||
return true
|
||||
}
|
||||
if _, ok := noOperandTable[upper]; ok {
|
||||
|
||||
+228
-2
@@ -5,6 +5,8 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
|
||||
type enc struct {
|
||||
out []byte
|
||||
patches []encPatch // disp32 fields awaiting static-symbol resolution
|
||||
|
||||
// FloatPool collects the pooled constants the floating-point
|
||||
// immediates reference, in first-use order.
|
||||
floatPool []floatPoolEntry
|
||||
floatPoolSeen map[string]bool
|
||||
}
|
||||
|
||||
// floatPoolEntry is one pooled floating-point constant: the symbol name
|
||||
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
|
||||
type floatPoolEntry struct {
|
||||
name string
|
||||
data []byte
|
||||
}
|
||||
|
||||
// addFloatPool records a pooled constant, deduplicated by symbol name.
|
||||
func (e *enc) addFloatPool(name string, bits uint64, width int) {
|
||||
if e.floatPoolSeen == nil {
|
||||
e.floatPoolSeen = map[string]bool{}
|
||||
}
|
||||
if e.floatPoolSeen[name] {
|
||||
return
|
||||
}
|
||||
e.floatPoolSeen[name] = true
|
||||
data := make([]byte, width)
|
||||
for i := range width {
|
||||
data[i] = byte(bits >> (8 * i))
|
||||
}
|
||||
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
|
||||
}
|
||||
|
||||
// floatPoolList returns the pooled constants in first-use order.
|
||||
func (e *enc) floatPoolList() []floatPoolEntry {
|
||||
return e.floatPool
|
||||
}
|
||||
|
||||
// encPatch marks a 4-byte displacement field in enc.out that must receive the
|
||||
@@ -29,6 +64,7 @@ type encPatch struct {
|
||||
off int
|
||||
name string
|
||||
addend int64
|
||||
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
|
||||
}
|
||||
|
||||
func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
@@ -102,6 +138,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
return e.encodeEnd(ops)
|
||||
case "ADJSP":
|
||||
return e.encodeAdjsp(ops)
|
||||
// The runtime's bookkeeping statements carry no text bytes: go tool asm
|
||||
// records FUNCDATA and PCDATA in the program list only, so the encoded
|
||||
// body shows nothing, on every architecture.
|
||||
case "FUNCDATA", "PCDATA":
|
||||
return e.encodeFuncdata(upper, ops)
|
||||
}
|
||||
|
||||
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
|
||||
@@ -141,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
}
|
||||
// Legacy SSE packed binaries dispatch on the full name: the packed
|
||||
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
|
||||
// which the size split must not eat.
|
||||
// which the size split must not eat. A floating-point immediate
|
||||
// rewrites into a pooled-constant read on the scalar members.
|
||||
if m, ok := sseBinTable[upper]; ok {
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatBin(upper, m, f, ops)
|
||||
}
|
||||
return e.encodeSSEBin(m, ops)
|
||||
}
|
||||
if m, ok := sseBinTable[base]; ok {
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatBin(upper, m, f, ops)
|
||||
}
|
||||
return e.encodeSSEBin(m, ops)
|
||||
}
|
||||
// The imm8-controlled legacy instructions, the lane extracts and inserts
|
||||
@@ -222,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
return e.encodeCvtInt(base, ops, size)
|
||||
case "FMOVD":
|
||||
return e.encodeFmov(ops)
|
||||
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
|
||||
case "MOVSD", "MOVSS":
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatMove(upper, f, ops)
|
||||
}
|
||||
return e.encodeSSEMove(sseMoveTable[base], ops)
|
||||
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
|
||||
return e.encodeSSEMove(sseMoveTable[base], ops)
|
||||
}
|
||||
return fmt.Errorf("unsupported instruction %q", mnem)
|
||||
@@ -285,6 +338,33 @@ func (e *enc) encodeData(mnem string, ops []Operand) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
|
||||
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
|
||||
// for either (the entries live in the object's ancillary tables, not the
|
||||
// function body), and the operand shapes it takes are exactly these: an
|
||||
// integer count first, then a symbol reference for FUNCDATA and an integer
|
||||
// value for PCDATA. The other architectures accept-and-ignore the same
|
||||
// statements; amd64 now matches.
|
||||
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
|
||||
if len(ops) != 2 {
|
||||
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
|
||||
}
|
||||
if _, ok := ops[0].(Imm); !ok {
|
||||
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
|
||||
}
|
||||
switch upper {
|
||||
case "FUNCDATA":
|
||||
if _, ok := ops[1].(sbMem); !ok {
|
||||
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
|
||||
}
|
||||
case "PCDATA":
|
||||
if _, ok := ops[1].(Imm); !ok {
|
||||
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
|
||||
// entirely: the AEND Prog is skipped when the program list is flushed, so
|
||||
// the statements after an END still belong to the same function and the
|
||||
@@ -319,6 +399,120 @@ func (e *enc) encodeAdjsp(ops []Operand) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// --- floating-point immediates ----------------------------------------------
|
||||
|
||||
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
|
||||
// immediate, the set go tool asm rewrites into a pooled-constant read: the
|
||||
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
|
||||
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
|
||||
// reject the immediate in the toolchain and are absent here on purpose.
|
||||
var sseFloatImm = map[string]bool{
|
||||
"MOVSD": true, "MOVSS": true,
|
||||
"ADDSD": true, "ADDSS": true,
|
||||
"SUBSD": true, "SUBSS": true,
|
||||
"MULSD": true, "MULSS": true,
|
||||
"DIVSD": true, "DIVSS": true,
|
||||
"COMISD": true, "COMISS": true,
|
||||
"UCOMISD": true, "UCOMISS": true,
|
||||
}
|
||||
|
||||
// floatImmOperand reports whether the operand list opens with a
|
||||
// floating-point immediate in the two-operand spelling (imm, dst).
|
||||
func floatImmOperand(ops []Operand) (FloatImm, bool) {
|
||||
if len(ops) != 2 {
|
||||
return FloatImm{}, false
|
||||
}
|
||||
f, ok := ops[0].(FloatImm)
|
||||
return f, ok
|
||||
}
|
||||
|
||||
// floatPoolValue evaluates a floating-point immediate at the width its
|
||||
// mnemonic encodes and names the pool constant the toolchain synthesises:
|
||||
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
|
||||
// rounding of the parsed value). The name carries the IEEE-754 bits; the
|
||||
// section holds them little-endian.
|
||||
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
|
||||
v, err := strconv.ParseFloat(f.Text, 64)
|
||||
if err != nil {
|
||||
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
|
||||
}
|
||||
if f.Neg {
|
||||
v = -v
|
||||
}
|
||||
if strings.HasSuffix(mnem, "D") {
|
||||
bits = math.Float64bits(v)
|
||||
return bits, fmt.Sprintf("$f64.%016x", bits), nil
|
||||
}
|
||||
bits = uint64(math.Float32bits(float32(v)))
|
||||
return bits, fmt.Sprintf("$f32.%08x", bits), nil
|
||||
}
|
||||
|
||||
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
|
||||
// source. A positive zero needs no memory read: the toolchain emits
|
||||
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
|
||||
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
|
||||
// file-level layout or the linker resolves.
|
||||
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
|
||||
if !sseFloatImm[mnem] {
|
||||
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
|
||||
}
|
||||
dst, ok := ops[1].(Reg)
|
||||
if !ok || !dst.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
bits, name, err := floatPoolValue(mnem, f)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
e.addFloatPool(name, bits, mwidth(mnem))
|
||||
if bits == 0 {
|
||||
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
|
||||
if err := setRM(i, dst, dst, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
m := sseMoveTable[mnem]
|
||||
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
|
||||
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
|
||||
// a floating-point immediate source: the constant is read from the pool into
|
||||
// the instruction's r/m side (reg = destination), the rewrite go tool asm
|
||||
// performs at the source level.
|
||||
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
|
||||
if !sseFloatImm[mnem] {
|
||||
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
|
||||
}
|
||||
dst, ok := ops[1].(Reg)
|
||||
if !ok || !dst.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
bits, name, err := floatPoolValue(mnem, f)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
e.addFloatPool(name, bits, mwidth(mnem))
|
||||
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
|
||||
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
|
||||
// spellings end in D, the single spellings in S.
|
||||
func mwidth(mnem string) int {
|
||||
if strings.HasSuffix(mnem, "D") {
|
||||
return 8
|
||||
}
|
||||
return 4
|
||||
}
|
||||
|
||||
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
|
||||
func splitSize(upper string) (base string, size int) {
|
||||
if upper == "" {
|
||||
@@ -385,6 +579,7 @@ type instr struct {
|
||||
disp []byte
|
||||
imm []byte
|
||||
sb *sbRef // static-symbol displacement in disp, awaiting resolution
|
||||
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
|
||||
}
|
||||
|
||||
// sbRef records that an instruction's displacement refers to a static symbol
|
||||
@@ -427,6 +622,9 @@ func (e *enc) emit(i *instr) error {
|
||||
if i.sb != nil {
|
||||
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
|
||||
}
|
||||
if i.tls {
|
||||
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
|
||||
}
|
||||
e.out = append(e.out, i.disp...)
|
||||
e.out = append(e.out, i.imm...)
|
||||
return nil
|
||||
@@ -481,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
|
||||
i.disp = le32(0)
|
||||
i.sb = &sbRef{name: r.name, addend: r.addend}
|
||||
return nil
|
||||
case TLSMem:
|
||||
// off(TLS): the segment-prefixed absolute access, mod=00 with the
|
||||
// SIB escape's disp32 absolute form. The displacement is the TLS
|
||||
// slot offset, patched by the linker's TLS relocation.
|
||||
i.prefix = r.Seg
|
||||
i.modrm = 0x04 | regField<<3
|
||||
i.sib = 0x25
|
||||
i.disp = le32(r.Disp)
|
||||
i.tls = true
|
||||
return nil
|
||||
case SegAbs:
|
||||
// 0x30(GS): the segment override with the SIB escape's disp32
|
||||
// absolute form, no relocation.
|
||||
setSegAbs(i, regField, r)
|
||||
return nil
|
||||
default:
|
||||
return fmt.Errorf("invalid r/m operand %T", rm)
|
||||
}
|
||||
}
|
||||
|
||||
func setMem(i *instr, regField int, m Mem) error {
|
||||
if m.Seg != 0 {
|
||||
i.prefix = m.Seg
|
||||
}
|
||||
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -499,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
|
||||
// override with the mod=00 SIB escape's disp32 absolute form and no
|
||||
// relocation.
|
||||
func setSegAbs(i *instr, regField int, m SegAbs) {
|
||||
i.prefix = m.Seg
|
||||
i.modrm = 0x04 | regField<<3
|
||||
i.sib = 0x25
|
||||
i.disp = le32(m.Disp)
|
||||
}
|
||||
|
||||
// memComponents computes the ModR/M byte (with the given reg field), the SIB
|
||||
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
|
||||
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
|
||||
|
||||
@@ -9,6 +9,9 @@ import (
|
||||
"testing"
|
||||
|
||||
"golang.org/x/arch/x86/x86asm"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// decode encodes an instruction and decodes it back, returning the decoded
|
||||
@@ -1064,3 +1067,157 @@ func TestAdjsp(t *testing.T) {
|
||||
t.Error("ADJSP AX assembled, want an error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
|
||||
// byte for byte against go tool asm: the scalar moves and the scalar
|
||||
// arithmetic read the constant from a synthesised read-only pool symbol
|
||||
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
|
||||
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
|
||||
func TestFloatImmediateGroundTruth(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
|
||||
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
|
||||
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
|
||||
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
|
||||
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
|
||||
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
|
||||
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
|
||||
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
|
||||
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
|
||||
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
|
||||
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
|
||||
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
|
||||
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
|
||||
// single spellings; negative zero keeps its sign bit and never takes the
|
||||
// XORPS shortcut.
|
||||
for _, c := range []struct {
|
||||
mnem string
|
||||
imm FloatImm
|
||||
want string
|
||||
}{
|
||||
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
|
||||
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
|
||||
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
|
||||
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
|
||||
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
|
||||
} {
|
||||
_, name, err := floatPoolValue(c.mnem, c.imm)
|
||||
if err != nil {
|
||||
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
|
||||
continue
|
||||
}
|
||||
if name != c.want {
|
||||
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
|
||||
}
|
||||
}
|
||||
|
||||
// The shapes the toolchain's parser rejects: the packed and uniform
|
||||
// forms, a non-vector destination, and the integer spellings.
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
|
||||
} {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
|
||||
// go tool asm emits no text bytes for either, on every architecture.
|
||||
func TestBookkeepingGroundTruth(t *testing.T) {
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
|
||||
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
|
||||
} {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if len(code) != 0 {
|
||||
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
|
||||
}
|
||||
}
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
|
||||
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
|
||||
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
|
||||
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
|
||||
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
|
||||
} {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
|
||||
// At the statement level the bookkeeping lines sit between real
|
||||
// instructions and contribute nothing to the body, symbol reference
|
||||
// included: the FUNCDATA operand never needs file-level resolution.
|
||||
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
want := "90c3"
|
||||
if got := hexCompact(img.Code); got != want {
|
||||
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
|
||||
}
|
||||
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
|
||||
t.Error("FUNCDATA $1, X0 assembled, want an error")
|
||||
}
|
||||
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
|
||||
t.Error("PCDATA $1, X0 assembled, want an error")
|
||||
}
|
||||
|
||||
// Encodable mirrors Encode for the names this work touched.
|
||||
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
|
||||
if !Encodable(mnem) {
|
||||
t.Errorf("Encodable(%s) = false, want true", mnem)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// mustParse parses src or fails the test.
|
||||
func mustParse(t *testing.T, src string) *ast.File {
|
||||
t.Helper()
|
||||
f, errs := parser.Parse("t_amd64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
return f
|
||||
}
|
||||
|
||||
+97
-2
@@ -737,6 +737,91 @@ var evexTable = map[string]evexSpec{
|
||||
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
}
|
||||
|
||||
// evexQuad describes one quad-register instruction: the opcode under
|
||||
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
|
||||
// registers the bracketed list and the destination take (512-bit ZMM for
|
||||
// the packed forms, 128-bit XMM for the scalar ones).
|
||||
type evexQuad struct {
|
||||
opcode byte
|
||||
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
|
||||
}
|
||||
|
||||
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
|
||||
// families) to their encoding. The operand shape is fixed: a single memory
|
||||
// source in r/m, the bracketed register list whose LOW register travels the
|
||||
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
|
||||
// destination in reg. The vector length follows the destination (512-bit
|
||||
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
|
||||
// multiplier stays 16 for every member, the toolchain's own tuple choice.
|
||||
var evexQuadTable = map[string]evexQuad{
|
||||
"V4FMADDPS": {0x9A, 64},
|
||||
"V4FMADDSS": {0x9B, 16},
|
||||
"V4FNMADDPS": {0xAA, 64},
|
||||
"V4FNMADDSS": {0xAB, 16},
|
||||
"VP4DPWSSD": {0x52, 64},
|
||||
"VP4DPWSSDS": {0x53, 64},
|
||||
}
|
||||
|
||||
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
|
||||
func isEvexQuad(upper string) bool {
|
||||
_, ok := evexQuadTable[upper]
|
||||
return ok
|
||||
}
|
||||
|
||||
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
|
||||
// The register list is the VVVV-side source: its low register fills the
|
||||
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
|
||||
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
|
||||
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
|
||||
// suffix applies.
|
||||
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
|
||||
if len(ops) != 3 && len(ops) != 4 {
|
||||
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
|
||||
}
|
||||
mem, lst := ops[0], ops[1]
|
||||
dst := ops[len(ops)-1]
|
||||
mask := 0
|
||||
if len(ops) == 4 {
|
||||
k, ok := ops[2].(Reg)
|
||||
if !ok || !k.mask {
|
||||
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
|
||||
}
|
||||
if k.idx == 0 {
|
||||
return fmt.Errorf("k0 is not a usable mask register")
|
||||
}
|
||||
mask = k.idx
|
||||
}
|
||||
list, ok := lst.(RegList)
|
||||
if !ok {
|
||||
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
|
||||
}
|
||||
if list.Lo.size != q.width {
|
||||
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
|
||||
}
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || !dstReg.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
if dstReg.size != q.width {
|
||||
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
|
||||
}
|
||||
if !memOperand(mem) {
|
||||
return fmt.Errorf("%s: the source must be a memory operand", mnem)
|
||||
}
|
||||
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
|
||||
// its fifth bit into the same field the list's low register occupies.
|
||||
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
|
||||
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
|
||||
}
|
||||
if sfx.zeroing && mask == 0 {
|
||||
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
|
||||
}
|
||||
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
|
||||
// The vector length follows the destination (512-bit for the ZMM forms,
|
||||
// 128-bit for the scalar ones), exactly as the oracle encodes it.
|
||||
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
|
||||
}
|
||||
|
||||
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
|
||||
// depends on the source kind, a GPR source uses opReg, a memory source uses
|
||||
// opMem with a disp8×N of n.
|
||||
@@ -812,8 +897,10 @@ func isEvex(mnemUpper string) bool {
|
||||
if _, ok := evexBcastTable[mnemUpper]; ok {
|
||||
return true
|
||||
}
|
||||
_, ok := evexMoveTable[mnemUpper]
|
||||
return ok
|
||||
if _, ok := evexMoveTable[mnemUpper]; ok {
|
||||
return true
|
||||
}
|
||||
return isEvexQuad(mnemUpper)
|
||||
}
|
||||
|
||||
// evexRequired reports whether the operands force the EVEX encoding of a
|
||||
@@ -995,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
|
||||
return e.encodeEvexRM(spec, ops, 0, sfx)
|
||||
}
|
||||
spec, inTable := evexTable[mnemUpper]
|
||||
if q, ok := evexQuadTable[mnemUpper]; ok {
|
||||
// The quad-register family carries no rounding, SAE or broadcast;
|
||||
// only masking and zeroing apply.
|
||||
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
|
||||
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
|
||||
}
|
||||
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
|
||||
}
|
||||
if inTable {
|
||||
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
|
||||
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
|
||||
|
||||
@@ -810,3 +810,127 @@ func TestAvx512CorpusFamilies(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
|
||||
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
|
||||
// source keeps r/m, the bracketed list's LOW register travels the inverted
|
||||
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
|
||||
// the vector length follows the destination (L'L=512 for the ZMM forms,
|
||||
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
|
||||
// member. The x86 decoder has no view of these forms, so no decode check
|
||||
// runs.
|
||||
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
|
||||
sp := vreg(t, "RSP")
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a9a842411000000"},
|
||||
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f22f4a9a842411000000"},
|
||||
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f25f429a842411000000"},
|
||||
{"V4FMADDPS Z8 dst", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
|
||||
"62727f4a9a842411000000"},
|
||||
{"V4FMADDPS disp8x16", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a9a442404"},
|
||||
{"V4FMADDPS unmasked", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
|
||||
"62f27f489a842411000000"},
|
||||
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0d9bb007000000"},
|
||||
{"V4FMADDSS (DI)", "V4FMADDSS",
|
||||
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0d9b37"},
|
||||
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e22f0d9bb007000000"},
|
||||
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e25f059bb007000000"},
|
||||
{"V4FMADDSS X30 dst", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
|
||||
"62627f0d9bb007000000"},
|
||||
{"V4FMADDSS X3 dst", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
|
||||
"62f27f0d9b9807000000"},
|
||||
{"V4FMADDSS disp8x16", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
|
||||
"62625f059b7001"},
|
||||
{"V4FNMADDPS", "V4FNMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4aaa842411000000"},
|
||||
{"V4FNMADDSS", "V4FNMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0dabb007000000"},
|
||||
{"VP4DPWSSD", "VP4DPWSSD",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a52842411000000"},
|
||||
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
|
||||
"62f27f4853842411000000"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
|
||||
// the register class the list and the destination take is fixed per
|
||||
// instruction, the source is memory only, the opmask slot is positional and
|
||||
// the list's low register owns V'VVVV.
|
||||
func TestEvexQuadRegisterErrors(t *testing.T) {
|
||||
sp := vreg(t, "RSP")
|
||||
list := func(lo, hi string) RegList {
|
||||
return RegList{vreg(t, lo), vreg(t, hi)}
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"X list on the PS form", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"Z list on the SS form", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
|
||||
{"Y destination", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
|
||||
{"register source", "V4FMADDPS",
|
||||
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"non-mask third operand", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
|
||||
{"k0 mask", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
|
||||
{"K after the destination", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
|
||||
{"zeroing without a mask", "V4FMADDPS.Z",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
|
||||
{"SAE suffix", "V4FMADDPS.SAE",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"high index source", "VP4DPWSSD",
|
||||
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"short operand list", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+1
-1
@@ -12,7 +12,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// goobjView is a minimal parsed view of a GOOBJ payload, enough to check
|
||||
|
||||
+6
-6
@@ -9,8 +9,8 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// The expected bytes are pinned from `go tool asm` output (Go 1.27, amd64,
|
||||
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
|
||||
func TestStackGuardGOObjInternalCall(t *testing.T) {
|
||||
for _, tt := range []struct {
|
||||
src string
|
||||
assemble func(*ast.File) (*Image, error)
|
||||
assemble func(*ast.File, ...AssembleOption) (*Image, error)
|
||||
}{
|
||||
{"g_amd64.s", AssembleFile},
|
||||
{"g_arm64.s", AssembleFileARM64},
|
||||
{"g_riscv64.s", AssembleFileRISCV},
|
||||
{"g_loong64.s", AssembleFileLOONG64},
|
||||
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
|
||||
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
|
||||
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
|
||||
} {
|
||||
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
|
||||
+74
-21
@@ -192,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case TLSMem:
|
||||
if !dstIsReg {
|
||||
return fmt.Errorf("MOV: two memory operands")
|
||||
}
|
||||
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
|
||||
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
|
||||
// offset with its R_TLSLE patch site.
|
||||
i := newInstr(size, []byte{movRR(size)})
|
||||
if err := setRM(i, dstReg, src, size); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case SegAbs:
|
||||
if !dstIsReg {
|
||||
return fmt.Errorf("MOV: two memory operands")
|
||||
}
|
||||
// MOV r, 0x30(GS): the segment-absolute load.
|
||||
i := newInstr(size, []byte{movRR(size)})
|
||||
if err := setRM(i, dstReg, src, size); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case Imm:
|
||||
if dstIsReg {
|
||||
v := int64(src)
|
||||
@@ -232,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
|
||||
i.imm = imm
|
||||
return e.emit(i)
|
||||
}
|
||||
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
|
||||
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
|
||||
// destination slot is the absolute-address crash-store spelling,
|
||||
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
|
||||
// as an immediate, and the store's disp32 carries the address.
|
||||
op := byte(0xC7)
|
||||
if size == 1 {
|
||||
op = 0xC6
|
||||
}
|
||||
if d, ok := dst.(Imm); ok {
|
||||
i := newInstr(size, []byte{op})
|
||||
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
|
||||
immBytes, err := immediate(int64(src), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{op})
|
||||
if err := setRMDigit(i, 0, dst, size); err != nil {
|
||||
return err
|
||||
@@ -627,7 +664,16 @@ func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
|
||||
func (e *enc) encodeImul(ops []Operand, size int) error {
|
||||
switch len(ops) {
|
||||
case 2:
|
||||
// IMUL r, r/m: 0x0F 0xAF.
|
||||
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
|
||||
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
|
||||
// Otherwise IMUL r, r/m: 0x0F 0xAF.
|
||||
if imm, ok := ops[0].(Imm); ok {
|
||||
dstReg, isReg := ops[1].(Reg)
|
||||
if !isReg {
|
||||
return fmt.Errorf("IMUL: destination must be a register")
|
||||
}
|
||||
return e.encodeImulImm(imm, dstReg, dstReg, size)
|
||||
}
|
||||
dstReg, ok := ops[1].(Reg)
|
||||
if !ok {
|
||||
return fmt.Errorf("IMUL: destination must be a register")
|
||||
@@ -647,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
|
||||
if !ok {
|
||||
return fmt.Errorf("IMUL: immediate operand expected first")
|
||||
}
|
||||
// Plan 9 order: IMUL $imm, src, dst.
|
||||
if fits8(int64(imm)) {
|
||||
i := newInstr(size, []byte{0x6B})
|
||||
if err := setRM(i, dstReg, ops[1], size); err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = []byte{byte(int8(imm))}
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{0x69})
|
||||
if err := setRM(i, dstReg, ops[1], size); err != nil {
|
||||
return err
|
||||
}
|
||||
immBytes, err := immediate(int64(imm), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
|
||||
// r/m operand (setRM takes registers and memory alike).
|
||||
return e.encodeImulImm(imm, ops[1], dstReg, size)
|
||||
}
|
||||
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
|
||||
}
|
||||
|
||||
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
|
||||
// when the value fits, 0x69 with a 32-bit immediate otherwise.
|
||||
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
|
||||
if fits8(int64(imm)) {
|
||||
i := newInstr(size, []byte{0x6B})
|
||||
if err := setRM(i, dst, rm, size); err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = []byte{byte(int8(imm))}
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{0x69})
|
||||
if err := setRM(i, dst, rm, size); err != nil {
|
||||
return err
|
||||
}
|
||||
immBytes, err := immediate(int64(imm), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// --- PUSH / POP -------------------------------------------------------------
|
||||
|
||||
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
|
||||
|
||||
@@ -16,7 +16,7 @@ import (
|
||||
|
||||
"golang.org/x/arch/x86/x86asm"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestAssembleGoFlacAVX2Kernel assembles the whole production AVX2 kernel;
|
||||
@@ -25,7 +25,7 @@ import (
|
||||
func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
|
||||
path := "../../go-libraries/go-flac/avx2_amd64.s"
|
||||
if _, err := os.Stat(path); err != nil {
|
||||
t.Skip("go-libraries repository not present next to gasm-devkit")
|
||||
t.Skip("go-libraries repository not present next to gasm-sdk")
|
||||
}
|
||||
src, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
@@ -86,7 +86,7 @@ func TestAssembleGoFlacAVX2Kernel(t *testing.T) {
|
||||
func TestAssembleGoFlacAVX512Kernel(t *testing.T) {
|
||||
path := "../../go-libraries/go-flac/avx512_amd64.s"
|
||||
if _, err := os.Stat(path); err != nil {
|
||||
t.Skip("go-libraries repository not present next to gasm-devkit")
|
||||
t.Skip("go-libraries repository not present next to gasm-sdk")
|
||||
}
|
||||
src, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
|
||||
@@ -13,7 +13,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// The differential kernels for the DATA-path and front-end gaps are kept in
|
||||
@@ -63,7 +63,10 @@ func toolAsmObject(t *testing.T, path, goarch string) []byte {
|
||||
}
|
||||
|
||||
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
|
||||
// toolchain object, keyed by the name the object records (pkg.name).
|
||||
// toolchain object, keyed by the name the object records (pkg.name). Each
|
||||
// function's span is its own symbol size: a toolchain object that follows
|
||||
// the text with data symbols (the synthesised float-constant pool) would
|
||||
// otherwise fold them into the last function's bytes.
|
||||
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
|
||||
t.Helper()
|
||||
v := openGoobj(t, obj)
|
||||
@@ -76,18 +79,13 @@ func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
|
||||
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
|
||||
preceding += len(v.blk(bi)) / symSize
|
||||
}
|
||||
total := preceding + len(nps)
|
||||
out := make(map[string][]byte, len(nps))
|
||||
for i, s := range nps {
|
||||
if s.typ != kindSTEXT {
|
||||
continue
|
||||
}
|
||||
start := le.Uint32(didx[4*(preceding+i):])
|
||||
end := uint32(len(data))
|
||||
if preceding+i+1 < total {
|
||||
end = le.Uint32(didx[4*(preceding+i+1):])
|
||||
}
|
||||
out[s.name] = data[start:end]
|
||||
out[s.name] = data[start : start+s.size]
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -129,6 +127,10 @@ func TestDifferentialKernels(t *testing.T) {
|
||||
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
|
||||
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
|
||||
} {
|
||||
|
||||
@@ -12,7 +12,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestGOObjectLOONG64Structure checks the emitted loong64 object's blocks:
|
||||
|
||||
+109
-9
@@ -5,10 +5,11 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"sort"
|
||||
"strconv"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// Image is an assembled file: the function bodies laid out in source order,
|
||||
@@ -149,6 +150,18 @@ func (img *Image) Bytes() []byte {
|
||||
return append(out, img.Data...)
|
||||
}
|
||||
|
||||
// AssembleOption adjusts the file-level assembly context.
|
||||
type AssembleOption func(*linkInfo)
|
||||
|
||||
// WithGOOS selects the target operating system for the forms that depend on
|
||||
// it, the TLS access shape above all: linux and freebsd take the
|
||||
// one-instruction form, windows and plan9 keep the two-instruction load.
|
||||
func WithGOOS(goos string) AssembleOption {
|
||||
return func(l *linkInfo) {
|
||||
l.goos = goos
|
||||
}
|
||||
}
|
||||
|
||||
// AssembleFile assembles every TEXT function of a parsed file and lays out
|
||||
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
|
||||
// reference to a file-local static symbol becomes a RIP-relative load whose
|
||||
@@ -156,7 +169,7 @@ func (img *Image) Bytes() []byte {
|
||||
// GLOBL defines is recorded as an external relocation (Externals) with its
|
||||
// displacement left zero, the object-file emitters resolve it at link
|
||||
// time, while the raw image (Bytes) cannot represent it.
|
||||
func AssembleFile(f *ast.File) (*Image, error) {
|
||||
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
|
||||
dataSyms, err := collectData(f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -165,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
|
||||
for _, d := range dataSyms {
|
||||
known[d.name] = true
|
||||
}
|
||||
// TEXT symbols are file-level definitions too: a symbol immediate
|
||||
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
|
||||
for _, d := range f.Decls {
|
||||
if t, ok := d.(*ast.Text); ok {
|
||||
known[t.Name.Name] = true
|
||||
}
|
||||
}
|
||||
link := &linkInfo{symbols: known, allowExternal: true}
|
||||
for _, o := range opts {
|
||||
o(link)
|
||||
}
|
||||
poolSeen := map[string]bool{}
|
||||
|
||||
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
|
||||
textOff := map[string]int{}
|
||||
@@ -179,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
code, patches, labels, steps, lines, err := assemble(t, link)
|
||||
code, patches, labels, steps, lines, pool, err := assemble(t, link)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
// The pooled floating-point constants join the declared data as
|
||||
// read-only symbols, deduplicated across the file (the toolchain
|
||||
// synthesises the same symbols into its rodata).
|
||||
for _, entry := range pool {
|
||||
if poolSeen[entry.name] {
|
||||
continue
|
||||
}
|
||||
poolSeen[entry.name] = true
|
||||
dataSyms = append(dataSyms, dataSym{
|
||||
name: entry.name,
|
||||
buf: entry.data,
|
||||
size: len(entry.data),
|
||||
rodata: true,
|
||||
dupok: true,
|
||||
})
|
||||
}
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
@@ -299,6 +342,10 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// The pooled $i64 constants the wide MOV immediate loads refer to join
|
||||
// the declared data as read-only symbols, deduplicated across the file
|
||||
// (the toolchain synthesises the same symbols into its rodata).
|
||||
litSeen := map[string]bool{}
|
||||
|
||||
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
|
||||
for _, d := range f.Decls {
|
||||
@@ -306,10 +353,23 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
code, labels, relocs, lines, spadj, err := assembleRISCV(t)
|
||||
code, labels, relocs, lines, spadj, lits, err := assembleRISCV(t)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
for _, lit := range lits {
|
||||
if litSeen[lit.Name] {
|
||||
continue
|
||||
}
|
||||
litSeen[lit.Name] = true
|
||||
dataSyms = append(dataSyms, dataSym{
|
||||
name: lit.Name,
|
||||
buf: lit.Data,
|
||||
size: len(lit.Data),
|
||||
rodata: true,
|
||||
dupok: true,
|
||||
})
|
||||
}
|
||||
fl := FuncLayout{
|
||||
Name: t.Name.Name,
|
||||
Pkg: t.Name.Pkg,
|
||||
@@ -561,11 +621,6 @@ func collectData(f *ast.File) ([]dataSym, error) {
|
||||
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
|
||||
}
|
||||
w := dd.Width
|
||||
switch w {
|
||||
case 1, 2, 4, 8:
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
off := dd.Name.Offset
|
||||
buf := syms[i].buf
|
||||
if off < 0 || off+int64(w) > int64(len(buf)) {
|
||||
@@ -587,9 +642,54 @@ func collectData(f *ast.File) ([]dataSym, error) {
|
||||
})
|
||||
continue
|
||||
}
|
||||
// A string or rune value ("DATA s+0(SB)/20, $"text"") writes its
|
||||
// bytes into the field and leaves the rest zero, the toolchain's
|
||||
// WriteString: the declared width must hold every byte, and any
|
||||
// width is legal.
|
||||
if s := dd.Value.Imm.Str; s != "" && !dd.Value.Imm.HasVal {
|
||||
text, err := strconv.Unquote(s)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("DATA %q: invalid string value %s", dd.Name.Name, s)
|
||||
}
|
||||
if len(text) > w {
|
||||
return nil, fmt.Errorf("DATA %q: string of %d bytes does not fit width %d", dd.Name.Name, len(text), w)
|
||||
}
|
||||
copy(buf[off:], text)
|
||||
continue
|
||||
}
|
||||
// A floating-point value stores its IEEE-754 bits: /4 the float32
|
||||
// rounding of the parsed double, /8 the full 64 bits, the
|
||||
// toolchain's WriteFloat32 and WriteFloat64.
|
||||
if f := dd.Value.Imm.Float; f != "" && !dd.Value.Imm.HasVal {
|
||||
num, err := strconv.ParseFloat(f, 64)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("DATA %q: invalid floating-point value %q", dd.Name.Name, f)
|
||||
}
|
||||
if dd.Value.Imm.Neg {
|
||||
num = -num
|
||||
}
|
||||
var v uint64
|
||||
switch w {
|
||||
case 4:
|
||||
v = uint64(math.Float32bits(float32(num)))
|
||||
case 8:
|
||||
v = math.Float64bits(num)
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d for a float (want 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
for j := range w {
|
||||
buf[off+int64(j)] = byte(v >> (8 * j))
|
||||
}
|
||||
continue
|
||||
}
|
||||
if !dd.Value.Imm.HasVal {
|
||||
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
|
||||
}
|
||||
switch w {
|
||||
case 1, 2, 4, 8:
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
v := dd.Value.Imm.Val
|
||||
if dd.Value.Imm.Neg {
|
||||
v = -v
|
||||
|
||||
+74
-1
@@ -5,13 +5,14 @@ package asm
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestAssembleFileStaticData checks the whole-image layout; code, padding
|
||||
@@ -418,3 +419,75 @@ func main() {
|
||||
t.Fatalf("linked program failed: %v\n%s", err, out)
|
||||
}
|
||||
}
|
||||
|
||||
// TestCollectDataFloatAndStringValues covers the non-integer DATA values the
|
||||
// runtime's math and asm files use: floating-point initialisers store their
|
||||
// IEEE-754 bits (/4 the float32 rounding, /8 the full double) and string
|
||||
// initialisers write their bytes zero-padded within the declared width.
|
||||
func TestCollectDataFloatAndStringValues(t *testing.T) {
|
||||
src := `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-8
|
||||
RET
|
||||
GLOBL vals<>(SB), RODATA, $44
|
||||
DATA vals<>+0(SB)/8, $0.5
|
||||
DATA vals<>+8(SB)/8, $-1.0
|
||||
DATA vals<>+16(SB)/4, $1.5
|
||||
DATA vals<>+20(SB)/16, $"call frame too "
|
||||
DATA vals<>+36(SB)/4, $"hi"
|
||||
`
|
||||
f, errs := parser.Parse("fvals_amd64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFile: %v", err)
|
||||
}
|
||||
byName := map[string]DataSymbol{}
|
||||
for _, d := range img.DataSyms {
|
||||
byName[d.Name] = d
|
||||
}
|
||||
d := byName["vals"]
|
||||
if d.Size != 44 {
|
||||
t.Fatalf("vals size = %d, want 44", d.Size)
|
||||
}
|
||||
buf := img.Data[d.Offset : d.Offset+44]
|
||||
// 0.5 = 0x3FE0000000000000, -1.0 = 0xBFF0000000000000 (float64);
|
||||
// 1.5 = 0x3FC00000 (float32).
|
||||
for _, c := range []struct {
|
||||
off int
|
||||
want []byte
|
||||
}{
|
||||
{0, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xE0, 0x3F}},
|
||||
{8, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xF0, 0xBF}},
|
||||
{16, []byte{0x00, 0x00, 0xC0, 0x3F}},
|
||||
{20, []byte("call frame too ")},
|
||||
{36, []byte{'h', 'i', 0x00, 0x00}},
|
||||
} {
|
||||
if string(buf[c.off:c.off+len(c.want)]) != string(c.want) {
|
||||
t.Errorf("vals+%d: got % x, want % x", c.off, buf[c.off:c.off+len(c.want)], c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestCollectDataValueErrors pins the value-kind width rules: a float needs
|
||||
// width 4 or 8, a string must fit its declared width, and a bad float
|
||||
// literal is diagnosed rather than stored.
|
||||
func TestCollectDataValueErrors(t *testing.T) {
|
||||
cases := []string{
|
||||
`GLOBL v<>(SB), RODATA, $4
|
||||
DATA v<>+0(SB)/1, $0.5`,
|
||||
`GLOBL v<>(SB), RODATA, $2
|
||||
DATA v<>+0(SB)/2, $"toolarge"`,
|
||||
}
|
||||
for i, src := range cases {
|
||||
full := "#include \"textflag.h\"\nTEXT ·Keep(SB), NOSPLIT, $0-8\n\tRET\n" + src
|
||||
f, errs := parser.Parse(fmt.Sprintf("verr%d_amd64.s", i), full)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("case %d parse: %v", i, errs)
|
||||
}
|
||||
if _, err := AssembleFile(f); err == nil {
|
||||
t.Errorf("case %d: expected an error, got none", i)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -9,7 +9,7 @@ import (
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// assembleLOONG64 assembles a LoongArch (loong64) TEXT function body into
|
||||
|
||||
@@ -8,8 +8,8 @@ import (
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// firstTextLOONG64 parses assembly source and returns the first TEXT body.
|
||||
|
||||
@@ -6,7 +6,7 @@ package asm
|
||||
import (
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// Loong64 frame mapping, matching the Go toolchain's loong64 backend.
|
||||
|
||||
@@ -7,7 +7,7 @@ import (
|
||||
"bytes"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestLOONG64_sys exercises the no-operand system instructions and the
|
||||
|
||||
@@ -6,7 +6,7 @@ package asm
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestLOONG64RelocOffsetsIncludePrologue pins the function-relative
|
||||
|
||||
@@ -14,6 +14,51 @@ type Imm int64
|
||||
|
||||
func (Imm) isOperand() {}
|
||||
|
||||
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
|
||||
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
|
||||
// low register through the inverted 5-bit V'VVVV field; the three higher
|
||||
// registers are implied by the instruction, so only the pair travels here.
|
||||
type RegList struct {
|
||||
Lo Reg
|
||||
Hi Reg // implied by the encoding; Lo.idx+3 by construction
|
||||
}
|
||||
|
||||
func (RegList) isOperand() {}
|
||||
|
||||
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
|
||||
// encoding takes an XMM/memory source at that position rewrite it as a read
|
||||
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
|
||||
// own behaviour; every other instruction rejects it.
|
||||
type FloatImm struct {
|
||||
Text string // the numeric text as written, sign excluded
|
||||
Neg bool // a leading minus
|
||||
}
|
||||
|
||||
func (FloatImm) isOperand() {}
|
||||
|
||||
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
|
||||
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
|
||||
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
|
||||
// (the linker fills the TLS slot offset).
|
||||
type TLSMem struct {
|
||||
Disp int64
|
||||
Size int
|
||||
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
|
||||
}
|
||||
|
||||
func (TLSMem) isOperand() {}
|
||||
|
||||
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
|
||||
// prefixes a disp32 absolute reference with no relocation. The base
|
||||
// register spellings GS and FS produce it.
|
||||
type SegAbs struct {
|
||||
Disp int64
|
||||
Size int
|
||||
Seg byte // 0x64 FS, 0x65 GS
|
||||
}
|
||||
|
||||
func (SegAbs) isOperand() {}
|
||||
|
||||
// Mem is a memory operand of the form disp(base)(index*scale).
|
||||
type Mem struct {
|
||||
Base Reg
|
||||
@@ -23,6 +68,7 @@ type Mem struct {
|
||||
Size int // operand width in bytes
|
||||
HasBase bool
|
||||
HasIndex bool
|
||||
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
|
||||
}
|
||||
|
||||
func (Mem) isOperand() {}
|
||||
|
||||
+287
-36
@@ -6,21 +6,23 @@ package asm
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"math/bits"
|
||||
"slices"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// assembleRISCV assembles a RISC-V TEXT function body into machine code.
|
||||
// It handles the full RV64IMAFDC instruction set including RVC compression.
|
||||
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
|
||||
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, []RiscvLiteral, error) {
|
||||
fi := riscvComputeFrame(t)
|
||||
prologue := riscvPrologue(fi)
|
||||
guardLen, err := riscvGuardLen(fi)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, err
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
lits := &riscvLiterals{}
|
||||
|
||||
var relocs []Reloc
|
||||
var spadj []SpadjStep
|
||||
@@ -74,9 +76,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
pc := len(prologue)
|
||||
for i := range recs {
|
||||
branchLike := isBranchLike(recs[i].instr.Mnemonic.Text) || riscvIsCondBranch(recs[i].instr.Mnemonic.Text)
|
||||
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil) // no relocs in Pass 2
|
||||
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil, lits) // no relocs in Pass 2
|
||||
if err != nil && !(branchLike && riscvIsRangeError(err)) {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
|
||||
}
|
||||
if err != nil {
|
||||
code = make([]byte, 4)
|
||||
@@ -178,13 +180,20 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
}
|
||||
}
|
||||
if !changed {
|
||||
// Capture the final pcs for the N(PC) branch forms: their target
|
||||
// is the instruction N source slots away, resolved by index.
|
||||
// Capture the final pcs for the N(PC) branch and jump forms: the
|
||||
// target is the instruction N source slots away (N=0 the branch
|
||||
// itself, N negative backwards), resolved by index against the
|
||||
// final layout.
|
||||
pcRelPcs = map[*ast.Instr]int{}
|
||||
for i := range recs {
|
||||
if _, ok := riscvPCRelOffset(recs[i].instr); ok {
|
||||
pcRelPcs[recs[i].instr] = pcs[i]
|
||||
n, ok := riscvPCRelOffset(recs[i].instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if i+n < 0 || i+n >= len(recs) {
|
||||
continue
|
||||
}
|
||||
pcRelPcs[recs[i].instr] = pcs[i+n]
|
||||
}
|
||||
break
|
||||
}
|
||||
@@ -198,7 +207,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
var out []byte
|
||||
guardBytes, guardReloc, err := riscvGuard(fi)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, err
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
if fi.needSplit {
|
||||
out = append(out, guardBytes...)
|
||||
@@ -220,11 +229,11 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
// The JMP a relaxation inserted: JAL X0 to the original target.
|
||||
targetOff, ok := offsets[r.jmpTo]
|
||||
if !ok {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
|
||||
}
|
||||
offset := int32(targetOff - pc)
|
||||
if err := riscvCheckJumpOffset(r.jmpTo, offset); err != nil {
|
||||
return nil, nil, nil, nil, nil, err
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
word := riscvJType(0, offset)
|
||||
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
|
||||
@@ -233,7 +242,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
// JMP, always the very next instruction (offset 4).
|
||||
enc, rs1, rs2, ok := riscvInvertedBranchEnc(strings.ToUpper(r.instr.Mnemonic.Text), r.instr.Operands)
|
||||
if !ok {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
|
||||
}
|
||||
word := riscvBType(enc, rs1, rs2, 4)
|
||||
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
|
||||
@@ -241,9 +250,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
code = r.code
|
||||
default:
|
||||
var err error
|
||||
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs)
|
||||
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs, lits)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, err
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
if c16, ok := tryCompressRVC(r.instr, fi); ok {
|
||||
code = []byte{byte(c16), byte(c16 >> 8)}
|
||||
@@ -270,7 +279,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
|
||||
if fi.needSplit {
|
||||
relocs = append(relocs, guardReloc)
|
||||
}
|
||||
return out, offsets, relocs, lines, spadj, nil
|
||||
return out, offsets, relocs, lines, spadj, lits.list(), nil
|
||||
}
|
||||
|
||||
// riscvImmAlias maps the R-type ALU mnemonics onto their I-type immediate
|
||||
@@ -348,6 +357,11 @@ func riscvPadBytes(pad int) []byte {
|
||||
func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
|
||||
mnem := instr.Mnemonic.Text
|
||||
ops := instr.Operands
|
||||
mnem = riscvNormalisePseudo(mnem)
|
||||
if mnem == "FUNCDATA" || mnem == "PCDATA" {
|
||||
// The bookkeeping statements contribute no bytes.
|
||||
return 0
|
||||
}
|
||||
var immNeg bool
|
||||
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
|
||||
if mnem == "RET" {
|
||||
@@ -368,7 +382,25 @@ func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
|
||||
}
|
||||
// MOV $imm, rd → size depends on the immediate and RVC compression.
|
||||
if isImmOperand(ops[0]) && ops[0].Imm.Sym == nil {
|
||||
return riscvMovImmSize(regFromOperand(ops[1]), immFromOperand(ops[0]))
|
||||
imm := riscvOperandImm64(ops[0])
|
||||
if int64(int32(imm)) != imm {
|
||||
return riscvMovImm64Size(regFromOperand(ops[1]), imm)
|
||||
}
|
||||
return riscvMovImmSize(regFromOperand(ops[1]), int32(imm))
|
||||
}
|
||||
// MOV $sym+off(FP|SP), rd → the frame-adjusted offset as an ADDI,
|
||||
// compressed like riscvSPAddiBytes encodes it.
|
||||
if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil &&
|
||||
(ops[0].Imm.Sym.Pseudo == "FP" || ops[0].Imm.Sym.Pseudo == "SP") {
|
||||
rd := regFromOperand(ops[1])
|
||||
_, off := riscvResolvePseudo(ops[0].Imm.Sym, fi)
|
||||
if rd > 0 && off == 0 {
|
||||
return 2 // C.MV rd, SP
|
||||
}
|
||||
if isRVCIntReg(rd) && off > 0 && off < 1024 && off%4 == 0 {
|
||||
return 2 // C.ADDI4SPN
|
||||
}
|
||||
return riscvItypeImmediateSize("ADDI", off)
|
||||
}
|
||||
// Frame-relative loads and stores: a frame offset beyond the signed
|
||||
// 12-bit range materialises the address in X31 first.
|
||||
@@ -664,9 +696,10 @@ func riscvCheckJumpOffset(target string, off int32) error {
|
||||
}
|
||||
|
||||
// encodeRISCVInstr encodes a single RISC-V instruction.
|
||||
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
|
||||
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int, lits *riscvLiterals) ([]byte, error) {
|
||||
mnem := instr.Mnemonic.Text
|
||||
ops := instr.Operands
|
||||
mnem = riscvNormalisePseudo(mnem)
|
||||
var immNeg bool
|
||||
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
|
||||
var word uint32
|
||||
@@ -677,6 +710,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
|
||||
// RET = epilogue (restore LR and close the frame when present) +
|
||||
// uncompressed JALR X0, 0(X1) (the toolchain never compresses RET).
|
||||
return riscvReturn(fi), nil
|
||||
case "FUNCDATA":
|
||||
// The assembler's bookkeeping statement, the expanded form of the
|
||||
// GO_ARGS and NO_LOCAL_POINTERS macros: FUNCDATA $n, sym(SB)
|
||||
// contributes no bytes, exactly as the toolchain's listing shows
|
||||
// (the FUNCDATA entries and the instruction after them share a PC).
|
||||
if len(ops) != 2 || !isImmOperand(ops[0]) {
|
||||
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
|
||||
}
|
||||
return nil, nil
|
||||
case "PCDATA":
|
||||
// The other bookkeeping statement, the expanded form of
|
||||
// GO_RESULTS_INITIALIZED: PCDATA $n, $m contributes no bytes too.
|
||||
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
|
||||
return nil, fmt.Errorf("PCDATA expects $n, $m")
|
||||
}
|
||||
return nil, nil
|
||||
case "WORD":
|
||||
// WORD $w lays down a raw 32-bit little-endian word.
|
||||
if len(ops) != 1 {
|
||||
@@ -739,6 +788,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
|
||||
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
|
||||
}
|
||||
target = labelFromOperand(ops[0])
|
||||
// JMP N(PC): the PC-relative slot form, resolved like the
|
||||
// branches (the toolchain counts source instructions at a
|
||||
// uniform 4 bytes, so JMP 0(PC) is a self-loop and JMP -3(PC)
|
||||
// reaches twelve bytes back). It must be recognised before the
|
||||
// indirect-register form, whose operand it resembles.
|
||||
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
offset := int32(off - pc)
|
||||
if err := riscvCheckJumpOffset("", offset); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
word = riscvJType(0, offset)
|
||||
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
|
||||
}
|
||||
// JMP (X5): an indirect branch, the toolchain's JALR X0, 0(X5).
|
||||
if ops[0].Addr.Sym == nil && ops[0].Addr.Base != "" {
|
||||
if ops[0].Addr.Offset != 0 || ops[0].Addr.Index != "" {
|
||||
@@ -751,17 +816,6 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
|
||||
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, rs1, 0)
|
||||
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
|
||||
}
|
||||
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
offset := int32(off - pc)
|
||||
if err := riscvCheckJumpOffset("", offset); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
word = riscvJType(0, offset)
|
||||
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
|
||||
}
|
||||
}
|
||||
targetOff, ok := offsets[target]
|
||||
if !ok {
|
||||
@@ -810,7 +864,7 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
|
||||
// (MOVB/MOVH/MOVW and unsigned forms) select the access width, and
|
||||
// MOVD/MOVF address the FP registers.
|
||||
case "MOV", "MOVB", "MOVBU", "MOVH", "MOVHU", "MOVW", "MOVWU", "MOVF", "MOVD":
|
||||
return encodeRISCVMov(instr, fi, relocs)
|
||||
return encodeRISCVMov(instr, fi, relocs, lits)
|
||||
|
||||
// JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1).
|
||||
case "JALR":
|
||||
@@ -1316,7 +1370,7 @@ func isImmOperand(op *ast.Operand) bool {
|
||||
// - MOV Rs, (Rd) register-relative store
|
||||
// - MOV Rs, Rd register-to-register move (ADDI $0)
|
||||
// - MOV $imm, Rd load immediate (ADDI or LUI+ADDIW)
|
||||
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) {
|
||||
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc, lits *riscvLiterals) ([]byte, error) {
|
||||
ops := instr.Operands
|
||||
if len(ops) != 2 {
|
||||
return nil, fmt.Errorf("MOV expects 2 operands, got %d", len(ops))
|
||||
@@ -1335,8 +1389,21 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
|
||||
}
|
||||
return encodeRISCVSBAddr(src.Imm.Sym, rd, relocs), nil
|
||||
}
|
||||
// MOV $sym+off(FP|SP), rd: the address of a frame slot as an
|
||||
// immediate is the frame-adjusted offset against the hardware SP,
|
||||
// the toolchain's ADDI $adj, SP, rd (argframe+0(FP) in the runtime's
|
||||
// reflect trampolines is the spelling).
|
||||
if src.Imm.Sym != nil && (src.Imm.Sym.Pseudo == "FP" || src.Imm.Sym.Pseudo == "SP") {
|
||||
rd := regFromOperand(dst)
|
||||
if rd < 0 {
|
||||
return nil, fmt.Errorf("MOV $%s(%s): invalid destination register", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
|
||||
}
|
||||
_, off := riscvResolvePseudo(src.Imm.Sym, fi)
|
||||
return riscvSPAddiBytes(rd, off), nil
|
||||
}
|
||||
// MOV $sym(FP/SP), rd, not supported: immediate symbol references
|
||||
// other than SB cannot be encoded as a simple immediate.
|
||||
// other than the frame pseudos cannot be encoded as a simple
|
||||
// immediate.
|
||||
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo != "" {
|
||||
return nil, fmt.Errorf("MOV $%s(%s): unsupported immediate symbol reference (only SB is supported)", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
|
||||
}
|
||||
@@ -1344,11 +1411,14 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
|
||||
if rd < 0 {
|
||||
return nil, fmt.Errorf("MOV $imm: invalid destination register")
|
||||
}
|
||||
imm, err := riscvImm32FromOperand(src, false)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
imm := riscvOperandImm64(src)
|
||||
if int64(int32(imm)) != imm {
|
||||
// Beyond the signed 32-bit span the toolchain either builds the
|
||||
// value from a shifted 32-bit part or loads it from the pooled
|
||||
// $i64 constant it synthesises for the purpose.
|
||||
return riscvLoadImm64(rd, imm, lits, relocs), nil
|
||||
}
|
||||
return encodeRISCVLoadImm(rd, imm), nil
|
||||
return encodeRISCVLoadImm(rd, int32(imm)), nil
|
||||
}
|
||||
|
||||
// Memory → register (load).
|
||||
@@ -1557,6 +1627,186 @@ func splitRISCV32Imm(imm int32) (low, high int32) {
|
||||
return low, high
|
||||
}
|
||||
|
||||
// riscvNormalisePseudo rewrites the toolchain's UNDEF spelling onto EBREAK:
|
||||
// the assembler accepts UNDEF where the hardware wants the trap instruction
|
||||
// and emits ebreak (compressed to C.EBREAK under RVC), so every pass sees the
|
||||
// canonical name.
|
||||
func riscvNormalisePseudo(mnem string) string {
|
||||
if strings.EqualFold(mnem, "UNDEF") {
|
||||
return "EBREAK"
|
||||
}
|
||||
return mnem
|
||||
}
|
||||
|
||||
// riscvOperandImm64 reads an immediate operand as a full signed 64-bit value,
|
||||
// where immFromOperand would truncate to int32; the MOV immediate path uses
|
||||
// it to classify the wide constants.
|
||||
func riscvOperandImm64(op *ast.Operand) int64 {
|
||||
if !op.Imm.HasVal {
|
||||
return 0
|
||||
}
|
||||
v := op.Imm.Val
|
||||
if op.Imm.Neg {
|
||||
v = -v
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
// riscvSplitShiftConst mirrors cmd/internal/obj/riscv's splitShiftConst: it
|
||||
// looks for the signed 32-bit integer a constant can be rebuilt from with a
|
||||
// left shift, a left-and-right shift pair (a run of ones), or a zero-extended
|
||||
// 32-bit pattern. A constant that fits none of the shapes is materialised
|
||||
// from the pooled $i64 data symbol instead.
|
||||
func riscvSplitShiftConst(v int64) (imm int64, lsh int, rsh int, ok bool) {
|
||||
// Rebuild from a signed 32-bit integer shifted left.
|
||||
lsh = bits.TrailingZeros64(uint64(v))
|
||||
c := v >> lsh
|
||||
if int64(int32(c)) == c {
|
||||
return c, lsh, 0, true
|
||||
}
|
||||
|
||||
// Rebuild from a small negative constant: shift left into place, then
|
||||
// shift the sign-extended ones run right.
|
||||
rsh = bits.LeadingZeros64(uint64(v))
|
||||
ones := bits.OnesCount64((uint64(v) >> lsh) >> 11)
|
||||
if rsh+ones+lsh+11 == 64 {
|
||||
c = (1<<11 | ((v >> lsh) & 0x7ff)) << 52 >> 52 // sign extend 12 bits
|
||||
if lsh > 0 || c != -1 {
|
||||
lsh += rsh
|
||||
}
|
||||
return c, lsh, rsh, true
|
||||
}
|
||||
|
||||
// Rebuild from a zero-extended signed 32-bit integer.
|
||||
if int64(uint32(c)) == c {
|
||||
c = int64(int32(c))
|
||||
lsh, rsh = 32, 32-lsh
|
||||
return c, lsh, rsh, true
|
||||
}
|
||||
|
||||
return 0, 0, 0, false
|
||||
}
|
||||
|
||||
// riscvSPAddiBytes encodes ADDI rd, SP, imm for the frame-address immediates
|
||||
// (the MOV $sym+off(FP|SP) form), using the compressed forms the toolchain
|
||||
// picks under RVC: C.ADDI4SPN for a positive 4-byte multiple that fits, C.MV
|
||||
// for the zero offset, the plain ADDI otherwise.
|
||||
func riscvSPAddiBytes(rd int, imm int32) []byte {
|
||||
if rd != 0 && imm == 0 {
|
||||
return word16(rvcCR(0x8, uint32(rd), 2)) // C.MV rd, SP
|
||||
}
|
||||
if isRVCIntReg(rd) && imm > 0 && imm < 1024 && imm%4 == 0 {
|
||||
return word16(rvcCIW(0x0, rvcReg3(rd), uint32(imm)))
|
||||
}
|
||||
return wordLE(riscvIType(riscvInstrTable["ADDI"], rd, 2, imm))
|
||||
}
|
||||
|
||||
// riscvMovImm64Size returns the encoded byte length of MOV $imm, rd when the
|
||||
// immediate sits outside the signed 32-bit span: the shifted-part sequences
|
||||
// of riscvLoadImm64, or the 8-byte AUIPC+LD pool load.
|
||||
func riscvMovImm64Size(rd int, imm int64) int {
|
||||
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
|
||||
if !ok {
|
||||
return 8 // AUIPC + LD against the $i64 pool symbol
|
||||
}
|
||||
size := riscvMovImmSize(rd, int32(c))
|
||||
if lsh > 0 {
|
||||
size += riscvShiftImmSize(rd, true)
|
||||
}
|
||||
if rsh > 0 {
|
||||
size += riscvShiftImmSize(rd, false)
|
||||
}
|
||||
return size
|
||||
}
|
||||
|
||||
// riscvShiftImmSize returns the encoded size of one SLLI/SRLI expansion
|
||||
// part: two bytes under RVC when the destination can carry a compressed
|
||||
// shift (C.SLLI admits every register but X0, C.SRLI only X8 to X15), four
|
||||
// otherwise.
|
||||
func riscvShiftImmSize(rd int, left bool) int {
|
||||
if rd != 0 && (left || isRVCIntReg(rd)) {
|
||||
return 2
|
||||
}
|
||||
return 4
|
||||
}
|
||||
|
||||
// riscvLoadImm64 encodes MOV $imm, rd for an immediate beyond the signed
|
||||
// 32-bit span, mirroring the toolchain's instructionsForMOVConst: when a
|
||||
// shifted 32-bit part rebuilds the value it emits that part (compressed like
|
||||
// any written MOV) followed by the SLLI and SRLI shifts; otherwise it loads
|
||||
// the constant from the pooled read-only $i64.<hex> symbol via AUIPC + LD
|
||||
// and registers the literal so the data section carries its bytes.
|
||||
func riscvLoadImm64(rd int, imm int64, lits *riscvLiterals, relocs *[]Reloc) []byte {
|
||||
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
|
||||
if !ok {
|
||||
name := fmt.Sprintf("$i64.%016x", uint64(imm))
|
||||
if lits != nil {
|
||||
lits.add(name, riscvLiteralBytes(imm))
|
||||
}
|
||||
return encodeRISCVSBLoad(&ast.Symbol{Name: name}, rd, relocs)
|
||||
}
|
||||
out := encodeRISCVLoadImm(rd, int32(c))
|
||||
if lsh > 0 {
|
||||
out = append(out, riscvShiftImmBytes(rd, lsh, true)...)
|
||||
}
|
||||
if rsh > 0 {
|
||||
out = append(out, riscvShiftImmBytes(rd, rsh, false)...)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// riscvShiftImmBytes encodes one SLLI (left) or SRLI expansion part, using
|
||||
// the compressed form the toolchain picks under RVC: C.SLLI admits every
|
||||
// register but X0, C.SRLI only X8 to X15.
|
||||
func riscvShiftImmBytes(rd, shamt int, left bool) []byte {
|
||||
if rd != 0 && shamt >= 1 && shamt <= 63 && (left || isRVCIntReg(rd)) {
|
||||
if left {
|
||||
return word16(rvcSLLI(uint32(rd), uint32(shamt)&0x3F))
|
||||
}
|
||||
return word16(rvcCBShift(0x0, rvcReg3(rd), uint32(shamt)&0x3F))
|
||||
}
|
||||
enc := riscvEnc{0x13, 0x1, 0x00} // SLLI
|
||||
imm := int32(shamt)
|
||||
if !left {
|
||||
enc = riscvEnc{0x13, 0x5, 0x00} // SRLI: funct6 000000, funct3 101
|
||||
}
|
||||
return wordLE(riscvIType(enc, rd, rd, imm))
|
||||
}
|
||||
|
||||
// riscvLiteralBytes renders a 64-bit constant as the little-endian bytes the
|
||||
// $i64 pool symbol holds.
|
||||
func riscvLiteralBytes(v int64) []byte {
|
||||
return []byte{byte(v), byte(v >> 8), byte(v >> 16), byte(v >> 24),
|
||||
byte(v >> 32), byte(v >> 40), byte(v >> 48), byte(v >> 56)}
|
||||
}
|
||||
|
||||
// RiscvLiteral is one pooled 64-bit constant: a MOV whose immediate sits
|
||||
// beyond both the 32-bit span and the shift sequences loads its bits from a
|
||||
// read-only data symbol named like the toolchain's $i64 pool.
|
||||
type RiscvLiteral struct {
|
||||
Name string
|
||||
Data []byte
|
||||
}
|
||||
|
||||
// riscvLiterals collects the pooled constants the MOV expansions refer to,
|
||||
// deduplicated by name, in first-use order.
|
||||
type riscvLiterals struct {
|
||||
order []RiscvLiteral
|
||||
seen map[string]bool
|
||||
}
|
||||
|
||||
func (l *riscvLiterals) add(name string, data []byte) {
|
||||
if l.seen == nil {
|
||||
l.seen = map[string]bool{}
|
||||
}
|
||||
if !l.seen[name] {
|
||||
l.seen[name] = true
|
||||
l.order = append(l.order, RiscvLiteral{Name: name, Data: data})
|
||||
}
|
||||
}
|
||||
|
||||
func (l *riscvLiterals) list() []RiscvLiteral { return l.order }
|
||||
|
||||
// encodeRISCVItypeImmediate encodes an I-type arithmetic instruction, expanding
|
||||
// large immediates for ADDI/ANDI/ORI/XORI into LUI+ADDIW+op (or two ADDIs for
|
||||
// ADDI), matching the Go assembler.
|
||||
@@ -1734,6 +1984,7 @@ func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
|
||||
// RVC form. It returns the compressed instruction word and true on success.
|
||||
func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
|
||||
mnem := riscvCompressMnem(instr)
|
||||
mnem = riscvNormalisePseudo(mnem)
|
||||
ops := instr.Operands
|
||||
// The immediate aliases fold onto their I-type mnemonics before
|
||||
// compression: the toolchain compresses ADD $imm, rd as c.addi, exactly
|
||||
|
||||
+1
-1
@@ -63,7 +63,7 @@ func riscvRegNum(name string) int {
|
||||
return 24
|
||||
case "X25", "S9":
|
||||
return 25
|
||||
case "X26", "S10":
|
||||
case "X26", "S10", "CTXT":
|
||||
return 26
|
||||
case "X27", "S11", "g":
|
||||
return 27
|
||||
|
||||
+154
-21
@@ -10,8 +10,8 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// firstTextRISCV parses assembly source and returns the first TEXT function body.
|
||||
@@ -33,7 +33,7 @@ func firstTextRISCV(t *testing.T, src string) *ast.Text {
|
||||
// assembleRISCVHelper assembles one TEXT function and returns its code bytes.
|
||||
func assembleRISCVHelper(t *testing.T, fn *ast.Text) []byte {
|
||||
t.Helper()
|
||||
code, _, _, _, _, err := assembleRISCV(fn)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
@@ -785,16 +785,140 @@ TEXT ·sys(SB), NOSPLIT, $0
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOV_sym_FP_error(t *testing.T) {
|
||||
// MOV $sym(FP), rd should return an error (unsupported).
|
||||
func TestRISCV_MOV_sym_FP(t *testing.T) {
|
||||
// MOV $sym(FP), rd lowers to the frame-adjusted ADDI against SP: the
|
||||
// toolchain's argframe spelling. A zero frame leaves the offset at the
|
||||
// 8-byte link slot, compressed to C.ADDI4SPN.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·badfp(SB), NOSPLIT, $0
|
||||
TEXT ·argfp(SB), NOSPLIT, $0
|
||||
MOV $arg(FP), X10
|
||||
RET
|
||||
`)
|
||||
_, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err == nil {
|
||||
t.Error("expected error for MOV $arg(FP), got nil")
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// prologue (0: leaf, zero frame) + C.ADDI4SPN (2) + RET (4) = 6
|
||||
want := []byte{0x28, 0x00, 0x67, 0x80, 0x00, 0x00}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_Bookkeeping(t *testing.T) {
|
||||
// FUNCDATA and PCDATA contribute no bytes; UNDEF is the toolchain's
|
||||
// ebreak, compressed to C.EBREAK under RVC.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·book(SB), NOSPLIT, $0-8
|
||||
FUNCDATA $0, marks<>(SB)
|
||||
PCDATA $1, $1
|
||||
UNDEF
|
||||
MOV $1, X10
|
||||
MOV X10, ret+0(FP)
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// C.EBREAK (2) + C.LI X10, 1 (2) + C.SWSP (2) + RET (4) = 10: the
|
||||
// FUNCDATA and PCDATA statements contribute nothing.
|
||||
want := []byte{0x02, 0x90, 0x05, 0x45, 0x2a, 0xe4, 0x67, 0x80, 0x00, 0x00}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_JMPPCRel(t *testing.T) {
|
||||
// JMP N(PC): the displacement tracks the instruction N source slots
|
||||
// away in the final layout (0 the jump itself, negative backwards).
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·slots(SB), NOSPLIT, $0-0
|
||||
JMP 2(PC)
|
||||
MOV $1, X11
|
||||
MOV $2, X12
|
||||
MOV X12, X11
|
||||
JMP -3(PC)
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// JMP 2(PC) lands on the C.MV six bytes ahead; JMP -3(PC) lands back on
|
||||
// the first C.LI, six bytes behind.
|
||||
want := []byte{
|
||||
0x6f, 0x00, 0x60, 0x00, // JAL X0, 6
|
||||
0x85, 0x45, // C.LI X11, 1
|
||||
0x09, 0x46, // C.LI X12, 2
|
||||
0xb2, 0x85, // C.MV X11, X12
|
||||
0x6f, 0xf0, 0xbf, 0xff, // JAL X0, -6
|
||||
0x67, 0x80, 0x00, 0x00, // RET
|
||||
}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOVWideImm(t *testing.T) {
|
||||
// Shift-sequence constants compress like the toolchain's expansion.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·wide(SB), NOSPLIT, $0-0
|
||||
MOV $0x8000000000000000, X5
|
||||
MOV $0x100000000, X5
|
||||
MOV $0x000fffffffffffda, X5
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// C.LI -1, C.SLLI 63; C.LI 1, C.SLLI 32; C.LI -19, C.SLLI 13, SRLI 12.
|
||||
want := []byte{
|
||||
0xfd, 0x52, 0xfe, 0x12,
|
||||
0x85, 0x42, 0x82, 0x12,
|
||||
0xb5, 0x52, 0xb6, 0x02, 0x93, 0xd2, 0xc2, 0x00,
|
||||
0x67, 0x80, 0x00, 0x00,
|
||||
}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOVImmPool(t *testing.T) {
|
||||
// A constant outside the shift shapes loads from the pooled $i64 data
|
||||
// symbol via AUIPC+LD, named like the toolchain's pool.
|
||||
src := `#include "textflag.h"
|
||||
TEXT ·pool(SB), NOSPLIT, $0-8
|
||||
MOV $0x0101010101010101, X16
|
||||
MOV X16, ret+0(FP)
|
||||
RET
|
||||
`
|
||||
f, errs := parser.Parse("pool_riscv64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileRISCV(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileRISCV: %v", err)
|
||||
}
|
||||
// AUIPC X16, 0 + LD X16, 0(X16): the relocation pair carries the symbol.
|
||||
wantCode := []byte{0x17, 0x08, 0x00, 0x00, 0x03, 0x38, 0x08, 0x00}
|
||||
if string(img.Code[0:8]) != string(wantCode) {
|
||||
t.Errorf("pool load: got % x", img.Code[0:8])
|
||||
}
|
||||
var lit *DataSymbol
|
||||
for i := range img.DataSyms {
|
||||
if img.DataSyms[i].Name == "$i64.0101010101010101" {
|
||||
lit = &img.DataSyms[i]
|
||||
}
|
||||
}
|
||||
if lit == nil {
|
||||
t.Fatalf("pool symbol missing: %v", img.DataSyms)
|
||||
}
|
||||
wantData := []byte{0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01}
|
||||
if string(img.Data[lit.Offset:lit.Offset+8]) != string(wantData) {
|
||||
t.Errorf("pool bytes: got % x", img.Data[lit.Offset:lit.Offset+8])
|
||||
}
|
||||
}
|
||||
|
||||
@@ -805,7 +929,7 @@ TEXT ·calltest(SB), NOSPLIT, $0
|
||||
CALL ext(SB)
|
||||
RET
|
||||
`)
|
||||
code, _, relocs, _, _, err := assembleRISCV(fn)
|
||||
code, _, relocs, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
@@ -834,7 +958,7 @@ TEXT ·calllocal(SB), NOSPLIT, $0
|
||||
sub:
|
||||
RET
|
||||
`)
|
||||
_, _, _, _, _, err := assembleRISCV(fn)
|
||||
_, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err == nil {
|
||||
t.Error("expected error for CALL to local label, got nil")
|
||||
}
|
||||
@@ -868,7 +992,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
|
||||
t.Helper()
|
||||
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
|
||||
instr := fn.Body[0].(*ast.Instr)
|
||||
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil)
|
||||
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil, nil)
|
||||
}
|
||||
|
||||
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
|
||||
@@ -917,7 +1041,7 @@ func TestRISCVBranchFarBody(t *testing.T) {
|
||||
}
|
||||
sb.WriteString("done:\n\tRET\n")
|
||||
fn := firstTextRISCV(t, sb.String())
|
||||
out, _, _, _, _, err := assembleRISCV(fn)
|
||||
out, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
@@ -943,7 +1067,7 @@ TEXT ·csrhi(SB), NOSPLIT, $0
|
||||
CSRRW $4096, X10, X11
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
t.Error("expected an out-of-range error for CSR $4096, got none")
|
||||
}
|
||||
fn = firstTextRISCV(t, `#include "textflag.h"
|
||||
@@ -951,25 +1075,24 @@ TEXT ·csrmax(SB), NOSPLIT, $0
|
||||
CSRRW $4095, X10, X11
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("CSR $4095 must assemble: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
|
||||
// span are diagnosed instead of silently truncated to their low 32 bits (the
|
||||
// toolchain materialises such constants via SLLI expansion, which this
|
||||
// assembler does not implement).
|
||||
// span are diagnosed instead of silently truncated to their low 32 bits for
|
||||
// the I-type arithmetic; the MOV forms materialise the wide constant instead
|
||||
// (shift sequence or pooled load), like the toolchain.
|
||||
func TestRISCV_Imm64Rejected(t *testing.T) {
|
||||
cases := []string{
|
||||
"MOV $0x123456789, X10",
|
||||
"ADDI $0x100000000, X10, X11",
|
||||
"ANDI $-0x800000001, X10, X11",
|
||||
"SUB $0x100000000, X10, X11",
|
||||
}
|
||||
for _, src := range cases {
|
||||
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
t.Errorf("%s: expected an out-of-range error, got none", src)
|
||||
}
|
||||
}
|
||||
@@ -982,9 +1105,19 @@ TEXT ·edge(SB), NOSPLIT, $0
|
||||
SUB $0x80000000, X12, X13
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("int32-span immediates must assemble: %v", err)
|
||||
}
|
||||
// Beyond the span the MOV forms materialise the constant like the
|
||||
// toolchain instead of diagnosing it.
|
||||
fn = firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·pool(SB), NOSPLIT, $0
|
||||
MOV $0x123456789, X10
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("MOV with a 64-bit immediate must assemble: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// riscvWants decodes code as little-endian words and pins each one; the
|
||||
|
||||
+1
-1
@@ -7,7 +7,7 @@ import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
)
|
||||
|
||||
// RISC-V frame mapping, matching the Go toolchain's riscv64 backend.
|
||||
|
||||
@@ -6,7 +6,7 @@ package asm
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestRISCVFrameSpadjAndLines checks that a framed function records its
|
||||
|
||||
@@ -13,7 +13,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// TestGOObjectRISCVCallReloc checks that CALL sym(SB) emits a single JAL
|
||||
|
||||
+1
-1
@@ -8,7 +8,7 @@
|
||||
// to the arch package; the AST records syntax only.
|
||||
package ast
|
||||
|
||||
import "sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
import "sourcedock.dev/petrbalvin/gasm-sdk/token"
|
||||
|
||||
// File is the parsed representation of one .s source file.
|
||||
type File struct {
|
||||
|
||||
+1
-1
@@ -6,7 +6,7 @@ package ast
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/token"
|
||||
)
|
||||
|
||||
func pos(line, col int) token.Position { return token.Position{Line: line, Column: col} }
|
||||
|
||||
+1
-1
@@ -17,7 +17,7 @@ import (
|
||||
"regexp"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
)
|
||||
|
||||
// go_asm.h is the header the Go compiler writes for every package that
|
||||
|
||||
+57
-3
@@ -5,6 +5,7 @@ package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
@@ -392,14 +393,67 @@ func TestRunCorpusAuditGOOS(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestRunCorpusAuditBuildConstraint covers the //go:build classification end
|
||||
// to end: a generic-named file whose constraint admits one target is
|
||||
// attempted there alone (cpu_x86.s on amd64), and a file whose constraint
|
||||
// admits none of the four targets is never attempted (the msan and
|
||||
// goexperiment trees).
|
||||
func TestRunCorpusAuditBuildConstraint(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
write := func(name, src string) {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
write("x86.s", "//go:build 386 || amd64\n\nTEXT \xc2\xb7f(SB), NOSPLIT, $0\n\tRET\n")
|
||||
write("racey.s", "//go:build race\n\nTEXT \xc2\xb7r(SB), NOSPLIT, $0\n\tRET\n")
|
||||
write("plain.s", "TEXT \xc2\xb7p(SB), NOSPLIT, $0\n\tRET\n")
|
||||
|
||||
stats, err := runCorpusAudit(dir, nil)
|
||||
if err != nil {
|
||||
t.Fatalf("runCorpusAudit: %v", err)
|
||||
}
|
||||
tally := func(name string) *corpusTally {
|
||||
for i, tg := range stats.targets {
|
||||
if tg.name == name {
|
||||
return stats.tallies[i]
|
||||
}
|
||||
}
|
||||
t.Fatalf("no tally for %s", name)
|
||||
return nil
|
||||
}
|
||||
if stats.narrowed != 1 || stats.excluded != 1 || stats.generic != 1 {
|
||||
t.Errorf("buckets = narrowed %d, excluded %d, generic %d; want 1, 1, 1", stats.narrowed, stats.excluded, stats.generic)
|
||||
}
|
||||
if a := tally("amd64"); a.attempted != 2 || a.assembled != 2 {
|
||||
t.Errorf("amd64 = %d/%d, want 2/2 (x86.s and plain.s)", a.assembled, a.attempted)
|
||||
}
|
||||
for _, name := range []string{"arm64", "riscv64", "loong64"} {
|
||||
if a := tally(name); a.attempted != 1 || a.assembled != 1 {
|
||||
t.Errorf("%s = %d/%d, want 1/1 (plain.s only)", name, a.assembled, a.attempted)
|
||||
}
|
||||
}
|
||||
if stats.full != 2 {
|
||||
t.Errorf("full = %d, want 2 (x86.s over its one target, plain.s over all four)", stats.full)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
|
||||
// the runtime package, whose header the toolchain's own -asmhdr output was
|
||||
// sampled from. Skipped in short mode: it type-checks the whole package.
|
||||
// the runtime package of the ambient toolchain, whose header the toolchain's
|
||||
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
|
||||
// the whole package. The GOROOT comes from the go command itself, so the
|
||||
// test follows whatever toolchain the host provides.
|
||||
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("type-checks the whole runtime package")
|
||||
}
|
||||
dir, err := generateGoAsmHeader("/usr/local/go/src/runtime", "", "amd64", t.TempDir())
|
||||
out, err := exec.Command("go", "env", "GOROOT").Output()
|
||||
if err != nil {
|
||||
t.Skipf("no Go toolchain: %v", err)
|
||||
}
|
||||
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
|
||||
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
|
||||
}
|
||||
|
||||
+170
-34
@@ -5,6 +5,7 @@ package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"go/build/constraint"
|
||||
"maps"
|
||||
"os"
|
||||
"os/exec"
|
||||
@@ -15,9 +16,9 @@ import (
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// cmdAuditInstructions cross-checks a gasm encoder against the Go toolchain's
|
||||
@@ -38,7 +39,7 @@ import (
|
||||
// construction and are excluded from the diff; the other architectures list
|
||||
// their conditional branches outright.
|
||||
func cmdAuditInstructions(args []string) error {
|
||||
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]", `
|
||||
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
|
||||
Compare the gasm encoder for the given architecture (default amd64) against
|
||||
go tool asm and print the diff: superset encodings (gasm-only, shippable via
|
||||
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
|
||||
@@ -55,16 +56,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
|
||||
attempted for that architecture; a file without one is attempted for all
|
||||
four, exactly as a GOARCH build would compile it. The report gives the
|
||||
per-architecture pass rates and the most common failure reasons, which drive
|
||||
the encodability backlog by frequency rather than by table order.
|
||||
the encodability backlog by frequency rather than by table order. With
|
||||
-list the report also prints every failing file with its reason, per
|
||||
architecture.
|
||||
`)
|
||||
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
|
||||
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
|
||||
var dirs includeDirs
|
||||
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return err
|
||||
}
|
||||
if *corpus {
|
||||
return cmdAuditCorpus(fs.Args(), dirs)
|
||||
return cmdAuditCorpus(fs.Args(), dirs, *list)
|
||||
}
|
||||
archName := "amd64"
|
||||
switch n := len(fs.Args()); {
|
||||
@@ -403,20 +407,30 @@ type corpusTally struct {
|
||||
assembled int
|
||||
reasons map[string]int // failure reason → count
|
||||
example map[string]string // failure reason → one representative file
|
||||
fails []corpusFailure // every failure, in file order, for --list
|
||||
}
|
||||
|
||||
func (t *corpusTally) fail(path, reason string) {
|
||||
// corpusFailure is one failed attempt, recorded for the --list report.
|
||||
type corpusFailure struct {
|
||||
path string
|
||||
reason string
|
||||
detail string
|
||||
}
|
||||
|
||||
func (t *corpusTally) fail(path string, err error) {
|
||||
reason := corpusReason(err)
|
||||
t.reasons[reason]++
|
||||
if t.example[reason] == "" {
|
||||
t.example[reason] = path
|
||||
}
|
||||
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
|
||||
}
|
||||
|
||||
// cmdAuditCorpus implements audit-instructions --corpus. The include
|
||||
// directories carry #include resolution over a corpus whose files refer to
|
||||
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
|
||||
// needs.
|
||||
func cmdAuditCorpus(args []string, dirs includeDirs) error {
|
||||
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
|
||||
if len(args) > 1 {
|
||||
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
|
||||
}
|
||||
@@ -454,7 +468,7 @@ func cmdAuditCorpus(args []string, dirs includeDirs) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
printCorpusStats(stats)
|
||||
printCorpusStats(stats, list)
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -463,8 +477,10 @@ type corpusStats struct {
|
||||
root string
|
||||
files int
|
||||
generic int // files attempted for all four architectures
|
||||
narrowed int // files whose //go:build admits a proper subset of the four
|
||||
excluded int // files whose //go:build admits none of the four: never compiled
|
||||
otherPort int // files named for another Go port: never attempted
|
||||
full int // files that assembled for every target architecture
|
||||
full int // files that assembled for every applicable target architecture
|
||||
targets []corpusTarget
|
||||
tallies []*corpusTally
|
||||
}
|
||||
@@ -475,8 +491,8 @@ type corpusStats struct {
|
||||
// set, even when gasm does not support the architecture.
|
||||
var goPortSuffixes = []string{
|
||||
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
|
||||
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
|
||||
"s390x", "wasm",
|
||||
"mips64le", "mipsle", "mips64x", "mipsx", "ppc64", "ppc64le",
|
||||
"ppc64x", "riscv", "riscv64", "s390x", "wasm",
|
||||
}
|
||||
|
||||
// otherPortFile reports whether the file belongs to a build no supported
|
||||
@@ -560,6 +576,57 @@ func otherGOOSFile(path string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// buildConstraint returns the file's leading //go:build expression, or nil
|
||||
// when the file carries none. The constraint governs the same header block
|
||||
// go/build reads: blank lines and comments may precede it, and the first
|
||||
// line that is neither ends the block. A constraint that does not parse
|
||||
// narrows nothing, so the file stays in the attempted set: the audit must
|
||||
// never exclude a file the toolchain would compile.
|
||||
func buildConstraint(src string) constraint.Expr {
|
||||
for line := range strings.SplitSeq(src, "\n") {
|
||||
t := strings.TrimSpace(line)
|
||||
switch {
|
||||
case t == "":
|
||||
continue
|
||||
case strings.HasPrefix(t, "//"):
|
||||
if constraint.IsGoBuild(t) {
|
||||
e, err := constraint.Parse(t)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
return e
|
||||
}
|
||||
continue
|
||||
default:
|
||||
return nil
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// unixOS is go/build's unixOS set: the GOOSes the unix build tag admits.
|
||||
var unixOS = map[string]bool{
|
||||
"aix": true, "android": true, "darwin": true, "dragonfly": true,
|
||||
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
|
||||
"linux": true, "netbsd": true, "openbsd": true, "solaris": true,
|
||||
}
|
||||
|
||||
// constraintTags answers the build tags a plain `go build` sets for a
|
||||
// target: the GOOS and GOARCH, gc, and unix on the unix-like GOOSes. No
|
||||
// experiment, sanitiser or cgo tag is ever true: the audit models the
|
||||
// default build, and no GOROOT assembly file's constraint hinges on cgo.
|
||||
func constraintTags(goarch, goos string) func(string) bool {
|
||||
return func(tag string) bool {
|
||||
switch tag {
|
||||
case goarch, goos, "gc":
|
||||
return true
|
||||
case "unix":
|
||||
return unixOS[goos]
|
||||
}
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
files, err := asmFiles(root)
|
||||
if err != nil {
|
||||
@@ -577,8 +644,8 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
tallies[i] = &corpusTally{reasons: map[string]int{}, example: map[string]string{}}
|
||||
}
|
||||
// full is the north-star number: a file counts when every architecture
|
||||
// its name allows assembles it.
|
||||
full, generic, otherPort := 0, 0, 0
|
||||
// its build admits assembles it.
|
||||
full, generic, otherPort, narrowedCount, excluded := 0, 0, 0, 0, 0
|
||||
|
||||
// Header generation is created on first use, so a corpus with no
|
||||
// go_asm.h includes never pays for a temp directory.
|
||||
@@ -601,28 +668,78 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
// invisible to a file-name rule).
|
||||
goos := goosFromFilename(path)
|
||||
|
||||
// The GOOS the header generation type-checks under follows the
|
||||
// file's name when the name carries one; the ambient GOOS is the
|
||||
// honest guess otherwise.
|
||||
namedArch := arch.FromFilename(path)
|
||||
var wanted []int // indexes into targets
|
||||
if a := arch.FromFilename(path); a != arch.Unknown {
|
||||
other := false
|
||||
switch {
|
||||
case namedArch != arch.Unknown:
|
||||
for i, tg := range targets {
|
||||
if tg.a == a {
|
||||
if tg.a == namedArch {
|
||||
wanted = append(wanted, i)
|
||||
}
|
||||
}
|
||||
} else if otherPortFile(path) {
|
||||
case otherPortFile(path):
|
||||
// A file named for a Go port gasm does not support (arm,
|
||||
// 386, s390x, ...) or for another GOOS is compiled by no
|
||||
// supported-arch build, so it is neither generic nor a
|
||||
// per-arch attempt: counting it as generic would make the
|
||||
// headline unreachably low for reasons no supported target
|
||||
// can fix.
|
||||
other = true
|
||||
otherPort++
|
||||
} else {
|
||||
generic++
|
||||
default:
|
||||
for i := range targets {
|
||||
wanted = append(wanted, i)
|
||||
}
|
||||
}
|
||||
|
||||
// A //go:build constraint narrows the set of targets the file is
|
||||
// assembled for, the way the go command compiles the file only for
|
||||
// the targets the expression admits: cpu_x86.s belongs to the x86
|
||||
// build alone, and a file whose constraint admits none of the four
|
||||
// targets (the goexperiment.runtimesecret and msan trees) is
|
||||
// compiled by no supported build. The tags mirror what a plain
|
||||
// `go build` sets: the GOOS and GOARCH, gc, and unix on the
|
||||
// unix-like GOOSes; no experiment, sanitiser or cgo tag is ever
|
||||
// true. The GOOS is the file's own when the name carries one,
|
||||
// else the ambient one.
|
||||
goosForEval := goos
|
||||
if goosForEval == "" {
|
||||
goosForEval = runtime.GOOS
|
||||
}
|
||||
narrowed := false
|
||||
if len(wanted) > 0 {
|
||||
if ce := buildConstraint(src); ce != nil {
|
||||
kept := make([]int, 0, len(wanted))
|
||||
for _, i := range wanted {
|
||||
tg := targets[i]
|
||||
if ce.Eval(constraintTags(goarchName(tg.a), goosForEval)) {
|
||||
kept = append(kept, i)
|
||||
}
|
||||
}
|
||||
if len(kept) < len(wanted) {
|
||||
narrowed = true
|
||||
}
|
||||
wanted = kept
|
||||
}
|
||||
}
|
||||
|
||||
switch {
|
||||
case other:
|
||||
// already tallied above
|
||||
case len(wanted) == 0:
|
||||
excluded++
|
||||
case namedArch != arch.Unknown:
|
||||
// a per-arch attempt over the constraint's subset
|
||||
case narrowed:
|
||||
narrowedCount++
|
||||
default:
|
||||
generic++
|
||||
}
|
||||
|
||||
// A file that includes go_asm.h parses against a per-target header:
|
||||
// the defines differ per architecture (internal/cpu's layout, for
|
||||
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
|
||||
@@ -645,21 +762,22 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
|
||||
if err != nil {
|
||||
ok = false
|
||||
t.fail(path, corpusReason(err))
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{
|
||||
Expand: true,
|
||||
IncludeDirs: append(slices.Clone(dirs), hdrDir),
|
||||
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
|
||||
})
|
||||
if len(errs) > 0 {
|
||||
ok = false
|
||||
t.fail(path, corpusReason(errs[0]))
|
||||
t.fail(path, errs[0])
|
||||
continue
|
||||
}
|
||||
if _, err := assembleFile(tg.a, f); err != nil {
|
||||
if _, err := assembleFile(tg.a, f, goos); err != nil {
|
||||
ok = false
|
||||
t.fail(path, corpusReason(err))
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
t.assembled++
|
||||
@@ -670,21 +788,28 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
continue
|
||||
}
|
||||
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
|
||||
|
||||
ok := true
|
||||
for _, i := range wanted {
|
||||
tg, t := targets[i], tallies[i]
|
||||
t.attempted++
|
||||
// The parse carries the target's platform predefines, so it
|
||||
// cannot be shared across targets the way a header-free file's
|
||||
// could: a #ifdef GOARCH_arm block must be live on arm64 and
|
||||
// dead everywhere else.
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{
|
||||
Expand: true,
|
||||
IncludeDirs: dirs,
|
||||
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
|
||||
})
|
||||
var err error
|
||||
if len(errs) > 0 {
|
||||
err = errs[0] // a parse failure is a failure for every target
|
||||
} else {
|
||||
_, err = assembleFile(tg.a, f)
|
||||
_, err = assembleFile(tg.a, f, goos)
|
||||
}
|
||||
if err != nil {
|
||||
ok = false
|
||||
t.fail(path, corpusReason(err))
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
t.assembled++
|
||||
@@ -698,6 +823,8 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
root: root,
|
||||
files: len(files),
|
||||
generic: generic,
|
||||
narrowed: narrowedCount,
|
||||
excluded: excluded,
|
||||
otherPort: otherPort,
|
||||
full: full,
|
||||
targets: targets,
|
||||
@@ -706,14 +833,16 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
}
|
||||
|
||||
// printCorpusStats renders the corpus audit report.
|
||||
func printCorpusStats(s *corpusStats) {
|
||||
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
|
||||
func printCorpusStats(s *corpusStats, list bool) {
|
||||
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d narrowed by //go:build; %d excluded by //go:build; %d named for other Go ports, never attempted)\n",
|
||||
s.root, s.files, s.generic, s.narrowed, s.excluded, s.otherPort)
|
||||
// The rate is over the files a supported build would attempt: the
|
||||
// other ports' files sit in the count for completeness but can never
|
||||
// assemble, so counting them in the denominator would report the gap
|
||||
// of architectures gasm deliberately does not target.
|
||||
attemptable := max(s.files-s.otherPort, 1)
|
||||
fmt.Printf(" assemble for every target architecture: %d of %d attemptable (%.1f%%)\n", s.full, attemptable, 100*float64(s.full)/float64(attemptable))
|
||||
// other ports' files and the ones no supported target compiles sit in
|
||||
// the count for completeness but can never assemble, so counting them
|
||||
// in the denominator would report the gap of platforms gasm
|
||||
// deliberately does not target.
|
||||
attemptable := max(s.files-s.otherPort-s.excluded, 1)
|
||||
fmt.Printf(" assemble for every applicable target: %d of %d attemptable (%.1f%%)\n", s.full, attemptable, 100*float64(s.full)/float64(attemptable))
|
||||
for i, tg := range s.targets {
|
||||
t := s.tallies[i]
|
||||
fmt.Printf(" %s: %d/%d attempted\n", tg.name, t.assembled, t.attempted)
|
||||
@@ -721,6 +850,13 @@ func printCorpusStats(s *corpusStats) {
|
||||
fmt.Printf(" %4d %s\n", t.reasons[r], r)
|
||||
fmt.Printf(" e.g. %s\n", t.example[r])
|
||||
}
|
||||
if !list {
|
||||
continue
|
||||
}
|
||||
for _, f := range t.fails {
|
||||
fmt.Printf(" FAIL %s\n", f.path)
|
||||
fmt.Printf(" %s: %s\n", f.reason, f.detail)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+61
-1
@@ -4,9 +4,10 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"runtime"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
)
|
||||
|
||||
func TestDerivedFamily(t *testing.T) {
|
||||
@@ -64,3 +65,62 @@ func TestGasmEncodable(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestBuildConstraint pins the //go:build reader: the constraint governs the
|
||||
// leading comment block, the first non-comment line ends it (a tag below a
|
||||
// #include governs nothing, exactly as go/build drops it), and a file
|
||||
// without one admits every target.
|
||||
func TestBuildConstraint(t *testing.T) {
|
||||
admits := func(src, goarch, goos string) bool {
|
||||
t.Helper()
|
||||
e := buildConstraint(src)
|
||||
if e == nil {
|
||||
return true
|
||||
}
|
||||
return e.Eval(constraintTags(goarch, goos))
|
||||
}
|
||||
const ret = "TEXT \xc2\xb7f(SB), NOSPLIT, $0\n\tRET\n"
|
||||
cases := []struct {
|
||||
name string
|
||||
src string
|
||||
amd64, arm64 bool
|
||||
}{
|
||||
{"no constraint", ret, true, true},
|
||||
{"x86 only", "//go:build 386 || amd64\n\n" + ret, true, false},
|
||||
{"arm64 and linux", "//go:build arm64 && linux\n\n" + ret, false, true},
|
||||
{"msan never", "//go:build msan\n\n" + ret, false, false},
|
||||
{"experiment never", "//go:build goexperiment.runtimesecret\n\n" + ret, false, false},
|
||||
{"below an include governs nothing", "#include \"textflag.h\"\n//go:build amd64\n" + ret, true, true},
|
||||
{"unparsable narrows nothing", "//go:build (amd64\n" + ret, true, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
if got := admits(c.src, "amd64", runtime.GOOS); got != c.amd64 {
|
||||
t.Errorf("amd64 admission = %v, want %v", got, c.amd64)
|
||||
}
|
||||
if got := admits(c.src, "arm64", runtime.GOOS); got != c.arm64 {
|
||||
t.Errorf("arm64 admission = %v, want %v", got, c.arm64)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestConstraintTags pins the tag set a plain `go build` sets: the GOOS and
|
||||
// GOARCH, gc, unix on the unix-like GOOSes; nothing else is ever true.
|
||||
func TestConstraintTags(t *testing.T) {
|
||||
ok := constraintTags("amd64", "linux")
|
||||
for _, tag := range []string{"amd64", "linux", "gc", "unix"} {
|
||||
if !ok(tag) {
|
||||
t.Errorf("tag %q = false, want true", tag)
|
||||
}
|
||||
}
|
||||
for _, tag := range []string{"arm64", "freebsd", "darwin", "cgo", "race", "msan", "goexperiment.runtimesecret"} {
|
||||
if ok(tag) {
|
||||
t.Errorf("tag %q = true, want false", tag)
|
||||
}
|
||||
}
|
||||
fb := constraintTags("arm64", "freebsd")
|
||||
if !fb("unix") {
|
||||
t.Error("unix on freebsd = false, want true")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build !linux
|
||||
//go:build !(linux || (freebsd && (amd64 || arm64 || riscv64)))
|
||||
|
||||
package main
|
||||
|
||||
@@ -11,6 +11,6 @@ import (
|
||||
)
|
||||
|
||||
func cmdDebug(args []string) int {
|
||||
fmt.Fprintln(os.Stderr, "gasm debug: the interactive debugger requires Linux (ptrace)")
|
||||
fmt.Fprintln(os.Stderr, "gasm debug: the interactive debugger requires Linux or FreeBSD (ptrace)")
|
||||
return 1
|
||||
}
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build linux
|
||||
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
|
||||
|
||||
package main
|
||||
|
||||
@@ -13,8 +13,8 @@ import (
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/debug"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/debug"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
func cmdDebug(args []string) int {
|
||||
+4
-4
@@ -9,9 +9,9 @@ import (
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// cmdDis disassembles machine code: either a raw binary (standard input with
|
||||
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
|
||||
if len(errs) > 0 {
|
||||
return 1
|
||||
}
|
||||
img, err := assembleFile(target, f)
|
||||
img, err := assembleFile(target, f, "")
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
|
||||
return 1
|
||||
|
||||
+43
-20
@@ -27,15 +27,15 @@ import (
|
||||
"sync"
|
||||
"syscall"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/format"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/lint"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/lsp"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/format"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/lexer"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/lint"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/lsp"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// version reports the release the toolchain recorded for this build: the
|
||||
@@ -574,7 +574,7 @@ naming the package.
|
||||
defer cleanup()
|
||||
dirs = append(dirs, hdrDir)
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
|
||||
for _, e := range errs {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
|
||||
}
|
||||
@@ -582,7 +582,7 @@ naming the package.
|
||||
return 1
|
||||
}
|
||||
|
||||
img, err := assembleFile(targetArch, f)
|
||||
img, err := assembleFile(targetArch, f, goos)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
|
||||
return 1
|
||||
@@ -789,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
|
||||
return 1
|
||||
}
|
||||
|
||||
// platformPredefines mirrors the go command's assembler invocation, which
|
||||
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
|
||||
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
|
||||
// exactly those names, so an assembler without them cannot see the platform
|
||||
// definitions at all.
|
||||
func platformPredefines(goarch, goos string) map[string]string {
|
||||
return map[string]string{
|
||||
"GOARCH_" + goarch: "1",
|
||||
"GOOS_" + goos: "1",
|
||||
}
|
||||
}
|
||||
|
||||
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
|
||||
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
|
||||
// and nothing else.
|
||||
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
|
||||
goos := fileGoos
|
||||
if goos == "" {
|
||||
goos = runtime.GOOS
|
||||
}
|
||||
return platformPredefines(goarch, goos)
|
||||
}
|
||||
|
||||
// assembleFile assembles a parsed file for the given architecture and returns the image.
|
||||
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
|
||||
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
|
||||
switch targetArch {
|
||||
case arch.AMD64:
|
||||
return asm.AssembleFile(f)
|
||||
return asm.AssembleFile(f, asm.WithGOOS(goos))
|
||||
case arch.RISCV:
|
||||
return asm.AssembleFileRISCV(f)
|
||||
case arch.ARM64:
|
||||
@@ -813,18 +836,18 @@ func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image,
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
|
||||
target := forced
|
||||
if target == arch.Unknown {
|
||||
target = arch.FromFilename(path)
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
|
||||
for _, e := range errs {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
|
||||
}
|
||||
if len(errs) > 0 {
|
||||
return nil, fmt.Errorf("parse errors")
|
||||
}
|
||||
target := forced
|
||||
if target == arch.Unknown {
|
||||
target = arch.FromFilename(path)
|
||||
}
|
||||
return assembleFile(target, f)
|
||||
return assembleFile(target, f, "")
|
||||
}
|
||||
|
||||
// printByteDiff shows the first few byte differences between two code blocks.
|
||||
@@ -920,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
|
||||
if len(errs) > 0 {
|
||||
return 1
|
||||
}
|
||||
img, err := assembleFile(targetArch, f)
|
||||
img, err := assembleFile(targetArch, f, "")
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
|
||||
return 1
|
||||
|
||||
@@ -14,8 +14,8 @@ import (
|
||||
"syscall"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
)
|
||||
|
||||
const clean = "#include \"textflag.h\"\n" +
|
||||
|
||||
@@ -11,8 +11,8 @@ import (
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
gasmast "sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
gasmparser "sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
gasmast "sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
gasmparser "sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
// cmdScaffold generates a differential test skeleton for every kernel in a
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build linux
|
||||
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
|
||||
|
||||
package debug
|
||||
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
// memory and returns its text representation and length in bytes.
|
||||
func (s *Session) Disassemble(addr uint64) (string, int, error) {
|
||||
mem, err := s.ReadMemory(addr, 15)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
ins, err := disasm.Decode(arch.AMD64, mem, addr)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
return ins.Text, ins.Len, nil
|
||||
}
|
||||
|
||||
// DisassembleN decodes up to n instructions starting at addr and returns
|
||||
// them as a formatted string with addresses and byte offsets.
|
||||
func (s *Session) DisassembleN(addr uint64, n int) string {
|
||||
var result strings.Builder
|
||||
pc := addr
|
||||
for range n {
|
||||
text, length, err := s.Disassemble(pc)
|
||||
if err != nil {
|
||||
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
|
||||
break
|
||||
}
|
||||
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
|
||||
if length == 0 {
|
||||
length = 1
|
||||
}
|
||||
pc += uint64(length)
|
||||
}
|
||||
return result.String()
|
||||
}
|
||||
|
||||
// isCallInsn reports whether disassembled text (x86asm.IntelSyntax) is a
|
||||
// call. The first token must match exactly: a prefix test would also catch
|
||||
// unrelated mnemonics.
|
||||
func isCallInsn(text string) bool {
|
||||
m, _, _ := strings.Cut(text, " ")
|
||||
return strings.ToLower(m) == "call"
|
||||
}
|
||||
@@ -0,0 +1,60 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && arm64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
// memory and returns its text representation and length in bytes.
|
||||
func (s *Session) Disassemble(addr uint64) (string, int, error) {
|
||||
mem, err := s.ReadMemory(addr, 4)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
ins, err := disasm.Decode(arch.ARM64, mem, addr)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
return ins.Text, ins.Len, nil
|
||||
}
|
||||
|
||||
// DisassembleN decodes up to n instructions starting at addr and returns
|
||||
// them as a formatted string with addresses and byte offsets.
|
||||
func (s *Session) DisassembleN(addr uint64, n int) string {
|
||||
var result strings.Builder
|
||||
pc := addr
|
||||
for range n {
|
||||
text, length, err := s.Disassemble(pc)
|
||||
if err != nil {
|
||||
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
|
||||
break
|
||||
}
|
||||
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
|
||||
if length == 0 {
|
||||
length = 1
|
||||
}
|
||||
pc += uint64(length)
|
||||
}
|
||||
return result.String()
|
||||
}
|
||||
|
||||
// isCallInsn reports whether disassembled text (arm64asm.GoSyntax) is a
|
||||
// call. GoSyntax renders bl as CALL; the native mnemonic is accepted too.
|
||||
// The first token must match exactly so branches never match.
|
||||
func isCallInsn(text string) bool {
|
||||
m, _, _ := strings.Cut(text, " ")
|
||||
switch strings.ToLower(m) {
|
||||
case "call", "bl":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && riscv64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
// memory and returns its text representation and length in bytes.
|
||||
func (s *Session) Disassemble(addr uint64) (string, int, error) {
|
||||
mem, err := s.ReadMemory(addr, 4)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
ins, err := disasm.Decode(arch.RISCV, mem, addr)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
return ins.Text, ins.Len, nil
|
||||
}
|
||||
|
||||
// DisassembleN decodes up to n instructions starting at addr and returns
|
||||
// them as a formatted string with addresses and byte offsets.
|
||||
func (s *Session) DisassembleN(addr uint64, n int) string {
|
||||
var result strings.Builder
|
||||
pc := addr
|
||||
for range n {
|
||||
text, length, err := s.Disassemble(pc)
|
||||
if err != nil {
|
||||
result.WriteString(fmt.Sprintf(" %#08x: <error: %v>\n", pc, err))
|
||||
break
|
||||
}
|
||||
result.WriteString(fmt.Sprintf(" %#08x: %s\n", pc, text))
|
||||
if length == 0 {
|
||||
length = 1
|
||||
}
|
||||
pc += uint64(length)
|
||||
}
|
||||
return result.String()
|
||||
}
|
||||
|
||||
// isCallInsn reports whether disassembled text (riscv64asm.GoSyntax) is a
|
||||
// call. GoSyntax renders jal and jalr calls as CALL; the native mnemonics
|
||||
// are accepted too. The first token must match exactly: a prefix test on
|
||||
// "bl" would catch branches on other architectures, and jalr as ret prints
|
||||
// RET, which must not be stepped over.
|
||||
func isCallInsn(text string) bool {
|
||||
m, _, _ := strings.Cut(text, " ")
|
||||
switch strings.ToLower(m) {
|
||||
case "call", "jal", "jalr":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -9,8 +9,8 @@ import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
|
||||
@@ -9,8 +9,8 @@ import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
|
||||
@@ -9,8 +9,8 @@ import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
|
||||
@@ -9,8 +9,8 @@ import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/disasm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/disasm"
|
||||
)
|
||||
|
||||
// Disassemble decodes the instruction at the given address in the debuggee's
|
||||
|
||||
@@ -0,0 +1,136 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
import "fmt"
|
||||
|
||||
func printRegs(regs *Regs, codeBase, funcOff uint64) {
|
||||
fmt.Printf(" RIP = %#016x (func+%#x)\n", regs.RIP, regs.RIP-codeBase-funcOff)
|
||||
fmt.Printf(" RSP = %#016x RBP = %#016x\n", regs.RSP, regs.RBP)
|
||||
fmt.Printf(" RAX = %#016x RBX = %#016x\n", regs.RAX, regs.RBX)
|
||||
fmt.Printf(" RCX = %#016x RDX = %#016x\n", regs.RCX, regs.RDX)
|
||||
fmt.Printf(" RSI = %#016x RDI = %#016x\n", regs.RSI, regs.RDI)
|
||||
fmt.Printf(" R8 = %#016x R9 = %#016x\n", regs.R8, regs.R9)
|
||||
fmt.Printf(" R10 = %#016x R11 = %#016x\n", regs.R10, regs.R11)
|
||||
fmt.Printf(" R12 = %#016x R13 = %#016x\n", regs.R12, regs.R13)
|
||||
fmt.Printf(" R14 = %#016x R15 = %#016x\n", regs.R14, regs.R15)
|
||||
fmt.Printf(" RFLAGS = %#x [%s]\n", regs.RFLAGS, decodeRflags(regs.RFLAGS))
|
||||
}
|
||||
|
||||
func printVectorRegs(v *VectorRegs) {
|
||||
fmt.Println("\n Vector registers (YMM):")
|
||||
for i := 0; i < 16; i += 2 {
|
||||
fmt.Printf(" YMM%-2d = ", i)
|
||||
printYMM(v.YMM[i][:])
|
||||
fmt.Printf(" YMM%-2d = ", i+1)
|
||||
printYMM(v.YMM[i+1][:])
|
||||
fmt.Println()
|
||||
}
|
||||
}
|
||||
|
||||
func printYMM(b []byte) {
|
||||
for j := 0; j < 32; j += 4 {
|
||||
v := uint32(b[j]) | uint32(b[j+1])<<8 | uint32(b[j+2])<<16 | uint32(b[j+3])<<24
|
||||
fmt.Printf("%08x ", v)
|
||||
}
|
||||
}
|
||||
|
||||
func decodeRflags(f uint64) string {
|
||||
var flags string
|
||||
if f&1 != 0 {
|
||||
flags += "CF "
|
||||
}
|
||||
if f&(1<<2) != 0 {
|
||||
flags += "PF "
|
||||
}
|
||||
if f&(1<<4) != 0 {
|
||||
flags += "AF "
|
||||
}
|
||||
if f&(1<<6) != 0 {
|
||||
flags += "ZF "
|
||||
}
|
||||
if f&(1<<7) != 0 {
|
||||
flags += "SF "
|
||||
}
|
||||
if f&(1<<8) != 0 {
|
||||
flags += "TF "
|
||||
}
|
||||
if f&(1<<9) != 0 {
|
||||
flags += "IF "
|
||||
}
|
||||
if f&(1<<10) != 0 {
|
||||
flags += "DF "
|
||||
}
|
||||
if f&(1<<11) != 0 {
|
||||
flags += "OF "
|
||||
}
|
||||
if flags == "" {
|
||||
return "none"
|
||||
}
|
||||
return flags[:len(flags)-1]
|
||||
}
|
||||
|
||||
// SetReg modifies a register value in the debuggee.
|
||||
func (s *Session) SetReg(name string, value uint64) error {
|
||||
regs, err := s.GetRegs()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
switch name {
|
||||
case "rax", "eax", "ax", "al":
|
||||
regs.RAX = value
|
||||
case "rbx", "ebx", "bx", "bl":
|
||||
regs.RBX = value
|
||||
case "rcx", "ecx", "cx", "cl":
|
||||
regs.RCX = value
|
||||
case "rdx", "edx", "dx", "dl":
|
||||
regs.RDX = value
|
||||
case "rsi", "esi", "si":
|
||||
regs.RSI = value
|
||||
case "rdi", "edi", "di":
|
||||
regs.RDI = value
|
||||
case "rbp", "ebp", "bp":
|
||||
regs.RBP = value
|
||||
case "rsp", "esp", "sp":
|
||||
regs.RSP = value
|
||||
case "r8":
|
||||
regs.R8 = value
|
||||
case "r9":
|
||||
regs.R9 = value
|
||||
case "r10":
|
||||
regs.R10 = value
|
||||
case "r11":
|
||||
regs.R11 = value
|
||||
case "r12":
|
||||
regs.R12 = value
|
||||
case "r13":
|
||||
regs.R13 = value
|
||||
case "r14":
|
||||
regs.R14 = value
|
||||
case "r15":
|
||||
regs.R15 = value
|
||||
case "rip", "eip":
|
||||
regs.RIP = value
|
||||
default:
|
||||
return fmt.Errorf("debug: unknown register %q", name)
|
||||
}
|
||||
return s.SetRegs(®s)
|
||||
}
|
||||
|
||||
// archReturnAddr reads the return address of the current frame (amd64
|
||||
// ABI0 convention). A function that contains a CALL (or has a frame) is
|
||||
// assembled with the prologue PUSHQ BP; MOVQ SP, BP, so mid-function the
|
||||
// word at SP is the saved caller BP, a stack address, and the return
|
||||
// address sits further up. Walk the stack from SP and take the first word
|
||||
// that lies in an executable mapping: stack and data words never do, a
|
||||
// return address always does. FreeBSD exposes no mapping list, so the
|
||||
// walk degenerates to the raw entry convention, [SP] before any push.
|
||||
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
|
||||
return s.Peek(regs.RSP)
|
||||
}
|
||||
|
||||
// archSPLabel returns the SP register name for display.
|
||||
func archSPLabel() string { return "RSP" }
|
||||
@@ -0,0 +1,127 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && arm64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
)
|
||||
|
||||
func printRegs(regs *Regs, codeBase, funcOff uint64) {
|
||||
fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff)
|
||||
fmt.Printf(" SP = %#016x FP = %#016x\n", regs.SP, regs.X29)
|
||||
fmt.Printf(" LR = %#016x\n", regs.X30)
|
||||
fmt.Printf(" X0 = %#016x X1 = %#016x\n", regs.X0, regs.X1)
|
||||
fmt.Printf(" X2 = %#016x X3 = %#016x\n", regs.X2, regs.X3)
|
||||
fmt.Printf(" X4 = %#016x X5 = %#016x\n", regs.X4, regs.X5)
|
||||
fmt.Printf(" X6 = %#016x X7 = %#016x\n", regs.X6, regs.X7)
|
||||
fmt.Printf(" X8 = %#016x X9 = %#016x\n", regs.X8, regs.X9)
|
||||
fmt.Printf(" X10 = %#016x X11 = %#016x\n", regs.X10, regs.X11)
|
||||
fmt.Printf(" X12 = %#016x X13 = %#016x\n", regs.X12, regs.X13)
|
||||
fmt.Printf(" X14 = %#016x X15 = %#016x\n", regs.X14, regs.X15)
|
||||
fmt.Printf(" X16 = %#016x X17 = %#016x\n", regs.X16, regs.X17)
|
||||
fmt.Printf(" X18 = %#016x X19 = %#016x\n", regs.X18, regs.X19)
|
||||
fmt.Printf(" X20 = %#016x X21 = %#016x\n", regs.X20, regs.X21)
|
||||
fmt.Printf(" X22 = %#016x X23 = %#016x\n", regs.X22, regs.X23)
|
||||
fmt.Printf(" X24 = %#016x X25 = %#016x\n", regs.X24, regs.X25)
|
||||
fmt.Printf(" X26 = %#016x X27 = %#016x\n", regs.X26, regs.X27)
|
||||
fmt.Printf(" X28 = %#016x PSTATE = %#x\n", regs.X28, regs.PSTATE)
|
||||
}
|
||||
|
||||
func printVectorRegs(v *VectorRegs) {
|
||||
fmt.Println("\n Vector registers (V0-V31):")
|
||||
for i := 0; i < 32; i += 2 {
|
||||
fmt.Printf(" V%-2d = %016x%016x\n", i, binary.LittleEndian.Uint64(v.V[i][8:16]), binary.LittleEndian.Uint64(v.V[i][0:8]))
|
||||
fmt.Printf(" V%-2d = %016x%016x\n", i+1, binary.LittleEndian.Uint64(v.V[i+1][8:16]), binary.LittleEndian.Uint64(v.V[i+1][0:8]))
|
||||
}
|
||||
}
|
||||
|
||||
// SetReg modifies a register value in the debuggee.
|
||||
func (s *Session) SetReg(name string, value uint64) error {
|
||||
regs, err := s.GetRegs()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
switch name {
|
||||
case "x0":
|
||||
regs.X0 = value
|
||||
case "x1":
|
||||
regs.X1 = value
|
||||
case "x2":
|
||||
regs.X2 = value
|
||||
case "x3":
|
||||
regs.X3 = value
|
||||
case "x4":
|
||||
regs.X4 = value
|
||||
case "x5":
|
||||
regs.X5 = value
|
||||
case "x6":
|
||||
regs.X6 = value
|
||||
case "x7":
|
||||
regs.X7 = value
|
||||
case "x8":
|
||||
regs.X8 = value
|
||||
case "x9":
|
||||
regs.X9 = value
|
||||
case "x10":
|
||||
regs.X10 = value
|
||||
case "x11":
|
||||
regs.X11 = value
|
||||
case "x12":
|
||||
regs.X12 = value
|
||||
case "x13":
|
||||
regs.X13 = value
|
||||
case "x14":
|
||||
regs.X14 = value
|
||||
case "x15":
|
||||
regs.X15 = value
|
||||
case "x16":
|
||||
regs.X16 = value
|
||||
case "x17":
|
||||
regs.X17 = value
|
||||
case "x18":
|
||||
regs.X18 = value
|
||||
case "x19":
|
||||
regs.X19 = value
|
||||
case "x20":
|
||||
regs.X20 = value
|
||||
case "x21":
|
||||
regs.X21 = value
|
||||
case "x22":
|
||||
regs.X22 = value
|
||||
case "x23":
|
||||
regs.X23 = value
|
||||
case "x24":
|
||||
regs.X24 = value
|
||||
case "x25":
|
||||
regs.X25 = value
|
||||
case "x26":
|
||||
regs.X26 = value
|
||||
case "x27":
|
||||
regs.X27 = value
|
||||
case "x28":
|
||||
regs.X28 = value
|
||||
case "x29", "fp":
|
||||
regs.X29 = value
|
||||
case "x30", "lr":
|
||||
regs.X30 = value
|
||||
case "sp":
|
||||
regs.SP = value
|
||||
case "pc":
|
||||
regs.PC = value
|
||||
default:
|
||||
return fmt.Errorf("debug: unknown register %q", name)
|
||||
}
|
||||
return s.SetRegs(®s)
|
||||
}
|
||||
|
||||
// archReturnAddr reads the return address from LR (arm64 convention).
|
||||
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
|
||||
return regs.X30, nil
|
||||
}
|
||||
|
||||
// archSPLabel returns the SP register name for display.
|
||||
func archSPLabel() string { return "SP" }
|
||||
@@ -0,0 +1,121 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && riscv64
|
||||
|
||||
package debug
|
||||
|
||||
import "fmt"
|
||||
|
||||
func printRegs(regs *Regs, codeBase, funcOff uint64) {
|
||||
fmt.Printf(" PC = %#016x (func+%#x)\n", regs.PC, regs.PC-codeBase-funcOff)
|
||||
fmt.Printf(" SP = %#016x FP = %#016x\n", regs.Sp, regs.S0)
|
||||
fmt.Printf(" RA = %#016x\n", regs.Ra)
|
||||
fmt.Printf(" A0 = %#016x A1 = %#016x\n", regs.A0, regs.A1)
|
||||
fmt.Printf(" A2 = %#016x A3 = %#016x\n", regs.A2, regs.A3)
|
||||
fmt.Printf(" A4 = %#016x A5 = %#016x\n", regs.A4, regs.A5)
|
||||
fmt.Printf(" A6 = %#016x A7 = %#016x\n", regs.A6, regs.A7)
|
||||
fmt.Printf(" T0 = %#016x T1 = %#016x\n", regs.T0, regs.T1)
|
||||
fmt.Printf(" T2 = %#016x T3 = %#016x\n", regs.T2, regs.T3)
|
||||
fmt.Printf(" T4 = %#016x T5 = %#016x\n", regs.T4, regs.T5)
|
||||
fmt.Printf(" T6 = %#016x\n", regs.T6)
|
||||
fmt.Printf(" S1 = %#016x S2 = %#016x\n", regs.S1, regs.S2)
|
||||
fmt.Printf(" S3 = %#016x S4 = %#016x\n", regs.S3, regs.S4)
|
||||
fmt.Printf(" S5 = %#016x S6 = %#016x\n", regs.S5, regs.S6)
|
||||
fmt.Printf(" S7 = %#016x S8 = %#016x\n", regs.S7, regs.S8)
|
||||
fmt.Printf(" S9 = %#016x S10 = %#016x\n", regs.S9, regs.S10)
|
||||
fmt.Printf(" S11 = %#016x\n", regs.S11)
|
||||
}
|
||||
|
||||
func printVectorRegs(v *VectorRegs) {
|
||||
fmt.Println("\n FP registers (F0-F31):")
|
||||
for i := 0; i < 32; i += 2 {
|
||||
fmt.Printf(" F%-2d = %#018x F%-2d = %#018x\n", i, v.F[i], i+1, v.F[i+1])
|
||||
}
|
||||
fmt.Printf(" FCSR = %#x\n", v.FCSR)
|
||||
}
|
||||
|
||||
// SetReg modifies a register value in the debuggee.
|
||||
func (s *Session) SetReg(name string, value uint64) error {
|
||||
regs, err := s.GetRegs()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
switch name {
|
||||
case "pc":
|
||||
regs.PC = value
|
||||
case "ra", "x1":
|
||||
regs.Ra = value
|
||||
case "sp", "x2":
|
||||
regs.Sp = value
|
||||
case "gp", "x3":
|
||||
regs.Gp = value
|
||||
case "tp", "x4":
|
||||
regs.Tp = value
|
||||
case "t0", "x5":
|
||||
regs.T0 = value
|
||||
case "t1", "x6":
|
||||
regs.T1 = value
|
||||
case "t2", "x7":
|
||||
regs.T2 = value
|
||||
case "s0", "fp", "x8":
|
||||
regs.S0 = value
|
||||
case "s1", "x9":
|
||||
regs.S1 = value
|
||||
case "a0", "x10":
|
||||
regs.A0 = value
|
||||
case "a1", "x11":
|
||||
regs.A1 = value
|
||||
case "a2", "x12":
|
||||
regs.A2 = value
|
||||
case "a3", "x13":
|
||||
regs.A3 = value
|
||||
case "a4", "x14":
|
||||
regs.A4 = value
|
||||
case "a5", "x15":
|
||||
regs.A5 = value
|
||||
case "a6", "x16":
|
||||
regs.A6 = value
|
||||
case "a7", "x17":
|
||||
regs.A7 = value
|
||||
case "s2", "x18":
|
||||
regs.S2 = value
|
||||
case "s3", "x19":
|
||||
regs.S3 = value
|
||||
case "s4", "x20":
|
||||
regs.S4 = value
|
||||
case "s5", "x21":
|
||||
regs.S5 = value
|
||||
case "s6", "x22":
|
||||
regs.S6 = value
|
||||
case "s7", "x23":
|
||||
regs.S7 = value
|
||||
case "s8", "x24":
|
||||
regs.S8 = value
|
||||
case "s9", "x25":
|
||||
regs.S9 = value
|
||||
case "s10", "x26":
|
||||
regs.S10 = value
|
||||
case "s11", "x27":
|
||||
regs.S11 = value
|
||||
case "t3", "x28":
|
||||
regs.T3 = value
|
||||
case "t4", "x29":
|
||||
regs.T4 = value
|
||||
case "t5", "x30":
|
||||
regs.T5 = value
|
||||
case "t6", "x31":
|
||||
regs.T6 = value
|
||||
default:
|
||||
return fmt.Errorf("debug: unknown register %q", name)
|
||||
}
|
||||
return s.SetRegs(®s)
|
||||
}
|
||||
|
||||
// archReturnAddr reads the return address from RA (riscv64 convention).
|
||||
func archReturnAddr(s *Session, regs *Regs) (uint64, error) {
|
||||
return regs.Ra, nil
|
||||
}
|
||||
|
||||
// archSPLabel returns the SP register name for display.
|
||||
func archSPLabel() string { return "SP" }
|
||||
@@ -17,8 +17,8 @@ import (
|
||||
"time"
|
||||
"unsafe"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// Integration tests beyond the basic entry breakpoint: hardware watchpoints,
|
||||
|
||||
@@ -0,0 +1,288 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && (amd64 || arm64 || riscv64)
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// Session is a ptrace debugging session controlling one debuggee process.
|
||||
// The FreeBSD implementation sits behind the same surface as the Linux one:
|
||||
// PT_TRACE_ME from the debuggee, PT_CONTINUE/PT_STEP from the tracer, and
|
||||
// tracee memory through PT_IO (FreeBSD has no /proc/pid/mem to fall back
|
||||
// on, so PT_IO is the only supported route).
|
||||
type Session struct {
|
||||
pid int
|
||||
cmd *exec.Cmd
|
||||
stopped bool
|
||||
exited bool
|
||||
codeBase uint64 // base address of the JIT code in the debuggee
|
||||
tmpDir string // scratch directory of the session, removed on Kill
|
||||
wpSlots [16]bool // hardware watchpoint slots in use (DR0-DR3, arm64 dbw 0-15)
|
||||
// lastSignal holds the signal of the most recent stop when that stop
|
||||
// was a genuine signal-delivery-stop the caller must see (a fault such
|
||||
// as SIGSEGV, SIGBUS, SIGFPE or SIGILL); 0 for breakpoint traps,
|
||||
// single-steps, SIGSTOP and suppressed runtime signals.
|
||||
lastSignal syscall.Signal
|
||||
}
|
||||
|
||||
// Launch starts the debuggee subprocess (gasm debug --target ...) and
|
||||
// attaches to it via ptrace.
|
||||
func Launch(gasmBin, asmPath, funcName string, args []byte) (*Session, error) {
|
||||
sess, _, err := LaunchWithBuffers(gasmBin, asmPath, funcName, args, "")
|
||||
return sess, err
|
||||
}
|
||||
|
||||
// LaunchWithBuffers is like Launch but also allocates buffers in the debuggee.
|
||||
//
|
||||
// It pins the calling goroutine to its OS thread and leaves it pinned: the
|
||||
// debuggee's PT_TRACE_ME binds the tracer relation to the forking thread,
|
||||
// and every ptrace request on the session must come from that same thread.
|
||||
// All Session methods must therefore be called from the goroutine that
|
||||
// launched the session (the REPL and coverage loops do exactly that).
|
||||
func LaunchWithBuffers(gasmBin, asmPath, funcName string, args []byte, bufSpec string) (*Session, []uint64, error) {
|
||||
runtime.LockOSThread() // ptrace requests must stay on the forking thread
|
||||
self, err := os.Executable()
|
||||
if err != nil {
|
||||
return nil, nil, fmt.Errorf("debug: cannot find gasm binary: %w", err)
|
||||
}
|
||||
if gasmBin != "" {
|
||||
self = gasmBin
|
||||
}
|
||||
|
||||
tmpDir, err := os.MkdirTemp("", "gasm-debug-*")
|
||||
if err != nil {
|
||||
return nil, nil, fmt.Errorf("debug: tempdir: %w", err)
|
||||
}
|
||||
argsFile := filepath.Join(tmpDir, "args.bin")
|
||||
if err := os.WriteFile(argsFile, args, 0o644); err != nil {
|
||||
os.RemoveAll(tmpDir)
|
||||
return nil, nil, fmt.Errorf("debug: write args: %w", err)
|
||||
}
|
||||
|
||||
if bufSpec != "" {
|
||||
if err := os.WriteFile(filepath.Join(tmpDir, "bufspec"), []byte(bufSpec), 0o644); err != nil {
|
||||
os.RemoveAll(tmpDir)
|
||||
return nil, nil, fmt.Errorf("debug: write bufspec: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
cmd := exec.Command(self, "debug", "--func", funcName, "--args", argsFile, asmPath)
|
||||
cmd.Env = append(os.Environ(), "GASM_DEBUG_TARGET=1", "GASM_DEBUG_TMP="+tmpDir)
|
||||
cmd.Stdout = nil
|
||||
cmd.Stderr = os.Stderr
|
||||
cmd.SysProcAttr = &syscall.SysProcAttr{}
|
||||
|
||||
if err := cmd.Start(); err != nil {
|
||||
os.RemoveAll(tmpDir)
|
||||
return nil, nil, fmt.Errorf("debug: start debuggee: %w", err)
|
||||
}
|
||||
|
||||
s := &Session{pid: cmd.Process.Pid, cmd: cmd, tmpDir: tmpDir}
|
||||
|
||||
readyFile := filepath.Join(tmpDir, "ready")
|
||||
for range 500 {
|
||||
if _, err := os.Stat(readyFile); err == nil {
|
||||
break
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
|
||||
// The debuggee parks itself with SIGSTOP once the JIT code is mapped.
|
||||
// A Go tracee also reports SIGURG preemption as signal-delivery-stops,
|
||||
// so the wait loops until a stop the debugger cares about instead of
|
||||
// assuming the first event is the SIGSTOP.
|
||||
if _, err := s.waitStopped(); err != nil {
|
||||
cmd.Process.Kill()
|
||||
os.RemoveAll(tmpDir)
|
||||
return nil, nil, fmt.Errorf("debug: wait for debuggee: %w", err)
|
||||
}
|
||||
s.stopped = true
|
||||
|
||||
// The debuggee reports its JIT mapping in the codebase file; that is
|
||||
// the supported path on FreeBSD, where no /proc/pid/maps exists to
|
||||
// scan for the RWX region as a fallback.
|
||||
if data, err := os.ReadFile(filepath.Join(tmpDir, "codebase")); err == nil {
|
||||
fmt.Sscanf(string(data), "%d", &s.codeBase)
|
||||
}
|
||||
|
||||
var bufAddrs []uint64
|
||||
if bufSpec != "" {
|
||||
addrFile := filepath.Join(tmpDir, "bufaddrs")
|
||||
if data, err := os.ReadFile(addrFile); err == nil {
|
||||
for line := range strings.SplitSeq(strings.TrimSpace(string(data)), "\n") {
|
||||
var addr uint64
|
||||
if _, err := fmt.Sscanf(line, "%d", &addr); err == nil {
|
||||
bufAddrs = append(bufAddrs, addr)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return s, bufAddrs, nil
|
||||
}
|
||||
|
||||
// waitStopped consumes ptrace-stop events until one the debugger cares
|
||||
// about arrives: SIGTRAP (a breakpoint or a completed single-step), the
|
||||
// debuggee's own SIGSTOP, or a genuine signal-delivery-stop. A Go tracee's
|
||||
// runtime raises SIGURG for asynchronous preemption, and every signal on a
|
||||
// traced thread surfaces as a signal-delivery-stop, so SIGURG is suppressed
|
||||
// and the tracee resumed without it. Every other signal (SIGSEGV, SIGBUS,
|
||||
// SIGFPE, SIGILL, ...) is returned to the caller: resuming with signal 0
|
||||
// would restart the faulting instruction and fault forever, so a faulting
|
||||
// kernel must surface as a stop the caller reports.
|
||||
func (s *Session) waitStopped() (syscall.Signal, error) {
|
||||
for {
|
||||
var ws syscall.WaitStatus
|
||||
if _, err := syscall.Wait4(s.pid, &ws, syscall.WUNTRACED, nil); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if ws.Exited() {
|
||||
s.exited = true
|
||||
return 0, fmt.Errorf("debuggee exited with status %d", ws.ExitStatus())
|
||||
}
|
||||
if ws.Signaled() {
|
||||
s.exited = true
|
||||
return 0, fmt.Errorf("debuggee killed by signal %v", ws.Signal())
|
||||
}
|
||||
switch sig := ws.StopSignal(); sig {
|
||||
case syscall.SIGTRAP, syscall.SIGSTOP:
|
||||
s.stopped = true
|
||||
s.lastSignal = 0
|
||||
return sig, nil
|
||||
case syscall.SIGURG:
|
||||
// Go runtime asynchronous preemption: resume the tracee
|
||||
// without delivering the signal.
|
||||
s.lastSignal = 0
|
||||
if err := unix.PtraceCont(s.pid, 0); err != nil {
|
||||
return 0, fmt.Errorf("debug: PT_CONTINUE: %w", err)
|
||||
}
|
||||
default:
|
||||
// A genuine signal-delivery-stop. Report it; the caller
|
||||
// decides how to proceed.
|
||||
s.stopped = true
|
||||
s.lastSignal = sig
|
||||
return sig, nil
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// LastSignal returns the signal of the most recent stop when that stop was
|
||||
// a genuine signal-delivery-stop (a fault such as SIGSEGV, SIGFPE, SIGILL
|
||||
// or SIGBUS), and 0 for breakpoint traps, single-steps, SIGSTOP and
|
||||
// suppressed runtime signals.
|
||||
func (s *Session) LastSignal() syscall.Signal { return s.lastSignal }
|
||||
|
||||
// Peek reads a word (8 bytes) from the debuggee's memory at addr, through
|
||||
// PT_IO with PIOD_READ_D.
|
||||
func (s *Session) Peek(addr uint64) (uint64, error) {
|
||||
var buf [8]byte
|
||||
if _, err := unix.PtraceIO(unix.PIOD_READ_D, s.pid, uintptr(addr), buf[:], len(buf)); err != nil {
|
||||
return 0, fmt.Errorf("debug: read mem %#x: %w", addr, err)
|
||||
}
|
||||
return uint64(buf[0]) | uint64(buf[1])<<8 | uint64(buf[2])<<16 | uint64(buf[3])<<24 |
|
||||
uint64(buf[4])<<32 | uint64(buf[5])<<40 | uint64(buf[6])<<48 | uint64(buf[7])<<56, nil
|
||||
}
|
||||
|
||||
// Poke writes a word (8 bytes) to the debuggee's memory at addr, through
|
||||
// PT_IO with PIOD_WRITE_D.
|
||||
func (s *Session) Poke(addr, val uint64) error {
|
||||
buf := []byte{byte(val), byte(val >> 8), byte(val >> 16), byte(val >> 24),
|
||||
byte(val >> 32), byte(val >> 40), byte(val >> 48), byte(val >> 56)}
|
||||
if _, err := unix.PtraceIO(unix.PIOD_WRITE_D, s.pid, uintptr(addr), buf, len(buf)); err != nil {
|
||||
return fmt.Errorf("debug: write mem %#x: %w", addr, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// ReadMemory reads len bytes from the debuggee's memory at addr in one
|
||||
// PT_IO request, the shape the request is built for.
|
||||
func (s *Session) ReadMemory(addr uint64, length int) ([]byte, error) {
|
||||
out := make([]byte, length)
|
||||
n, err := unix.PtraceIO(unix.PIOD_READ_D, s.pid, uintptr(addr), out, length)
|
||||
return out[:n], err
|
||||
}
|
||||
|
||||
// WriteMemory writes bytes to the debuggee's memory at addr in one PT_IO
|
||||
// request.
|
||||
func (s *Session) WriteMemory(addr uint64, data []byte) error {
|
||||
_, err := unix.PtraceIO(unix.PIOD_WRITE_D, s.pid, uintptr(addr), data, len(data))
|
||||
return err
|
||||
}
|
||||
|
||||
// Step executes a single instruction in the debuggee.
|
||||
func (s *Session) Step() error {
|
||||
if s.exited {
|
||||
return fmt.Errorf("debug: debuggee has exited")
|
||||
}
|
||||
if err := unix.PtraceSingleStep(s.pid); err != nil {
|
||||
return fmt.Errorf("debug: PT_STEP: %w", err)
|
||||
}
|
||||
_, err := s.waitStopped()
|
||||
return err
|
||||
}
|
||||
|
||||
// Continue resumes execution until the next breakpoint or exit.
|
||||
func (s *Session) Continue() error {
|
||||
if s.exited {
|
||||
return fmt.Errorf("debug: debuggee has exited")
|
||||
}
|
||||
if err := unix.PtraceCont(s.pid, 0); err != nil {
|
||||
return fmt.Errorf("debug: PT_CONTINUE: %w", err)
|
||||
}
|
||||
_, err := s.waitStopped()
|
||||
return err
|
||||
}
|
||||
|
||||
// Exited returns true if the debuggee has terminated.
|
||||
func (s *Session) Exited() bool { return s.exited }
|
||||
|
||||
// Pid returns the debuggee's process ID.
|
||||
func (s *Session) Pid() int { return s.pid }
|
||||
|
||||
// CodeBase returns the base address of the JIT code in the debuggee.
|
||||
func (s *Session) CodeBase() uint64 { return s.codeBase }
|
||||
|
||||
// Kill terminates the debuggee and removes the session's scratch
|
||||
// directory, so a successful session leaves no gasm-debug-* debris behind.
|
||||
func (s *Session) Kill() {
|
||||
if !s.exited {
|
||||
syscall.Kill(s.pid, syscall.SIGKILL)
|
||||
syscall.Wait4(s.pid, nil, 0, nil)
|
||||
s.exited = true
|
||||
}
|
||||
if s.cmd != nil && s.cmd.Process != nil {
|
||||
s.cmd.Wait()
|
||||
}
|
||||
if s.tmpDir != "" {
|
||||
os.RemoveAll(s.tmpDir)
|
||||
s.tmpDir = ""
|
||||
}
|
||||
}
|
||||
|
||||
// execRange is one executable mapping of the debuggee.
|
||||
type execRange struct {
|
||||
lo, hi uint64
|
||||
}
|
||||
|
||||
// execRanges is a stub on FreeBSD: there is no /proc/pid/maps to parse,
|
||||
// and procfs(5) is not guaranteed to be mounted. The callers degrade
|
||||
// gracefully: archReturnAddr falls back to the raw stack convention and
|
||||
// the mapping scan is skipped.
|
||||
func execRanges(pid int) []execRange { return nil }
|
||||
|
||||
// findRWXMapping is a stub on FreeBSD for the same reason: the codebase
|
||||
// handshake file is the supported way the JIT region is located.
|
||||
func findRWXMapping(pid int) uint64 { return 0 }
|
||||
@@ -0,0 +1,160 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
"unsafe"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// GetRegs reads the general-purpose registers of the stopped debuggee and
|
||||
// converts the FreeBSD struct reg into the portable layout.
|
||||
func (s *Session) GetRegs() (Regs, error) {
|
||||
var ur unix.Reg
|
||||
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
|
||||
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
|
||||
}
|
||||
return Regs{
|
||||
R15: uint64(ur.R15),
|
||||
R14: uint64(ur.R14),
|
||||
R13: uint64(ur.R13),
|
||||
R12: uint64(ur.R12),
|
||||
R11: uint64(ur.R11),
|
||||
R10: uint64(ur.R10),
|
||||
R9: uint64(ur.R9),
|
||||
R8: uint64(ur.R8),
|
||||
RDI: uint64(ur.Rdi),
|
||||
RSI: uint64(ur.Rsi),
|
||||
RBP: uint64(ur.Rbp),
|
||||
RBX: uint64(ur.Rbx),
|
||||
RDX: uint64(ur.Rdx),
|
||||
RCX: uint64(ur.Rcx),
|
||||
RAX: uint64(ur.Rax),
|
||||
RIP: uint64(ur.Rip),
|
||||
CS: uint64(ur.Cs),
|
||||
RFLAGS: uint64(ur.Rflags),
|
||||
RSP: uint64(ur.Rsp),
|
||||
SS: uint64(ur.Ss),
|
||||
FS: uint64(ur.Fs),
|
||||
GS: uint64(ur.Gs),
|
||||
DS: uint64(ur.Ds),
|
||||
ES: uint64(ur.Es),
|
||||
}, nil
|
||||
}
|
||||
|
||||
// SetRegs writes the general-purpose registers of the stopped debuggee.
|
||||
func (s *Session) SetRegs(regs *Regs) error {
|
||||
// Read-modify-write keeps the fields FreeBSD owns (trapno, err) intact.
|
||||
var ur unix.Reg
|
||||
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
|
||||
return fmt.Errorf("debug: PT_GETREGS: %w", err)
|
||||
}
|
||||
ur.R15 = int64(regs.R15)
|
||||
ur.R14 = int64(regs.R14)
|
||||
ur.R13 = int64(regs.R13)
|
||||
ur.R12 = int64(regs.R12)
|
||||
ur.R11 = int64(regs.R11)
|
||||
ur.R10 = int64(regs.R10)
|
||||
ur.R9 = int64(regs.R9)
|
||||
ur.R8 = int64(regs.R8)
|
||||
ur.Rdi = int64(regs.RDI)
|
||||
ur.Rsi = int64(regs.RSI)
|
||||
ur.Rbp = int64(regs.RBP)
|
||||
ur.Rbx = int64(regs.RBX)
|
||||
ur.Rdx = int64(regs.RDX)
|
||||
ur.Rcx = int64(regs.RCX)
|
||||
ur.Rax = int64(regs.RAX)
|
||||
ur.Rip = int64(regs.RIP)
|
||||
ur.Cs = int64(regs.CS)
|
||||
ur.Rflags = int64(regs.RFLAGS)
|
||||
ur.Rsp = int64(regs.RSP)
|
||||
ur.Ss = int64(regs.SS)
|
||||
return unix.PtraceSetRegs(s.pid, &ur)
|
||||
}
|
||||
|
||||
// FPRegs holds the x87 FPU and SSE (XMM) register state, the FXSAVE image
|
||||
// the FreeBSD struct fpreg mirrors: XMM0-15 at the same offsets.
|
||||
type FPRegs struct {
|
||||
XMM [16][16]byte // XMM0-15
|
||||
}
|
||||
|
||||
// GetFPRegs retrieves the FPU/SSE register state via PT_GETFPREGS. The
|
||||
// FreeBSD struct fpreg mirrors the FXSAVE image: the x87 environment and
|
||||
// stack in Env/Acc, XMM0-15 in Xacc.
|
||||
func (s *Session) GetFPRegs() (FPRegs, error) {
|
||||
var fp FPRegs
|
||||
var fr unix.FpReg
|
||||
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
|
||||
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
|
||||
}
|
||||
for i := range 16 {
|
||||
copy(fp.XMM[i][:], fr.Xacc[i][:])
|
||||
}
|
||||
return fp, nil
|
||||
}
|
||||
|
||||
// VectorRegs holds the YMM register state.
|
||||
type VectorRegs struct {
|
||||
YMM [16][32]byte // YMM0-15 (full 256-bit values)
|
||||
}
|
||||
|
||||
// The XSAVE area the PT_GETXSTATE request returns follows the architectural
|
||||
// layout (Intel SDM vol 1, "XSAVE"): the 512-byte legacy FXSAVE image (x87
|
||||
// state in 0-159, XMM0-15 in 160-511), then the 64-byte xsave header whose
|
||||
// first 8 bytes are xstate_bv, then one component per set feature bit, each
|
||||
// 64-byte aligned. The YMM high halves are the first extended component,
|
||||
// at offset 576; XFEATURE_STATE_BIT_AVX is bit 2 of xstate_bv.
|
||||
const (
|
||||
xsaveXMMOffset = 160
|
||||
xsaveHeaderOffset = 512
|
||||
xsaveBVOffset = xsaveHeaderOffset
|
||||
ymmOffset = xsaveHeaderOffset + 64 // 576
|
||||
ymmSize = 256 // 16 registers, 16 bytes each
|
||||
xfeatureMaskYMM = 1 << 2
|
||||
xstateMaxBuffer = 4096 // PT_GETXSTATE_INFO bounds the size far below this
|
||||
)
|
||||
|
||||
// GetVectorRegs retrieves the YMM registers via PT_GETXSTATE. The low
|
||||
// (XMM) halves always come from the legacy image; the high halves are
|
||||
// copied only when xstate_bv reports the AVX state, and read as zero
|
||||
// otherwise. When the request fails the FP image still provides correct
|
||||
// XMM halves, so that is the fallback.
|
||||
func (s *Session) GetVectorRegs() (VectorRegs, error) {
|
||||
var v VectorRegs
|
||||
buf := make([]byte, xstateMaxBuffer)
|
||||
n, _, errno := unix.Syscall6(
|
||||
unix.SYS_PTRACE,
|
||||
uintptr(unix.PT_GETXSTATE),
|
||||
uintptr(s.pid),
|
||||
0,
|
||||
uintptr(unsafe.Pointer(&buf[0])),
|
||||
0, 0,
|
||||
)
|
||||
if errno != 0 {
|
||||
fp, err := s.GetFPRegs()
|
||||
if err != nil {
|
||||
return v, err
|
||||
}
|
||||
for i := range 16 {
|
||||
copy(v.YMM[i][:16], fp.XMM[i][:])
|
||||
}
|
||||
return v, nil
|
||||
}
|
||||
for i := range 16 {
|
||||
copy(v.YMM[i][:16], buf[xsaveXMMOffset+16*i:xsaveXMMOffset+16*i+16])
|
||||
}
|
||||
if int(n) >= ymmOffset+ymmSize {
|
||||
if binary.LittleEndian.Uint64(buf[xsaveBVOffset:xsaveBVOffset+8])&xfeatureMaskYMM != 0 {
|
||||
for i := range 16 {
|
||||
copy(v.YMM[i][16:], buf[ymmOffset+16*i:ymmOffset+16*i+16])
|
||||
}
|
||||
}
|
||||
}
|
||||
return v, nil
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && arm64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// GetRegs reads the general-purpose registers of the stopped debuggee and
|
||||
// converts the FreeBSD struct reg (x[30], lr, sp, elr, spsr) into the
|
||||
// portable layout.
|
||||
func (s *Session) GetRegs() (Regs, error) {
|
||||
var ur unix.Reg
|
||||
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
|
||||
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
|
||||
}
|
||||
return Regs{
|
||||
X0: ur.X[0],
|
||||
X1: ur.X[1],
|
||||
X2: ur.X[2],
|
||||
X3: ur.X[3],
|
||||
X4: ur.X[4],
|
||||
X5: ur.X[5],
|
||||
X6: ur.X[6],
|
||||
X7: ur.X[7],
|
||||
X8: ur.X[8],
|
||||
X9: ur.X[9],
|
||||
X10: ur.X[10],
|
||||
X11: ur.X[11],
|
||||
X12: ur.X[12],
|
||||
X13: ur.X[13],
|
||||
X14: ur.X[14],
|
||||
X15: ur.X[15],
|
||||
X16: ur.X[16],
|
||||
X17: ur.X[17],
|
||||
X18: ur.X[18],
|
||||
X19: ur.X[19],
|
||||
X20: ur.X[20],
|
||||
X21: ur.X[21],
|
||||
X22: ur.X[22],
|
||||
X23: ur.X[23],
|
||||
X24: ur.X[24],
|
||||
X25: ur.X[25],
|
||||
X26: ur.X[26],
|
||||
X27: ur.X[27],
|
||||
X28: ur.X[28],
|
||||
X29: ur.X[29],
|
||||
X30: ur.Lr,
|
||||
SP: ur.Sp,
|
||||
PC: ur.Elr,
|
||||
PSTATE: uint64(ur.Spsr),
|
||||
}, nil
|
||||
}
|
||||
|
||||
// SetRegs writes the general-purpose registers of the stopped debuggee.
|
||||
func (s *Session) SetRegs(regs *Regs) error {
|
||||
var ur unix.Reg
|
||||
ur.X = [30]uint64{
|
||||
regs.X0, regs.X1, regs.X2, regs.X3, regs.X4, regs.X5, regs.X6,
|
||||
regs.X7, regs.X8, regs.X9, regs.X10, regs.X11, regs.X12, regs.X13,
|
||||
regs.X14, regs.X15, regs.X16, regs.X17, regs.X18, regs.X19, regs.X20,
|
||||
regs.X21, regs.X22, regs.X23, regs.X24, regs.X25, regs.X26, regs.X27,
|
||||
regs.X28, regs.X29,
|
||||
}
|
||||
ur.Lr = regs.X30
|
||||
ur.Sp = regs.SP
|
||||
ur.Elr = regs.PC
|
||||
ur.Spsr = uint32(regs.PSTATE)
|
||||
return unix.PtraceSetRegs(s.pid, &ur)
|
||||
}
|
||||
|
||||
// FPRegs holds the arm64 FP/NEON register state: the 32 128-bit V
|
||||
// registers, then FPSR and FPCR (the user_fpsimd shape).
|
||||
type FPRegs struct {
|
||||
V [32][16]byte // V0-V31 (128-bit NEON/FP registers)
|
||||
FPSR uint32
|
||||
FPCR uint32
|
||||
}
|
||||
|
||||
// GetFPRegs retrieves the FP/NEON register state via PT_GETFPREGS. The
|
||||
// FreeBSD struct fpreg holds the 32 128-bit V registers followed by FPSR
|
||||
// and FPCR, the user_fpsimd shape.
|
||||
func (s *Session) GetFPRegs() (FPRegs, error) {
|
||||
var fp FPRegs
|
||||
var fr unix.FpReg
|
||||
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
|
||||
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
|
||||
}
|
||||
for i := range 32 {
|
||||
copy(fp.V[i][:], fr.Q[i][:])
|
||||
}
|
||||
return fp, nil
|
||||
}
|
||||
|
||||
// VectorRegs holds the full SIMD register state.
|
||||
type VectorRegs struct {
|
||||
V [32][16]byte // V0-V31 (128-bit)
|
||||
}
|
||||
|
||||
// GetVectorRegs retrieves the SIMD registers.
|
||||
func (s *Session) GetVectorRegs() (VectorRegs, error) {
|
||||
var v VectorRegs
|
||||
fp, err := s.GetFPRegs()
|
||||
if err != nil {
|
||||
return v, err
|
||||
}
|
||||
copy(v.V[:][:], fp.V[:][:])
|
||||
return v, nil
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && riscv64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// GetRegs reads the general-purpose registers of the stopped debuggee and
|
||||
// converts the FreeBSD struct reg into the portable layout. Sstatus rides
|
||||
// the kernel's struct but the portable surface carries the GPRs and PC.
|
||||
func (s *Session) GetRegs() (Regs, error) {
|
||||
var ur unix.Reg
|
||||
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
|
||||
return Regs{}, fmt.Errorf("debug: PT_GETREGS: %w", err)
|
||||
}
|
||||
return Regs{
|
||||
PC: ur.Sepc,
|
||||
Ra: ur.Ra,
|
||||
Sp: ur.Sp,
|
||||
Gp: ur.Gp,
|
||||
Tp: ur.Tp,
|
||||
T0: ur.T[0],
|
||||
T1: ur.T[1],
|
||||
T2: ur.T[2],
|
||||
S0: ur.S[0],
|
||||
S1: ur.S[1],
|
||||
A0: ur.A[0],
|
||||
A1: ur.A[1],
|
||||
A2: ur.A[2],
|
||||
A3: ur.A[3],
|
||||
A4: ur.A[4],
|
||||
A5: ur.A[5],
|
||||
A6: ur.A[6],
|
||||
A7: ur.A[7],
|
||||
S2: ur.S[2],
|
||||
S3: ur.S[3],
|
||||
S4: ur.S[4],
|
||||
S5: ur.S[5],
|
||||
S6: ur.S[6],
|
||||
S7: ur.S[7],
|
||||
S8: ur.S[8],
|
||||
S9: ur.S[9],
|
||||
S10: ur.S[10],
|
||||
S11: ur.S[11],
|
||||
T3: ur.T[3],
|
||||
T4: ur.T[4],
|
||||
T5: ur.T[5],
|
||||
T6: ur.T[6],
|
||||
}, nil
|
||||
}
|
||||
|
||||
// SetRegs writes the general-purpose registers of the stopped debuggee.
|
||||
// Read-modify-write keeps sstatus, which the kernel owns, intact.
|
||||
func (s *Session) SetRegs(regs *Regs) error {
|
||||
var ur unix.Reg
|
||||
if err := unix.PtraceGetRegs(s.pid, &ur); err != nil {
|
||||
return fmt.Errorf("debug: PT_GETREGS: %w", err)
|
||||
}
|
||||
ur.Sepc = regs.PC
|
||||
ur.Ra = regs.Ra
|
||||
ur.Sp = regs.Sp
|
||||
ur.Gp = regs.Gp
|
||||
ur.Tp = regs.Tp
|
||||
ur.T = [7]uint64{regs.T0, regs.T1, regs.T2, regs.T3, regs.T4, regs.T5, regs.T6}
|
||||
ur.S = [12]uint64{regs.S0, regs.S1, regs.S2, regs.S3, regs.S4, regs.S5,
|
||||
regs.S6, regs.S7, regs.S8, regs.S9, regs.S10, regs.S11}
|
||||
ur.A = [8]uint64{regs.A0, regs.A1, regs.A2, regs.A3, regs.A4, regs.A5, regs.A6, regs.A7}
|
||||
return unix.PtraceSetRegs(s.pid, &ur)
|
||||
}
|
||||
|
||||
// FPRegs holds the RISC-V FP register state (32 64-bit FP registers plus
|
||||
// fcsr).
|
||||
type FPRegs struct {
|
||||
F [32]uint64 // F0-F31 (64-bit FP registers)
|
||||
FCSR uint32
|
||||
}
|
||||
|
||||
// GetFPRegs retrieves the FP register state via PT_GETFPREGS. The FreeBSD
|
||||
// struct fpreg carries each 64-bit FP register in a 128-bit slot (fp_x is
|
||||
// the flat [64]-word area the x/sys type renders as [32][2]); the low word
|
||||
// holds the register, and FCSR rides the tail.
|
||||
func (s *Session) GetFPRegs() (FPRegs, error) {
|
||||
var fp FPRegs
|
||||
var fr unix.FpReg
|
||||
if err := unix.PtraceGetFpRegs(s.pid, &fr); err != nil {
|
||||
return fp, fmt.Errorf("debug: PT_GETFPREGS: %w", err)
|
||||
}
|
||||
for i := range 32 {
|
||||
fp.F[i] = fr.X[i][0]
|
||||
}
|
||||
fp.FCSR = uint32(fr.Fcsr)
|
||||
return fp, nil
|
||||
}
|
||||
|
||||
// VectorRegs holds the FP register state shown by the regs command
|
||||
// (riscv64 has 32 64-bit FP registers and fcsr).
|
||||
type VectorRegs struct {
|
||||
F [32]uint64
|
||||
FCSR uint32
|
||||
}
|
||||
|
||||
// GetVectorRegs retrieves the FP registers.
|
||||
func (s *Session) GetVectorRegs() (VectorRegs, error) {
|
||||
fp, err := s.GetFPRegs()
|
||||
if err != nil {
|
||||
return VectorRegs{}, err
|
||||
}
|
||||
return VectorRegs{F: fp.F, FCSR: fp.FCSR}, nil
|
||||
}
|
||||
@@ -0,0 +1,91 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// TestLaunchAndBreakpoint is the FreeBSD twin of the Linux integration
|
||||
// test: it drives the whole launch, breakpoint, trap and register-rewind
|
||||
// flow end to end. It needs a real FreeBSD kernel (ptrace does not work
|
||||
// under emulation), so it only runs where it can.
|
||||
func TestLaunchAndBreakpoint(t *testing.T) {
|
||||
if runtime.GOARCH != "amd64" {
|
||||
t.Skip("runs only on amd64 hosts")
|
||||
}
|
||||
// The tracer is the OS thread that forked the debuggee (PT_TRACE_ME
|
||||
// binds the relation to that thread); every ptrace request must come
|
||||
// from the same thread, so pin the test goroutine to one thread.
|
||||
runtime.LockOSThread()
|
||||
defer runtime.UnlockOSThread()
|
||||
|
||||
bin := filepath.Join(t.TempDir(), "gasm")
|
||||
out, err := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm").CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("build gasm: %v: %s", err, out)
|
||||
}
|
||||
|
||||
const kernelPath = "../testdata/verify/basic_amd64.s"
|
||||
k, err := verify.Load(kernelPath)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
t.Cleanup(k.Close)
|
||||
fl, err := k.Func("wideCopy")
|
||||
if err != nil {
|
||||
t.Fatalf("Func: %v", err)
|
||||
}
|
||||
|
||||
sess, err := Launch(bin, kernelPath, "wideCopy", make([]byte, fl.Args))
|
||||
if err != nil {
|
||||
t.Fatalf("Launch: %v", err)
|
||||
}
|
||||
t.Cleanup(sess.Kill)
|
||||
|
||||
bm := NewBreakpoints(sess)
|
||||
entry := sess.CodeBase() + uint64(fl.Offset)
|
||||
if _, err := bm.Set(entry, "entry"); err != nil {
|
||||
t.Fatalf("Set: %v", err)
|
||||
}
|
||||
|
||||
// The INT3 must be visible in the debuggee's memory.
|
||||
word, err := sess.Peek(entry)
|
||||
if err != nil {
|
||||
t.Fatalf("Peek: %v", err)
|
||||
}
|
||||
if b := word & 0xFF; b != 0xCC {
|
||||
t.Fatalf("int3 not patched: first byte %#02x at %#x", b, entry)
|
||||
}
|
||||
|
||||
// The debuggee raises a second SIGSTOP after the launch barrier (the
|
||||
// child's RunTarget marks its entry), so like the REPL and the cover
|
||||
// mode the test keeps resuming until the breakpoint trap arrives.
|
||||
for range 10 {
|
||||
if err := sess.Continue(); err != nil {
|
||||
t.Fatalf("Continue: %v", err)
|
||||
}
|
||||
if sess.Exited() {
|
||||
t.Fatal("debuggee exited instead of trapping on the breakpoint")
|
||||
}
|
||||
regs, err := sess.GetRegs()
|
||||
if err != nil {
|
||||
t.Fatalf("GetRegs: %v", err)
|
||||
}
|
||||
if bp := bm.HandleTrap(®s); bp != nil {
|
||||
if bp.Addr != entry {
|
||||
t.Fatalf("trap at %#x, want %#x", bp.Addr, entry)
|
||||
}
|
||||
return // trap on the entry breakpoint: the whole flow works
|
||||
}
|
||||
}
|
||||
t.Fatal("no breakpoint trap after 10 resumes")
|
||||
}
|
||||
@@ -14,7 +14,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// buildGasm produces the gasm binary the debugger spawns as its debuggee.
|
||||
@@ -24,7 +24,7 @@ func buildGasm(t *testing.T) string {
|
||||
return p
|
||||
}
|
||||
bin := filepath.Join(t.TempDir(), "gasm")
|
||||
cmd := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-devkit/cmd/gasm")
|
||||
cmd := exec.Command("go", "build", "-o", bin, "sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm")
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("build gasm: %v: %s", err, out)
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
// Regs holds the full general-purpose register set of a traced process
|
||||
// (the FreeBSD amd64 struct reg layout, sys/x86/include/reg.h). FreeBSD
|
||||
// reports segment selectors (FS/GS/ES/DS), not the bases the Linux ptrace
|
||||
// surface carries, and has no ORIG_RAX slot.
|
||||
type Regs struct {
|
||||
R15 uint64
|
||||
R14 uint64
|
||||
R13 uint64
|
||||
R12 uint64
|
||||
RBP uint64
|
||||
RBX uint64
|
||||
R11 uint64
|
||||
R10 uint64
|
||||
R9 uint64
|
||||
R8 uint64
|
||||
RAX uint64
|
||||
RCX uint64
|
||||
RDX uint64
|
||||
RSI uint64
|
||||
RDI uint64
|
||||
RIP uint64
|
||||
CS uint64
|
||||
RFLAGS uint64
|
||||
RSP uint64
|
||||
SS uint64
|
||||
FS uint64
|
||||
GS uint64
|
||||
DS uint64
|
||||
ES uint64
|
||||
}
|
||||
|
||||
// GetPC returns the program counter.
|
||||
func (r *Regs) GetPC() uint64 { return r.RIP }
|
||||
|
||||
// SetPC sets the program counter.
|
||||
func (r *Regs) SetPC(pc uint64) { r.RIP = pc }
|
||||
|
||||
// GetSP returns the stack pointer.
|
||||
func (r *Regs) GetSP() uint64 { return r.RSP }
|
||||
|
||||
// RegValue returns the value of the named register, or false if unknown.
|
||||
func (r *Regs) RegValue(name string) (uint64, bool) {
|
||||
switch name {
|
||||
case "rax", "eax", "ax", "al":
|
||||
return r.RAX, true
|
||||
case "rbx", "ebx", "bx", "bl":
|
||||
return r.RBX, true
|
||||
case "rcx", "ecx", "cx", "cl":
|
||||
return r.RCX, true
|
||||
case "rdx", "edx", "dx", "dl":
|
||||
return r.RDX, true
|
||||
case "rsi", "esi", "si":
|
||||
return r.RSI, true
|
||||
case "rdi", "edi", "di":
|
||||
return r.RDI, true
|
||||
case "rbp", "ebp", "bp":
|
||||
return r.RBP, true
|
||||
case "rsp", "esp", "sp":
|
||||
return r.RSP, true
|
||||
case "r8":
|
||||
return r.R8, true
|
||||
case "r9":
|
||||
return r.R9, true
|
||||
case "r10":
|
||||
return r.R10, true
|
||||
case "r11":
|
||||
return r.R11, true
|
||||
case "r12":
|
||||
return r.R12, true
|
||||
case "r13":
|
||||
return r.R13, true
|
||||
case "r14":
|
||||
return r.R14, true
|
||||
case "r15":
|
||||
return r.R15, true
|
||||
case "rip", "eip":
|
||||
return r.RIP, true
|
||||
default:
|
||||
return 0, false
|
||||
}
|
||||
}
|
||||
|
||||
// breakpointInsn is the software breakpoint instruction.
|
||||
var breakpointInsn = []byte{0xCC} // INT3
|
||||
|
||||
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
|
||||
// a trap. INT3 leaves the hardware PC on the following instruction (Intel
|
||||
// SDM vol 3, "Debug Exceptions") and the FreeBSD T_BPTFLT path delivers
|
||||
// that frame unmodified (sys/amd64/amd64/trap.c), so the trap address is
|
||||
// PC-1, the same correction the Linux side applies.
|
||||
const breakpointPCAdjust = 1
|
||||
@@ -0,0 +1,139 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && arm64
|
||||
|
||||
package debug
|
||||
|
||||
// Regs holds the full general-purpose register set of a traced process
|
||||
// (the FreeBSD arm64 struct reg layout, sys/arm64/include/reg.h: x[30], lr,
|
||||
// sp, elr, spsr).
|
||||
type Regs struct {
|
||||
X0 uint64
|
||||
X1 uint64
|
||||
X2 uint64
|
||||
X3 uint64
|
||||
X4 uint64
|
||||
X5 uint64
|
||||
X6 uint64
|
||||
X7 uint64
|
||||
X8 uint64
|
||||
X9 uint64
|
||||
X10 uint64
|
||||
X11 uint64
|
||||
X12 uint64
|
||||
X13 uint64
|
||||
X14 uint64
|
||||
X15 uint64
|
||||
X16 uint64
|
||||
X17 uint64
|
||||
X18 uint64
|
||||
X19 uint64
|
||||
X20 uint64
|
||||
X21 uint64
|
||||
X22 uint64
|
||||
X23 uint64
|
||||
X24 uint64
|
||||
X25 uint64
|
||||
X26 uint64
|
||||
X27 uint64
|
||||
X28 uint64
|
||||
X29 uint64 // FP (frame pointer)
|
||||
X30 uint64 // LR (link register)
|
||||
SP uint64
|
||||
PC uint64
|
||||
PSTATE uint64
|
||||
}
|
||||
|
||||
// GetPC returns the program counter.
|
||||
func (r *Regs) GetPC() uint64 { return r.PC }
|
||||
|
||||
// SetPC sets the program counter.
|
||||
func (r *Regs) SetPC(pc uint64) { r.PC = pc }
|
||||
|
||||
// GetSP returns the stack pointer.
|
||||
func (r *Regs) GetSP() uint64 { return r.SP }
|
||||
|
||||
// RegValue returns the value of the named register, or false if unknown.
|
||||
func (r *Regs) RegValue(name string) (uint64, bool) {
|
||||
switch name {
|
||||
case "x0":
|
||||
return r.X0, true
|
||||
case "x1":
|
||||
return r.X1, true
|
||||
case "x2":
|
||||
return r.X2, true
|
||||
case "x3":
|
||||
return r.X3, true
|
||||
case "x4":
|
||||
return r.X4, true
|
||||
case "x5":
|
||||
return r.X5, true
|
||||
case "x6":
|
||||
return r.X6, true
|
||||
case "x7":
|
||||
return r.X7, true
|
||||
case "x8":
|
||||
return r.X8, true
|
||||
case "x9":
|
||||
return r.X9, true
|
||||
case "x10":
|
||||
return r.X10, true
|
||||
case "x11":
|
||||
return r.X11, true
|
||||
case "x12":
|
||||
return r.X12, true
|
||||
case "x13":
|
||||
return r.X13, true
|
||||
case "x14":
|
||||
return r.X14, true
|
||||
case "x15":
|
||||
return r.X15, true
|
||||
case "x16":
|
||||
return r.X16, true
|
||||
case "x17":
|
||||
return r.X17, true
|
||||
case "x18":
|
||||
return r.X18, true
|
||||
case "x19":
|
||||
return r.X19, true
|
||||
case "x20":
|
||||
return r.X20, true
|
||||
case "x21":
|
||||
return r.X21, true
|
||||
case "x22":
|
||||
return r.X22, true
|
||||
case "x23":
|
||||
return r.X23, true
|
||||
case "x24":
|
||||
return r.X24, true
|
||||
case "x25":
|
||||
return r.X25, true
|
||||
case "x26":
|
||||
return r.X26, true
|
||||
case "x27":
|
||||
return r.X27, true
|
||||
case "x28":
|
||||
return r.X28, true
|
||||
case "x29", "fp":
|
||||
return r.X29, true
|
||||
case "x30", "lr":
|
||||
return r.X30, true
|
||||
case "sp":
|
||||
return r.SP, true
|
||||
case "pc":
|
||||
return r.PC, true
|
||||
default:
|
||||
return 0, false
|
||||
}
|
||||
}
|
||||
|
||||
// breakpointInsn is the software breakpoint instruction (BRK #0).
|
||||
var breakpointInsn = []byte{0x00, 0x00, 0x20, 0xD4} // BRK #0
|
||||
|
||||
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
|
||||
// a trap: 0. The BRK synchronous exception leaves ELR_EL0 on the BRK
|
||||
// itself (ARM DDI 0487), and the FreeBSD EXCP_BRKPT_EL0 handler delivers
|
||||
// the frame's elr unmodified (sys/arm64/arm64/trap.c), so the trap address
|
||||
// is the PC as reported.
|
||||
const breakpointPCAdjust = 0
|
||||
@@ -0,0 +1,134 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && riscv64
|
||||
|
||||
package debug
|
||||
|
||||
// Regs holds the full general-purpose register set of a traced process
|
||||
// (the FreeBSD riscv64 struct reg layout: ra, sp, gp, tp, t0-t6, s0-s11,
|
||||
// a0-a7, sepc, sstatus).
|
||||
type Regs struct {
|
||||
PC uint64 // sepc
|
||||
Ra uint64 // x1 (return address)
|
||||
Sp uint64 // x2
|
||||
Gp uint64 // x3
|
||||
Tp uint64 // x4
|
||||
T0 uint64 // x5
|
||||
T1 uint64 // x6
|
||||
T2 uint64 // x7
|
||||
S0 uint64 // x8 (frame pointer)
|
||||
S1 uint64 // x9
|
||||
A0 uint64 // x10
|
||||
A1 uint64 // x11
|
||||
A2 uint64 // x12
|
||||
A3 uint64 // x13
|
||||
A4 uint64 // x14
|
||||
A5 uint64 // x15
|
||||
A6 uint64 // x16
|
||||
A7 uint64 // x17
|
||||
S2 uint64 // x18
|
||||
S3 uint64 // x19
|
||||
S4 uint64 // x20
|
||||
S5 uint64 // x21
|
||||
S6 uint64 // x22
|
||||
S7 uint64 // x23
|
||||
S8 uint64 // x24
|
||||
S9 uint64 // x25
|
||||
S10 uint64 // x26
|
||||
S11 uint64 // x27
|
||||
T3 uint64 // x28
|
||||
T4 uint64 // x29
|
||||
T5 uint64 // x30
|
||||
T6 uint64 // x31
|
||||
}
|
||||
|
||||
// GetPC returns the program counter.
|
||||
func (r *Regs) GetPC() uint64 { return r.PC }
|
||||
|
||||
// SetPC sets the program counter.
|
||||
func (r *Regs) SetPC(pc uint64) { r.PC = pc }
|
||||
|
||||
// GetSP returns the stack pointer.
|
||||
func (r *Regs) GetSP() uint64 { return r.Sp }
|
||||
|
||||
// RegValue returns the value of the named register, or false if unknown.
|
||||
func (r *Regs) RegValue(name string) (uint64, bool) {
|
||||
switch name {
|
||||
case "pc":
|
||||
return r.PC, true
|
||||
case "ra", "x1":
|
||||
return r.Ra, true
|
||||
case "sp", "x2":
|
||||
return r.Sp, true
|
||||
case "gp", "x3":
|
||||
return r.Gp, true
|
||||
case "tp", "x4":
|
||||
return r.Tp, true
|
||||
case "t0", "x5":
|
||||
return r.T0, true
|
||||
case "t1", "x6":
|
||||
return r.T1, true
|
||||
case "t2", "x7":
|
||||
return r.T2, true
|
||||
case "s0", "fp", "x8":
|
||||
return r.S0, true
|
||||
case "s1", "x9":
|
||||
return r.S1, true
|
||||
case "a0", "x10":
|
||||
return r.A0, true
|
||||
case "a1", "x11":
|
||||
return r.A1, true
|
||||
case "a2", "x12":
|
||||
return r.A2, true
|
||||
case "a3", "x13":
|
||||
return r.A3, true
|
||||
case "a4", "x14":
|
||||
return r.A4, true
|
||||
case "a5", "x15":
|
||||
return r.A5, true
|
||||
case "a6", "x16":
|
||||
return r.A6, true
|
||||
case "a7", "x17":
|
||||
return r.A7, true
|
||||
case "s2", "x18":
|
||||
return r.S2, true
|
||||
case "s3", "x19":
|
||||
return r.S3, true
|
||||
case "s4", "x20":
|
||||
return r.S4, true
|
||||
case "s5", "x21":
|
||||
return r.S5, true
|
||||
case "s6", "x22":
|
||||
return r.S6, true
|
||||
case "s7", "x23":
|
||||
return r.S7, true
|
||||
case "s8", "x24":
|
||||
return r.S8, true
|
||||
case "s9", "x25":
|
||||
return r.S9, true
|
||||
case "s10", "x26":
|
||||
return r.S10, true
|
||||
case "s11", "x27":
|
||||
return r.S11, true
|
||||
case "t3", "x28":
|
||||
return r.T3, true
|
||||
case "t4", "x29":
|
||||
return r.T4, true
|
||||
case "t5", "x30":
|
||||
return r.T5, true
|
||||
case "t6", "x31":
|
||||
return r.T6, true
|
||||
default:
|
||||
return 0, false
|
||||
}
|
||||
}
|
||||
|
||||
// breakpointInsn is the software breakpoint instruction (EBREAK).
|
||||
var breakpointInsn = []byte{0x73, 0x00, 0x10, 0x00} // ebreak
|
||||
|
||||
// breakpointPCAdjust is how far PC is past the breakpoint instruction after
|
||||
// a trap: 0. The EBREAK synchronous exception leaves sepc on the ebreak
|
||||
// itself (RISC-V privileged architecture), so the trap address is the PC as
|
||||
// reported.
|
||||
const breakpointPCAdjust = 0
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build linux
|
||||
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
|
||||
|
||||
package debug
|
||||
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && (amd64 || arm64 || riscv64)
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// FreeBSD TRAP_* si_code values (sys/signal.h). A breakpoint (INT3, BRK,
|
||||
// EBREAK) arrives as TRAP_BRKPT on every supported architecture; TRAP_TRACE
|
||||
// is shared by the completed single-step and the hardware watchpoint hit,
|
||||
// so the watchpoint layer disambiguates from the debug registers.
|
||||
const (
|
||||
trapBRKPT = 1 // TRAP_BRKPT
|
||||
trapTRACE = 2 // TRAP_TRACE
|
||||
)
|
||||
|
||||
// StopReason describes why the debuggee stopped.
|
||||
type StopReason int
|
||||
|
||||
const (
|
||||
StopNone StopReason = iota
|
||||
StopBreakpoint // software breakpoint hit
|
||||
StopWatchpoint // hardware watchpoint triggered
|
||||
StopSingleStep // single-step completed
|
||||
StopSignal // stopped by a signal
|
||||
StopExited // process exited
|
||||
)
|
||||
|
||||
// StopInfo returns the reason the debuggee stopped and the faulting address
|
||||
// (for watchpoints, the watched address that was accessed). FreeBSD has no
|
||||
// PTRACE_GETSIGINFO; the stop's signal information comes from PT_LWPINFO,
|
||||
// whose pl_siginfo carries the siginfo the kernel delivered. A ptrace stop
|
||||
// with no signal behind it (a completed single-step, the initial attach)
|
||||
// fills no siginfo at all.
|
||||
func (s *Session) StopInfo() (StopReason, uint64) {
|
||||
if s.exited {
|
||||
return StopExited, 0
|
||||
}
|
||||
var info unix.PtraceLwpInfoStruct
|
||||
if err := unix.PtraceLwpInfo(s.pid, &info); err != nil {
|
||||
return StopNone, 0
|
||||
}
|
||||
// The siginfo layout is the FreeBSD siginfo_t: three leading ints
|
||||
// (signo, errno, code), then the union, 8-byte aligned, whose _fault
|
||||
// member puts the address at byte offset 16. The read is byte-wise
|
||||
// because the blob's alignment is not guaranteed.
|
||||
si := (*[64]byte)(unsafe.Pointer(&info.Siginfo))
|
||||
signo := int32(binary.LittleEndian.Uint32(si[0:4]))
|
||||
code := int32(binary.LittleEndian.Uint32(si[8:12]))
|
||||
switch {
|
||||
case signo == 0:
|
||||
// A pure ptrace stop: single-step completion, attach, or the
|
||||
// events the kernel resolves internally.
|
||||
return StopSingleStep, 0
|
||||
case signo != int32(syscall.SIGTRAP):
|
||||
return StopSignal, uint64(code)
|
||||
case code == trapBRKPT:
|
||||
return StopBreakpoint, 0
|
||||
case code == trapTRACE:
|
||||
addr := binary.LittleEndian.Uint64(si[16:24])
|
||||
return archStopTrace(s, addr)
|
||||
default:
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,195 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && (amd64 || arm64 || riscv64)
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"os"
|
||||
"runtime"
|
||||
"strconv"
|
||||
"strings"
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// mapRWX maps code into a read-write-execute region.
|
||||
func mapRWX(code []byte) ([]byte, error) {
|
||||
const pageSize = 4096
|
||||
size := (len(code) + pageSize - 1) &^ (pageSize - 1)
|
||||
mem, err := syscall.Mmap(-1, 0, size,
|
||||
syscall.PROT_READ|syscall.PROT_WRITE|syscall.PROT_EXEC,
|
||||
syscall.MAP_PRIVATE|syscall.MAP_ANON)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
copy(mem, code)
|
||||
return mem, nil
|
||||
}
|
||||
|
||||
// setupBuffers allocates buffers in the debuggee's memory.
|
||||
func setupBuffers(spec string, args []byte, tmpDir string) ([]byte, error) {
|
||||
type bufSpec struct {
|
||||
name string
|
||||
size int
|
||||
pattern string
|
||||
}
|
||||
var specs []bufSpec
|
||||
for part := range strings.SplitSeq(spec, ",") {
|
||||
fields := strings.SplitN(part, ":", 3)
|
||||
if len(fields) != 3 {
|
||||
continue
|
||||
}
|
||||
size, err := strconv.Atoi(fields[1])
|
||||
if err != nil || size <= 0 {
|
||||
continue
|
||||
}
|
||||
specs = append(specs, bufSpec{name: fields[0], size: size, pattern: fields[2]})
|
||||
}
|
||||
|
||||
if len(specs) == 0 {
|
||||
return args, nil
|
||||
}
|
||||
|
||||
var bufAddrs []uint64
|
||||
for _, s := range specs {
|
||||
buf, err := syscall.Mmap(-1, 0, s.size,
|
||||
syscall.PROT_READ|syscall.PROT_WRITE,
|
||||
syscall.MAP_PRIVATE|syscall.MAP_ANON)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("mmap buffer %s: %w", s.name, err)
|
||||
}
|
||||
fillBuffer(buf, s.pattern)
|
||||
bufAddrs = append(bufAddrs, uint64(uintptr(unsafe.Pointer(&buf[0]))))
|
||||
}
|
||||
|
||||
addrFile, err := os.Create(tmpDir + "/bufaddrs")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for _, addr := range bufAddrs {
|
||||
fmt.Fprintf(addrFile, "%d\n", addr)
|
||||
}
|
||||
addrFile.Close()
|
||||
|
||||
return args, nil
|
||||
}
|
||||
|
||||
// fillBuffer fills a buffer with the specified pattern.
|
||||
func fillBuffer(buf []byte, pattern string) {
|
||||
switch pattern {
|
||||
case "zero":
|
||||
case "ones":
|
||||
for i := range buf {
|
||||
buf[i] = 0xFF
|
||||
}
|
||||
case "seq":
|
||||
for i := range buf {
|
||||
buf[i] = byte(i)
|
||||
}
|
||||
default:
|
||||
if data, err := hex.DecodeString(pattern); err == nil && len(data) > 0 {
|
||||
for i := range buf {
|
||||
buf[i] = data[i%len(data)]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// RunTarget is the debuggee entry point (gasm debug --target).
|
||||
func RunTarget(asmPath, funcName, argsFile, tmpDir string) error {
|
||||
src, err := os.ReadFile(asmPath)
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug target: %w", err)
|
||||
}
|
||||
file, errs := parser.Parse(asmPath, string(src))
|
||||
if len(errs) > 0 {
|
||||
return fmt.Errorf("debug target: parse: %v", errs[0])
|
||||
}
|
||||
img, err := asm.AssembleFile(file)
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug target: assemble: %w", err)
|
||||
}
|
||||
|
||||
var fl *asm.FuncLayout
|
||||
for i := range img.Funcs {
|
||||
if img.Funcs[i].Name == funcName {
|
||||
fl = &img.Funcs[i]
|
||||
break
|
||||
}
|
||||
}
|
||||
if fl == nil {
|
||||
return fmt.Errorf("debug target: function %q not found", funcName)
|
||||
}
|
||||
|
||||
code := img.Bytes()
|
||||
exec, err := mapRWX(code)
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug target: mmap: %w", err)
|
||||
}
|
||||
|
||||
codeBase := uintptr(unsafe.Pointer(&exec[0]))
|
||||
if err := os.WriteFile(tmpDir+"/codebase", []byte(fmt.Sprintf("%d", codeBase)), 0o644); err != nil {
|
||||
return fmt.Errorf("debug target: write codebase: %w", err)
|
||||
}
|
||||
|
||||
meta := fmt.Sprintf("%d %d %d", fl.Offset, fl.Size, fl.Args)
|
||||
os.WriteFile(tmpDir+"/funcmeta", []byte(meta), 0o644)
|
||||
|
||||
labelsFile, _ := os.Create(tmpDir + "/labels")
|
||||
if labelsFile != nil {
|
||||
for label, off := range fl.Labels {
|
||||
fmt.Fprintf(labelsFile, "%s %d\n", label, off)
|
||||
}
|
||||
labelsFile.Close()
|
||||
}
|
||||
|
||||
args, err := os.ReadFile(argsFile)
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug target: read args: %w", err)
|
||||
}
|
||||
if len(args) < fl.Args {
|
||||
padded := make([]byte, fl.Args)
|
||||
copy(padded, args)
|
||||
args = padded
|
||||
}
|
||||
|
||||
bufSpecFile := tmpDir + "/bufspec"
|
||||
if bufSpec, err := os.ReadFile(bufSpecFile); err == nil && len(bufSpec) > 0 {
|
||||
args, err = setupBuffers(string(bufSpec), args, tmpDir)
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug target: setup buffers: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
runtime.LockOSThread()
|
||||
|
||||
if _, _, errno := unix.RawSyscall(unix.SYS_PTRACE, uintptr(unix.PT_TRACE_ME), 0, 0); errno != 0 {
|
||||
return fmt.Errorf("debug target: PT_TRACE_ME: %v", errno)
|
||||
}
|
||||
os.WriteFile(tmpDir+"/ready", []byte("ok"), 0o644)
|
||||
syscall.Kill(syscall.Getpid(), syscall.SIGSTOP)
|
||||
|
||||
os.WriteFile(tmpDir+"/entry", []byte("ok"), 0o644)
|
||||
syscall.Kill(syscall.Getpid(), syscall.SIGSTOP)
|
||||
|
||||
fnAddr := codeBase + uintptr(fl.Offset)
|
||||
stackArgs := make([]byte, fl.Args)
|
||||
copy(stackArgs, args)
|
||||
|
||||
if _, callErr := verify.Call(fnAddr, stackArgs); callErr != nil {
|
||||
os.Exit(1)
|
||||
}
|
||||
// Success returns to the caller, which exits with status 0; the JIT
|
||||
// code has already run to its own trampoline by the time Call returns.
|
||||
return nil
|
||||
}
|
||||
@@ -12,9 +12,9 @@ import (
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// RunTarget is the debuggee entry point (gasm debug --target).
|
||||
|
||||
@@ -12,9 +12,9 @@ import (
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// RunTarget is the debuggee entry point (gasm debug --target).
|
||||
|
||||
@@ -12,9 +12,9 @@ import (
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// RunTarget is the debuggee entry point (gasm debug --target).
|
||||
|
||||
@@ -12,9 +12,9 @@ import (
|
||||
"syscall"
|
||||
"unsafe"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/verify"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/verify"
|
||||
)
|
||||
|
||||
// RunTarget is the debuggee entry point (gasm debug --target).
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build linux
|
||||
//go:build linux || (freebsd && (amd64 || arm64 || riscv64))
|
||||
|
||||
package debug
|
||||
|
||||
|
||||
@@ -0,0 +1,190 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && amd64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"unsafe"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// Hardware watchpoint support via x86-64 debug registers (DR0-DR3, DR7),
|
||||
// read and written as one blob through PT_GETDBREGS/PT_SETDBREGS. The
|
||||
// FreeBSD struct dbreg is the raw DR file: dr[16], where DR0-DR3 are the
|
||||
// address registers, DR6 the status and DR7 the control (sys/x86/include/
|
||||
// reg.h; the DBREG_DRX accessor indexes the same array).
|
||||
|
||||
// dbreg mirrors FreeBSD's struct dbreg for PT_GETDBREGS/PT_SETDBREGS.
|
||||
type dbreg struct {
|
||||
Dr [16]uint64
|
||||
}
|
||||
|
||||
// dbreg indices of the registers the watchpoint layer drives.
|
||||
const (
|
||||
drStatus = 6 // DR6: the trap status register
|
||||
drControl = 7 // DR7: the debug control register
|
||||
)
|
||||
|
||||
// WatchpointType selects what triggers the watchpoint.
|
||||
type WatchpointType int
|
||||
|
||||
const (
|
||||
WatchWrite WatchpointType = 1 // trigger on write
|
||||
WatchRead WatchpointType = 3 // trigger on read or write
|
||||
)
|
||||
|
||||
// maxWatchpoints reports the number of hardware watchpoint slots the
|
||||
// architecture provides: four address registers, DR0-DR3.
|
||||
func maxWatchpoints() int { return 4 }
|
||||
|
||||
// getDbRegs reads the debug register file of the stopped debuggee.
|
||||
func (s *Session) getDbRegs() (*dbreg, error) {
|
||||
var dr dbreg
|
||||
if _, _, errno := unix.Syscall6(
|
||||
unix.SYS_PTRACE,
|
||||
uintptr(unix.PT_GETDBREGS),
|
||||
uintptr(s.pid),
|
||||
0,
|
||||
uintptr(unsafe.Pointer(&dr)),
|
||||
0, 0,
|
||||
); errno != 0 {
|
||||
return nil, errno
|
||||
}
|
||||
return &dr, nil
|
||||
}
|
||||
|
||||
// setDbRegs writes the debug register file of the stopped debuggee.
|
||||
func (s *Session) setDbRegs(dr *dbreg) error {
|
||||
if _, _, errno := unix.Syscall6(
|
||||
unix.SYS_PTRACE,
|
||||
uintptr(unix.PT_SETDBREGS),
|
||||
uintptr(s.pid),
|
||||
0,
|
||||
uintptr(unsafe.Pointer(dr)),
|
||||
0, 0,
|
||||
); errno != 0 {
|
||||
return errno
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// archStopTrace classifies a TRAP_TRACE stop. On amd64 the kernel
|
||||
// delivers both the completed single-step and the debug-register hit
|
||||
// through T_TRCTRAP with TRAP_TRACE (sys/amd64/amd64/trap.c), and DR6's
|
||||
// B0-B3 bits name the watchpoint that fired.
|
||||
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
if status := dr.Dr[drStatus]; status&0xF != 0 {
|
||||
for slot := range 4 {
|
||||
if status&(1<<slot) != 0 && dr.Dr[slot] != 0 {
|
||||
return StopWatchpoint, dr.Dr[slot]
|
||||
}
|
||||
}
|
||||
}
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
|
||||
// FindFreeWatchpointSlot returns the index of the first free watchpoint slot
|
||||
// (0-3), or -1 if all four hardware watchpoints are in use.
|
||||
func (s *Session) FindFreeWatchpointSlot() int {
|
||||
for i := range 4 {
|
||||
if !s.wpSlots[i] {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
// IsWatchpointSlotUsed reports whether slot (0-3) currently holds a watchpoint.
|
||||
func (s *Session) IsWatchpointSlotUsed(slot int) bool {
|
||||
if slot < 0 || slot > 3 {
|
||||
return false
|
||||
}
|
||||
return s.wpSlots[slot]
|
||||
}
|
||||
|
||||
// SetWatchpoint installs a hardware watchpoint on the given address.
|
||||
// DR7's encoding is architectural: a 2-bit local/global enable pair per
|
||||
// slot at bit 2*slot, the R/W field at 16+4*slot and the length field at
|
||||
// 18+4*slot (Intel SDM vol 3, "Debug Registers").
|
||||
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
|
||||
if slot < 0 || slot > 3 {
|
||||
return fmt.Errorf("debug: watchpoint slot must be 0-3")
|
||||
}
|
||||
if s.wpSlots[slot] {
|
||||
return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
|
||||
}
|
||||
|
||||
var lenBits uint64
|
||||
switch size {
|
||||
case 1:
|
||||
lenBits = 0
|
||||
case 2:
|
||||
lenBits = 1
|
||||
case 4:
|
||||
lenBits = 3
|
||||
case 8:
|
||||
lenBits = 2
|
||||
default:
|
||||
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
|
||||
}
|
||||
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug: read debug registers: %w", err)
|
||||
}
|
||||
dr.Dr[slot] = addr
|
||||
|
||||
dr7 := dr.Dr[drControl]
|
||||
enableBit := uint64(1) << (2 * slot)
|
||||
rwBits := uint64(typ) << (16 + 4*slot)
|
||||
lenField := lenBits << (18 + 4*slot)
|
||||
mask := ^((uint64(1) << (2 * slot)) | (uint64(3) << (16 + 4*slot)) | (uint64(3) << (18 + 4*slot)))
|
||||
dr.Dr[drControl] = (dr7 & mask) | enableBit | rwBits | lenField
|
||||
|
||||
if err := s.setDbRegs(dr); err != nil {
|
||||
return fmt.Errorf("debug: set debug registers: %w", err)
|
||||
}
|
||||
s.wpSlots[slot] = true
|
||||
return nil
|
||||
}
|
||||
|
||||
// ClearWatchpoint removes a hardware watchpoint.
|
||||
func (s *Session) ClearWatchpoint(slot int) error {
|
||||
if slot < 0 || slot > 3 {
|
||||
return fmt.Errorf("debug: watchpoint slot must be 0-3")
|
||||
}
|
||||
if !s.wpSlots[slot] {
|
||||
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
|
||||
}
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
dr.Dr[slot] = 0
|
||||
dr.Dr[drControl] &^= uint64(1) << (2 * slot)
|
||||
if err := s.setDbRegs(dr); err != nil {
|
||||
return err
|
||||
}
|
||||
s.wpSlots[slot] = false
|
||||
return nil
|
||||
}
|
||||
|
||||
// ClearAllWatchpoints removes all hardware watchpoints.
|
||||
func (s *Session) ClearAllWatchpoints() error {
|
||||
for slot := range maxWatchpoints() {
|
||||
if s.wpSlots[slot] {
|
||||
if err := s.ClearWatchpoint(slot); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,205 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && arm64
|
||||
|
||||
package debug
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"unsafe"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// Hardware watchpoint support via arm64 debug registers, read and written
|
||||
// as one blob through PT_GETDBREGS/PT_SETDBREGS. The FreeBSD struct dbreg
|
||||
// (sys/arm64/include/reg.h) opens with the debug-facility header and then
|
||||
// carries 16 breakpoint and 16 watchpoint pairs of {address, control}.
|
||||
|
||||
// dbreg mirrors FreeBSD's struct dbreg for PT_GETDBREGS/PT_SETDBREGS.
|
||||
type dbreg struct {
|
||||
DbDebugVer uint8
|
||||
DbNbkpts uint8
|
||||
DbNwtpts uint8
|
||||
_ [5]byte
|
||||
DbBreakregs [16]struct {
|
||||
Addr uint64
|
||||
Ctrl uint32
|
||||
_ uint32
|
||||
}
|
||||
DbWatchregs [16]struct {
|
||||
Addr uint64
|
||||
Ctrl uint32
|
||||
_ uint32
|
||||
}
|
||||
}
|
||||
|
||||
// WatchpointType selects what triggers the watchpoint.
|
||||
type WatchpointType int
|
||||
|
||||
const (
|
||||
WatchWrite WatchpointType = 1 // trigger on write
|
||||
WatchRead WatchpointType = 3 // trigger on read or write
|
||||
)
|
||||
|
||||
// maxWatchpoints reports the number of hardware watchpoint slots the
|
||||
// architecture provides: DBGWVR0-DBGWCR15.
|
||||
func maxWatchpoints() int { return 16 }
|
||||
|
||||
// getDbRegs reads the debug register file of the stopped debuggee.
|
||||
func (s *Session) getDbRegs() (*dbreg, error) {
|
||||
var dr dbreg
|
||||
if _, _, errno := unix.Syscall6(
|
||||
unix.SYS_PTRACE,
|
||||
uintptr(unix.PT_GETDBREGS),
|
||||
uintptr(s.pid),
|
||||
0,
|
||||
uintptr(unsafe.Pointer(&dr)),
|
||||
0, 0,
|
||||
); errno != 0 {
|
||||
return nil, errno
|
||||
}
|
||||
return &dr, nil
|
||||
}
|
||||
|
||||
// setDbRegs writes the debug register file of the stopped debuggee.
|
||||
func (s *Session) setDbRegs(dr *dbreg) error {
|
||||
if _, _, errno := unix.Syscall6(
|
||||
unix.SYS_PTRACE,
|
||||
uintptr(unix.PT_SETDBREGS),
|
||||
uintptr(s.pid),
|
||||
0,
|
||||
uintptr(unsafe.Pointer(dr)),
|
||||
0, 0,
|
||||
); errno != 0 {
|
||||
return errno
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// archStopTrace classifies a TRAP_TRACE stop. On arm64 the kernel
|
||||
// delivers both the software single step and the watchpoint hit through
|
||||
// EXCP_SOFTSTP_EL0/EXCP_WATCHPT_EL0 with TRAP_TRACE (sys/arm64/arm64/
|
||||
// trap.c); the watchpoint address rides the FAR register, so a stop whose
|
||||
// reported address falls inside an armed watchpoint's byte range is a
|
||||
// watchpoint and everything else is a single step.
|
||||
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
for slot := range 16 {
|
||||
ctrl := uint64(dr.DbWatchregs[slot].Ctrl)
|
||||
if ctrl&1 == 0 || dr.DbWatchregs[slot].Addr == 0 {
|
||||
continue
|
||||
}
|
||||
if bas := (ctrl >> 5) & 0xFF; bas != 0 && siAddr >= dr.DbWatchregs[slot].Addr && siAddr < dr.DbWatchregs[slot].Addr+8 {
|
||||
return StopWatchpoint, siAddr
|
||||
}
|
||||
}
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
|
||||
// FindFreeWatchpointSlot returns the index of the first free watchpoint
|
||||
// slot, or -1 if all of them are in use.
|
||||
func (s *Session) FindFreeWatchpointSlot() int {
|
||||
for i := range maxWatchpoints() {
|
||||
if !s.wpSlots[i] {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
// IsWatchpointSlotUsed reports whether slot currently holds a watchpoint.
|
||||
func (s *Session) IsWatchpointSlotUsed(slot int) bool {
|
||||
if slot < 0 || slot >= maxWatchpoints() {
|
||||
return false
|
||||
}
|
||||
return s.wpSlots[slot]
|
||||
}
|
||||
|
||||
// SetWatchpoint installs a hardware watchpoint on the given address. The
|
||||
// control word is the architectural DBGWCR (ARM DDI 0487): bit 0 enables,
|
||||
// bits 3-4 select the access type (10 store, 11 load+store) and bits 5-12
|
||||
// are the byte-address select, so the watch stays 8-byte aligned and names
|
||||
// its watched bytes through BAS.
|
||||
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
|
||||
if slot < 0 || slot >= maxWatchpoints() {
|
||||
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
|
||||
}
|
||||
if s.wpSlots[slot] {
|
||||
return fmt.Errorf("debug: watchpoint slot %d already in use", slot)
|
||||
}
|
||||
var bas uint64
|
||||
switch size {
|
||||
case 1:
|
||||
bas = 0x01
|
||||
case 2:
|
||||
bas = 0x03
|
||||
case 4:
|
||||
bas = 0x0F
|
||||
case 8:
|
||||
bas = 0xFF
|
||||
default:
|
||||
return fmt.Errorf("debug: watchpoint size must be 1, 2, 4, or 8")
|
||||
}
|
||||
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return fmt.Errorf("debug: read debug registers: %w", err)
|
||||
}
|
||||
if uint8(slot) >= dr.DbNwtpts && dr.DbNwtpts != 0 {
|
||||
return fmt.Errorf("debug: slot %d exceeds available watchpoints (%d)", slot, dr.DbNwtpts)
|
||||
}
|
||||
ctrl := uint64(1) // enable
|
||||
switch typ {
|
||||
case WatchWrite:
|
||||
ctrl |= 2 << 3 // store only
|
||||
case WatchRead:
|
||||
ctrl |= 3 << 3 // load+store
|
||||
}
|
||||
ctrl |= bas << 5
|
||||
dr.DbWatchregs[slot].Addr = addr
|
||||
dr.DbWatchregs[slot].Ctrl = uint32(ctrl)
|
||||
|
||||
if err := s.setDbRegs(dr); err != nil {
|
||||
return fmt.Errorf("debug: set debug registers: %w", err)
|
||||
}
|
||||
s.wpSlots[slot] = true
|
||||
return nil
|
||||
}
|
||||
|
||||
// ClearWatchpoint removes a hardware watchpoint.
|
||||
func (s *Session) ClearWatchpoint(slot int) error {
|
||||
if slot < 0 || slot >= maxWatchpoints() {
|
||||
return fmt.Errorf("debug: watchpoint slot must be 0-%d", maxWatchpoints()-1)
|
||||
}
|
||||
if !s.wpSlots[slot] {
|
||||
return fmt.Errorf("debug: watchpoint slot %d is not in use", slot)
|
||||
}
|
||||
dr, err := s.getDbRegs()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
dr.DbWatchregs[slot].Addr = 0
|
||||
dr.DbWatchregs[slot].Ctrl = 0
|
||||
if err := s.setDbRegs(dr); err != nil {
|
||||
return err
|
||||
}
|
||||
s.wpSlots[slot] = false
|
||||
return nil
|
||||
}
|
||||
|
||||
// ClearAllWatchpoints removes all hardware watchpoints.
|
||||
func (s *Session) ClearAllWatchpoints() error {
|
||||
for slot := range maxWatchpoints() {
|
||||
if s.wpSlots[slot] {
|
||||
if err := s.ClearWatchpoint(slot); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
//go:build freebsd && riscv64
|
||||
|
||||
package debug
|
||||
|
||||
import "fmt"
|
||||
|
||||
// The architecture has hardware watchpoint triggers, but FreeBSD exposes
|
||||
// no PT_GETDBREGS request for riscv64, so there is no supported way to arm
|
||||
// one: the watchpoint layer is honestly empty here.
|
||||
|
||||
// WatchpointType selects what triggers the watchpoint.
|
||||
type WatchpointType int
|
||||
|
||||
const (
|
||||
WatchWrite WatchpointType = 1 // trigger on write
|
||||
WatchRead WatchpointType = 3 // trigger on read or write
|
||||
)
|
||||
|
||||
// maxWatchpoints reports the number of hardware watchpoint slots the
|
||||
// platform provides: FreeBSD exposes none for riscv64.
|
||||
func maxWatchpoints() int { return 0 }
|
||||
|
||||
// archStopTrace classifies a TRAP_TRACE stop; with no watchpoint layer a
|
||||
// trace stop is always a completed single step.
|
||||
func archStopTrace(s *Session, siAddr uint64) (StopReason, uint64) {
|
||||
return StopSingleStep, 0
|
||||
}
|
||||
|
||||
// FindFreeWatchpointSlot returns -1: no slots exist.
|
||||
func (s *Session) FindFreeWatchpointSlot() int { return -1 }
|
||||
|
||||
// IsWatchpointSlotUsed reports whether slot currently holds a watchpoint.
|
||||
func (s *Session) IsWatchpointSlotUsed(slot int) bool { return false }
|
||||
|
||||
// SetWatchpoint is unsupported: FreeBSD exposes no debug register request
|
||||
// for riscv64.
|
||||
func (s *Session) SetWatchpoint(slot int, addr uint64, typ WatchpointType, size int) error {
|
||||
return fmt.Errorf("debug: hardware watchpoints are not supported on freebsd/riscv64")
|
||||
}
|
||||
|
||||
// ClearWatchpoint is unsupported for the same reason.
|
||||
func (s *Session) ClearWatchpoint(slot int) error {
|
||||
return fmt.Errorf("debug: hardware watchpoints are not supported on freebsd/riscv64")
|
||||
}
|
||||
|
||||
// ClearAllWatchpoints is a no-op: no watchpoint can be armed.
|
||||
func (s *Session) ClearAllWatchpoints() error { return nil }
|
||||
+1
-1
@@ -15,7 +15,7 @@ import (
|
||||
"golang.org/x/arch/riscv64/riscv64asm"
|
||||
"golang.org/x/arch/x86/x86asm"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
)
|
||||
|
||||
// Instruction is one decoded instruction: its text form, its length in bytes
|
||||
|
||||
@@ -7,10 +7,10 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-sdk/parser"
|
||||
)
|
||||
|
||||
func TestDecodeKnownBytes(t *testing.T) {
|
||||
|
||||
+30
-7
@@ -1,8 +1,8 @@
|
||||
# Architecture
|
||||
|
||||
How gasm-devkit is put together and why.
|
||||
How gasm-sdk is put together and why.
|
||||
|
||||
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
|
||||
Repository: [sourcedock.dev/petrbalvin/gasm-sdk](https://sourcedock.dev/petrbalvin/gasm-sdk)
|
||||
|
||||
## Overview
|
||||
|
||||
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
|
||||
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
|
||||
arch-agnostic and its output deterministic.
|
||||
|
||||
### Optional preprocessing
|
||||
|
||||
With `Options{Expand: true}` the parser runs a pre-parse pass
|
||||
(`preproc.go`) that splices `#include` files (the source directory, then the
|
||||
`-I` directories), expands object and parameterised `#define` macros,
|
||||
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
|
||||
folds constant expressions left in operands. The go command's platform
|
||||
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
|
||||
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
|
||||
language server read the raw file. The command layer adds the go_asm.h
|
||||
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
|
||||
defines type-checked out of its Go files for the target architecture and
|
||||
GOOS, with no compiler in the loop.
|
||||
|
||||
### `arch`
|
||||
|
||||
Register files are generated programmatically (the regular `R8`-`R15`,
|
||||
@@ -232,11 +246,18 @@ document store, republishes diagnostics on every change, and provides:
|
||||
pseudo-registers, labels, immediates, comments, directives, textflag macros);
|
||||
- **navigation**: go-to-definition from a label reference to its definition,
|
||||
find references, document highlights of every use of the symbol under the
|
||||
cursor, rename, and workspace symbol search over the open documents;
|
||||
cursor, rename, and workspace symbol search over the open documents and
|
||||
the indexed workspace files: the `.s` files under the workspace root that
|
||||
the editor has never opened, where an open buffer shadows its disk copy
|
||||
and watched-file events plus a per-query freshness check keep the index
|
||||
current;
|
||||
- **assists**: document formatting through the `format` package, inlay hints
|
||||
(the frame size after the TEXT argument area), signature help (the callee's
|
||||
`// func` signature while the cursor is on a `CALL`), and code actions
|
||||
offering quick fixes for the `missing-ret` and `unused-label` diagnostics.
|
||||
offering quick fixes for the `missing-ret` and `unused-label` diagnostics,
|
||||
the `missing-textflag-include` warning (the include after the last one in
|
||||
the file) and the `abi-argsize` warning (the argument area set to the size
|
||||
the `// func` signature implies).
|
||||
- **document information**: pull diagnostics (`textDocument/diagnostic`),
|
||||
#include document links (resolved against the document directory, then
|
||||
`$GOROOT/pkg/include`) and folding ranges (one collapsible region per
|
||||
@@ -558,7 +579,9 @@ AST, so neither depends on an encoding.
|
||||
--ground-truth`, `go list -json -export` locates the archives of the packages
|
||||
a GOOBJ object references, and `_gen` parses
|
||||
`$GOROOT/src/cmd/internal/obj/<arch>/anames.go` to rebuild the tables.
|
||||
- **Linux process interfaces** for the dynamic work: `mmap` and `mprotect` for
|
||||
the JIT mapping, ptrace with `/proc/pid/mem` for the debugger. That is why
|
||||
- **Linux and FreeBSD process interfaces** for the dynamic work: `mmap` and
|
||||
`mprotect` for the JIT mapping, ptrace for the debugger — with tracee
|
||||
memory through `/proc/pid/mem` on Linux and through `PT_IO` on FreeBSD,
|
||||
and the tracee's stop reports read from `PT_LWPINFO` there. That is why
|
||||
`verify` runs a JIT check only when the host architecture matches the
|
||||
kernel's, and why `debug` is Linux-only.
|
||||
kernel's, and why `debug` is bounded to those two kernels.
|
||||
|
||||
+13
-4
@@ -287,7 +287,8 @@ Usage: gasm debug <file.s> --func <name>
|
||||
The debugger re-executes the binary it is running as (`os.Executable()`) for the
|
||||
traced child, so the child is the same `gasm`, whether it is installed on `$PATH`
|
||||
or run with `go run ./cmd/gasm`; nothing has to be installed first. Requires
|
||||
Linux (ptrace), and all four architectures are supported.
|
||||
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64, arm64 and
|
||||
riscv64 on FreeBSD.
|
||||
|
||||
REPL commands:
|
||||
|
||||
@@ -365,7 +366,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
|
||||
## audit-instructions
|
||||
|
||||
```text
|
||||
Usage: gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]
|
||||
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
|
||||
```
|
||||
|
||||
Compare the gasm encoder for the given architecture (default amd64) against the
|
||||
@@ -403,7 +404,9 @@ architecture; a file without one is attempted for all four, exactly as a
|
||||
The report gives the headline number (files
|
||||
that assemble for every target architecture), the per-architecture pass rates
|
||||
and the most common failure reasons with one representative file each, which
|
||||
drive the encodability backlog by frequency rather than by table order. A run
|
||||
drive the encodability backlog by frequency rather than by table order. With
|
||||
`--list` the report additionally prints every failing file with its failure
|
||||
reason, per architecture. A run
|
||||
over GOROOT takes under a second.
|
||||
|
||||
```sh
|
||||
@@ -453,7 +456,13 @@ semantic tokens, go-to-definition, find references, rename, document
|
||||
formatting, inlay hints, code actions, signature help, document highlights,
|
||||
workspace symbol search, #include document links, and folding ranges for
|
||||
function bodies. Definition, references and rename work across every open
|
||||
document.
|
||||
document and the wider workspace on disk: the server indexes the `.s` files
|
||||
under the workspace root that the editor has never opened, an open buffer
|
||||
always shadows its disk copy, and watched-file events together with a
|
||||
per-query freshness check keep the index current. The quick fixes add the
|
||||
missing `#include "textflag.h"`, set the TEXT argument area to the size the
|
||||
`// func` signature implies, add a missing `RET`, and remove an unused
|
||||
label.
|
||||
|
||||
## version
|
||||
|
||||
|
||||
+3
-3
@@ -1,6 +1,6 @@
|
||||
# Development Guide
|
||||
|
||||
Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
|
||||
Repository: [sourcedock.dev/petrbalvin/gasm-sdk](https://sourcedock.dev/petrbalvin/gasm-sdk)
|
||||
|
||||
## Prerequisites
|
||||
|
||||
@@ -19,8 +19,8 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
|
||||
## Setup
|
||||
|
||||
```sh
|
||||
git clone https://sourcedock.dev/petrbalvin/gasm-devkit.git
|
||||
cd gasm-devkit
|
||||
git clone https://sourcedock.dev/petrbalvin/gasm-sdk.git
|
||||
cd gasm-sdk
|
||||
just build # compile bin/gasm, zero errors and zero warnings
|
||||
just gates # build, fmt-check, vet, test, race: the definition of done
|
||||
```
|
||||
|
||||
+657
@@ -0,0 +1,657 @@
|
||||
# The GOOBJ object file format
|
||||
|
||||
This document is a complete specification of GOOBJ, the object file format
|
||||
that the Go toolchain's assembler, compiler and linker exchange, written for
|
||||
implementers of independent producers and consumers. It documents the format
|
||||
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
|
||||
|
||||
No comparable document exists upstream. The format is defined only by the
|
||||
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
|
||||
internal interface with no stability promise, and it can change in any
|
||||
release. This specification was therefore produced by reverse engineering
|
||||
that source and by parsing real objects produced by `go tool asm` and
|
||||
`go tool compile`, byte for byte, against the layout described here. Within
|
||||
gasm-sdk it is kept honest by the differential tests in `asm/goobj_test.go`
|
||||
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
|
||||
the toolchain's own products and feed gasm objects to `go build`.
|
||||
|
||||
Every numeric value in this document, every block index, structure size, flag
|
||||
bit, type code and relocation number, was read from the Go 1.27.1 source at
|
||||
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
|
||||
`cmd/internal/objabi`, and exercised against assembled objects.
|
||||
|
||||
## Containers
|
||||
|
||||
The unit this document specifies is the **object**: one package's worth of
|
||||
symbols, relocations and data. An object is never consumed naked. Two
|
||||
wrappers exist in practice, and the linker dispatches on the first bytes of
|
||||
the file.
|
||||
|
||||
**The bare object**, written by `go tool asm`:
|
||||
|
||||
```text
|
||||
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
|
||||
"!\n"
|
||||
<GOOBJ blob>
|
||||
```
|
||||
|
||||
The first line is the toolchain configuration string, produced by
|
||||
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
|
||||
version, an optional architecture qualifier such as `GOAMD64=v1`, and
|
||||
`X:` followed by the enabled experiments, comma separated. The linker requires
|
||||
this line to match its own configuration exactly and rejects the file
|
||||
otherwise; the `-f` linker flag waives the check. Header lines may be
|
||||
followed by export data delimited by `$$` markers; the header region always
|
||||
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
|
||||
immediately after that line.
|
||||
|
||||
**The package archive**, written by the compiler output pipeline and consumed
|
||||
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
|
||||
data in a `__.PKGDEF` member and one or more objects as further members, each
|
||||
carrying the bare-object structure above. `go tool pack` creates and
|
||||
inspects such archives.
|
||||
|
||||
| Consumer | Role |
|
||||
|---|---|
|
||||
| `cmd/asm` | writes objects from `.s` files |
|
||||
| `cmd/compile` | writes objects from Go source |
|
||||
| `cmd/link` | reads objects and archives, produces executables |
|
||||
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
|
||||
|
||||
## Conventions
|
||||
|
||||
- All integers are **little endian**.
|
||||
- There is **no alignment or padding** anywhere in the file; structures follow
|
||||
one another byte by byte.
|
||||
- Every offset stored in the file is **relative to the first byte of the GOOBJ
|
||||
blob**, not to the start of the container.
|
||||
- The blob opens with a 96 byte header that carries the byte offset of every
|
||||
block. A block's length is the difference between its own offset and the
|
||||
next block's, so the offset array is the only index the format needs.
|
||||
|
||||
### Layout overview
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
A[Container header line and ! terminator] --> B[File header, 96 bytes]
|
||||
B --> C[String table, implicit region]
|
||||
C --> D[Autolib]
|
||||
D --> E[PkgIndex]
|
||||
E --> F[Files]
|
||||
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
|
||||
G --> H[RefFlags]
|
||||
H --> I[Hash64 and Hash]
|
||||
I --> J[RelocIndex, AuxIndex, DataIndex]
|
||||
J --> K[Relocs]
|
||||
K --> L[Aux]
|
||||
L --> M[Data]
|
||||
M --> N[RefNames]
|
||||
N --> O[BlkEnd marks the end of the blob]
|
||||
```
|
||||
|
||||
## The file header
|
||||
|
||||
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
|
||||
offsets.
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
|
||||
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
|
||||
| 16 | 4 | Flags | Bit field, see below. |
|
||||
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
|
||||
|
||||
Header flags:
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | ObjFlagShared | built with `-shared` |
|
||||
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
|
||||
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
|
||||
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
|
||||
| 4 | 16 | ObjFlagStd | standard library package |
|
||||
|
||||
### Block indices
|
||||
|
||||
The offset array is indexed by these constants, in file order:
|
||||
|
||||
| Index | Constant | Contents |
|
||||
|---|---|---|
|
||||
| 0 | BlkAutolib | imported packages |
|
||||
| 1 | BlkPkgIndex | referenced packages, indexed |
|
||||
| 2 | BlkFile | source file names |
|
||||
| 3 | BlkSymdef | symbol definitions, package scope |
|
||||
| 4 | BlkHashed64def | short hashed definitions |
|
||||
| 5 | BlkHasheddef | hashed definitions |
|
||||
| 6 | BlkNonpkgdef | non-package definitions |
|
||||
| 7 | BlkNonpkgref | non-package references |
|
||||
| 8 | BlkRefFlags | flags of referenced symbols |
|
||||
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
|
||||
| 10 | BlkHash | 16 byte hashes for hashed definitions |
|
||||
| 11 | BlkRelocIndex | per symbol relocation start index |
|
||||
| 12 | BlkAuxIndex | per symbol aux start index |
|
||||
| 13 | BlkDataIndex | per symbol data offset |
|
||||
| 14 | BlkReloc | relocations |
|
||||
| 15 | BlkAux | aux symbol entries |
|
||||
| 16 | BlkData | symbol payloads |
|
||||
| 17 | BlkRefName | names of referenced symbols, for tools |
|
||||
| 18 | BlkEnd | no contents; its offset is the end of the blob |
|
||||
|
||||
## The string table
|
||||
|
||||
There is no block index for strings. The table occupies the implicit region
|
||||
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
|
||||
string offset in the file points into that region. The writer de-duplicates:
|
||||
each distinct string is stored once, in first-use order, and the empty string
|
||||
is always the first entry, so its reference is length 0 and offset 96.
|
||||
|
||||
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
|
||||
offset of the bytes. The bytes are stored raw, with no terminator.
|
||||
|
||||
## Symbol references and the package index
|
||||
|
||||
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
|
||||
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
|
||||
lives in:
|
||||
|
||||
| Value | Constant | SymIdx indexes |
|
||||
|---|---|---|
|
||||
| 0 | PkgIdxInvalid | never valid in a written file |
|
||||
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
|
||||
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
|
||||
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
|
||||
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
|
||||
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
|
||||
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
|
||||
|
||||
Assignment rules, as the toolchain performs them:
|
||||
|
||||
- Every definition a package exports to the linker by index lands in Symdefs
|
||||
with PkgIdxSelf. The compiler puts its functions and data here; the
|
||||
assembler puts only its file-local static symbols here, everything else by
|
||||
name, see below.
|
||||
- External package references take indices 1, 2, 3, in order of first
|
||||
reference during assembly; the package names go into PkgIndex at those
|
||||
indices, entry 0 is the empty package and is never referenced.
|
||||
- References to the compiler's builtin functions become PkgIdxBuiltin with
|
||||
SymIdx set to the builtin's index.
|
||||
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
|
||||
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
|
||||
A producer must emit the definitions it made in NonPkgDefs and the pure
|
||||
references in NonPkgRefs.
|
||||
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
|
||||
is referenced by name, PkgIdxNone, **except** file-local static symbols,
|
||||
whose names carry `<>` and which are referenced by index. The compiler also
|
||||
forces references by name for symbols marked `//go:linkname` and for any
|
||||
symbol with the DUPOK attribute, which the linker de-duplicates by name.
|
||||
|
||||
## Symbol definition entries
|
||||
|
||||
The five definition and reference arrays (block indices 3 to 7) share one
|
||||
element layout, 21 bytes:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 8 | Name | string reference |
|
||||
| 8 | 2 | ABI | see table below |
|
||||
| 10 | 1 | Type | symbol kind, see the kind table |
|
||||
| 11 | 1 | Flag | bit field, see below |
|
||||
| 12 | 1 | Flag2 | second bit field, see below |
|
||||
| 13 | 4 | Siz | payload size in bytes, `uint32` |
|
||||
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
|
||||
|
||||
The Name is a real string reference for hand-written symbols. The auxiliary
|
||||
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
|
||||
entries, have empty names: length 0, and their identity is only via the Aux
|
||||
entries that point at them by index.
|
||||
|
||||
### The ABI field
|
||||
|
||||
| Value | Meaning |
|
||||
|---|---|
|
||||
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
|
||||
| 1 | ABIInternal, the register ABI of compiler-generated functions |
|
||||
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
|
||||
|
||||
### The Flag byte
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
|
||||
| 1 | 2 | SymFlagLocal | file-local |
|
||||
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
|
||||
| 3 | 8 | SymFlagLeaf | leaf function |
|
||||
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
|
||||
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
|
||||
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
|
||||
|
||||
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
|
||||
amd64 the assembler itself marks any function whose frame is below
|
||||
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
|
||||
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
|
||||
|
||||
### The Flag2 byte
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
|
||||
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
|
||||
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
|
||||
| 3 | 8 | SymFlagPkgInit | package initialisation function |
|
||||
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
|
||||
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
|
||||
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
|
||||
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
|
||||
|
||||
### The Type byte: symbol kinds
|
||||
|
||||
Values of `objabi.SymKind`, in numeric order:
|
||||
|
||||
| Value | Name | Meaning |
|
||||
|---|---|---|
|
||||
| 0 | Sxxx | invalid zero value |
|
||||
| 1 | STEXT | executable code |
|
||||
| 2 | STEXTFIPS | executable code, FIPS section |
|
||||
| 3 | SRODATA | read only data |
|
||||
| 4 | SRODATAFIPS | read only data, FIPS section |
|
||||
| 5 | SNOPTRDATA | data without pointers |
|
||||
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
|
||||
| 7 | SDATA | data, may contain pointers |
|
||||
| 8 | SDATAFIPS | data, FIPS section |
|
||||
| 9 | SBSS | zero initialised data |
|
||||
| 10 | SNOPTRBSS | zero initialised data without pointers |
|
||||
| 11 | STLSBSS | thread local zero initialised data |
|
||||
| 12 | SDWARFCUINFO | DWARF compile unit information |
|
||||
| 13 | SDWARFCONST | DWARF constants |
|
||||
| 14 | SDWARFFCN | DWARF function entry |
|
||||
| 15 | SDWARFABSFCN | DWARF absolute function entry |
|
||||
| 16 | SDWARFTYPE | DWARF type information |
|
||||
| 17 | SDWARFVAR | DWARF variable information |
|
||||
| 18 | SDWARFRANGE | DWARF range lists |
|
||||
| 19 | SDWARFLOC | DWARF location lists |
|
||||
| 20 | SDWARFLINES | DWARF line programs |
|
||||
| 21 | SDWARFADDR | DWARF address table |
|
||||
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
|
||||
| 23 | SCOVERAGE_COUNTER | coverage counter |
|
||||
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
|
||||
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
|
||||
|
||||
## Referenced symbol flags (RefFlags)
|
||||
|
||||
Element size 10 bytes, one per referenced external indexed symbol that
|
||||
carries a non-zero Flag2:
|
||||
|
||||
| Offset | Size | Field |
|
||||
|---|---|---|
|
||||
| 0 | 8 | Sym, a SymRef into another package |
|
||||
| 8 | 1 | Flag, always 0 in current writers |
|
||||
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
|
||||
|
||||
The linker uses these to preserve reachability of interface conversions
|
||||
across package boundaries. Entries with no flags are omitted entirely.
|
||||
|
||||
## Hashes
|
||||
|
||||
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
|
||||
a hash at all: the writer copies the **first 8 bytes of the symbol's
|
||||
payload**. Only symbols whose content-hash section byte is 0 may use the
|
||||
short form.
|
||||
|
||||
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
|
||||
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
|
||||
input, from `cmd/internal/obj/objfile.go`:
|
||||
|
||||
1. the payload size, little endian `uint64`;
|
||||
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
|
||||
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
|
||||
symbols, otherwise 0;
|
||||
3. for text symbols, the symbol name, which keeps distinct functions from
|
||||
merging;
|
||||
4. the payload with trailing zero bytes trimmed;
|
||||
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
|
||||
low type byte `uint8`, addend `int64`, followed by an encoding of the
|
||||
target: tag byte 0 then the target's short hash, tag 1 then its full
|
||||
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
|
||||
for PkgIdxSelf and imported packages, no tag, then the package path
|
||||
and the symbol index.
|
||||
|
||||
Two symbols with equal hashes are interchangeable at link time, which is what
|
||||
makes content addressing work. A producer that computes these hashes wrongly
|
||||
produces objects that link but de-duplicate wrongly; gasm verifies them by
|
||||
byte comparison against `go tool asm`.
|
||||
|
||||
## The index arrays
|
||||
|
||||
Three arrays of `uint32`, one element per **defined** symbol plus one final
|
||||
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
|
||||
defined symbols, each array holds N + 1 entries, and the entry at N is the
|
||||
total.
|
||||
|
||||
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
|
||||
entry i + 1 minus entry i is its count.
|
||||
- AuxIndex: the same construction over BlkAux.
|
||||
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
|
||||
the count is the difference of neighbours.
|
||||
|
||||
The toolchain writes relocations grouped per symbol in definition order, and
|
||||
sorts each symbol's relocations by their Off field first. A producer that
|
||||
skips the sort produces objects the linker still accepts, but that no longer
|
||||
compare byte-for-byte with the toolchain's output.
|
||||
|
||||
## Relocations
|
||||
|
||||
Element size 23 bytes:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
|
||||
| 4 | 1 | Siz | patch width in bytes |
|
||||
| 5 | 2 | Type | relocation type, `uint16`, see the table |
|
||||
| 7 | 8 | Add | addend, `int64` |
|
||||
| 15 | 8 | Sym | target SymRef |
|
||||
|
||||
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
|
||||
flavour the type prescribes is the linker's job; the object only records the
|
||||
request. A size 0 relocation patches nothing and exists purely as a marker
|
||||
for the linker's reachability analysis.
|
||||
|
||||
### Relocation types
|
||||
|
||||
Values of `objabi.RelocType`. The assembler and compiler emit the generic
|
||||
ones plus their own architecture's family; the rest exist for other ports and
|
||||
for the linker itself.
|
||||
|
||||
| Value | Name | Meaning |
|
||||
|---|---|---|
|
||||
| 1 | R_ADDR | absolute address |
|
||||
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
|
||||
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
|
||||
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
|
||||
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
|
||||
| 6 | R_SIZE | size of the referenced symbol |
|
||||
| 7 | R_CALL | direct call, PC relative |
|
||||
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
|
||||
| 9 | R_CALLARM64 | arm64: BL |
|
||||
| 10 | R_CALLIND | indirect call marker |
|
||||
| 11 | R_CALLPOWER | ppc64: call |
|
||||
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
|
||||
| 13 | R_CONST | constant value of the symbol |
|
||||
| 14 | R_PCREL | PC relative displacement |
|
||||
| 15 | R_TLS_LE | thread local, local exec offset |
|
||||
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
|
||||
| 17 | R_GOTOFF | offset from the GOT base |
|
||||
| 18 | R_PLT0 | PLT sequence, first instruction |
|
||||
| 19 | R_PLT1 | PLT sequence, second instruction |
|
||||
| 20 | R_PLT2 | PLT sequence, third instruction |
|
||||
| 21 | R_USEFIELD | field reachability marker |
|
||||
| 22 | R_USETYPE | type reachability marker, no bytes patched |
|
||||
| 23 | R_USEIFACE | interface conversion marker, size 0 |
|
||||
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
|
||||
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
|
||||
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
|
||||
| 27 | R_KEEP | keeps the target alive if the source survives |
|
||||
| 28 | R_POWER_TOC | ppc64: TOC relative |
|
||||
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
|
||||
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
|
||||
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
|
||||
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
|
||||
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
|
||||
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
|
||||
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
|
||||
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
|
||||
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
|
||||
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
|
||||
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
|
||||
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
|
||||
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
|
||||
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
|
||||
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
|
||||
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
|
||||
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
|
||||
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
|
||||
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
|
||||
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
|
||||
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
|
||||
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
|
||||
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
|
||||
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
|
||||
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
|
||||
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
|
||||
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
|
||||
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
|
||||
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
|
||||
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
|
||||
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
|
||||
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
|
||||
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
|
||||
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
|
||||
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
|
||||
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
|
||||
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
|
||||
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
|
||||
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
|
||||
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
|
||||
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
|
||||
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
|
||||
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
|
||||
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
|
||||
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
|
||||
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
|
||||
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
|
||||
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
|
||||
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
|
||||
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
|
||||
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
|
||||
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
|
||||
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
|
||||
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
|
||||
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
|
||||
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
|
||||
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
|
||||
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
|
||||
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
|
||||
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
|
||||
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
|
||||
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
|
||||
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
|
||||
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
|
||||
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
|
||||
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
|
||||
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
|
||||
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
|
||||
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
|
||||
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
|
||||
| 99 | R_WASMIMPORT | wasm: import module and name indices |
|
||||
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
|
||||
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
|
||||
| 102 | R_INITORDER | orders inittask records, patches nothing |
|
||||
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
|
||||
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
|
||||
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
|
||||
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
|
||||
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
|
||||
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
|
||||
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
|
||||
|
||||
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
|
||||
type's value with bit 15 set. The linker strips the bit before dispatch.
|
||||
|
||||
## Aux symbol entries
|
||||
|
||||
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
|
||||
auxiliary symbols to a definition; the arrays run per symbol in the order
|
||||
given by AuxIndex.
|
||||
|
||||
| Value | Name | Attaches |
|
||||
|---|---|---|
|
||||
| 0 | AuxGotype | the Go type of a data symbol |
|
||||
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
|
||||
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
|
||||
| 3 | AuxDwarfInfo | DWARF debug info for the function |
|
||||
| 4 | AuxDwarfLoc | DWARF location lists |
|
||||
| 5 | AuxDwarfRanges | DWARF range lists |
|
||||
| 6 | AuxDwarfLines | DWARF line program |
|
||||
| 7 | AuxPcsp | pc-value table: SP adjustments |
|
||||
| 8 | AuxPcfile | pc-value table: source file indices |
|
||||
| 9 | AuxPcline | pc-value table: line numbers |
|
||||
| 10 | AuxPcinline | pc-value table: inlining tree positions |
|
||||
| 11 | AuxPcdata | one pc-value table per live variable slot |
|
||||
| 12 | AuxWasmImport | wasm import description |
|
||||
| 13 | AuxWasmType | wasm export type description |
|
||||
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
|
||||
|
||||
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
|
||||
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
|
||||
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
|
||||
payload would be empty. A function assembled from `.s` source by Go 1.27.1
|
||||
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
|
||||
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
|
||||
same set.
|
||||
|
||||
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
|
||||
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
|
||||
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
|
||||
resolve to definitions in the package's compiled Go code when there is any.
|
||||
|
||||
## Symbol payloads (BlkData)
|
||||
|
||||
The payloads of all defined symbols, in definition order, concatenated with
|
||||
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
|
||||
its machine code, with the stack-split preamble and any morestack block
|
||||
already included. A data symbol's payload is the bytes laid down by its DATA
|
||||
directives, zero filled to its declared size. If a symbol was created from an
|
||||
embedded file, the file's bytes follow the payload and count towards its
|
||||
DataIndex extent; assembly producers never write this extension.
|
||||
|
||||
### The FuncInfo payload
|
||||
|
||||
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
|
||||
little endian:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
|
||||
| 4 | 4 | Locals | frame size in bytes |
|
||||
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
|
||||
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
|
||||
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
|
||||
| 12 | 4 | StartLine | source line of the TEXT declaration |
|
||||
| 16 | 4 | NumFile | count of file indices that follow |
|
||||
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
|
||||
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
|
||||
| then | 24 × NumInlTree | InlTree | nodes, see below |
|
||||
|
||||
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
|
||||
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
|
||||
`int32` parent PC.
|
||||
|
||||
The assembler derives FuncID from the symbol name through
|
||||
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
|
||||
specially gets that classification even when defined in assembly; an ordinary
|
||||
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
|
||||
|
||||
### The pc-value tables
|
||||
|
||||
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
|
||||
pc-value tables, each a sequence of value deltas and PC deltas:
|
||||
|
||||
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
|
||||
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
|
||||
raw delta divided by the architecture's minimum instruction length;
|
||||
- the table ends with a final PC delta to the end of the function followed by
|
||||
a zero byte.
|
||||
|
||||
The first value applies from function entry. The encoding is the one
|
||||
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
|
||||
final runtime pclntable carries.
|
||||
|
||||
### The DWARF payloads
|
||||
|
||||
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
|
||||
symbols whose payloads are DWARF byte streams. The object format treats them
|
||||
as opaque: the linker concatenates them into the final `.debug_*` sections
|
||||
and resolves the relocations recorded inside them. The compiler produces
|
||||
DWARF content per its own generation; gasm produces DWARF5 streams in
|
||||
`asm/goobj_dwarf.go`.
|
||||
|
||||
## Builtins
|
||||
|
||||
Frequently referenced runtime functions are referenced by index rather than
|
||||
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
|
||||
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
|
||||
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
|
||||
ABI 0. Builtin names never enter the string table. The mapping only applies
|
||||
while the object is not linked against shared libraries, and a linkname'd
|
||||
symbol never counts as a builtin even when its name matches.
|
||||
|
||||
## Fingerprints
|
||||
|
||||
The 8 byte fingerprint identifies one build of a package. The compiler fills
|
||||
it with a hash of the package's export data; the assembler leaves it zero.
|
||||
The linker checks a package's fingerprint against the fingerprints its
|
||||
importers recorded in their Autolib entries and rejects a mismatched build,
|
||||
which is how stale objects are caught.
|
||||
|
||||
## What a producer must do
|
||||
|
||||
The checklist a third-party writer must satisfy for `go build` to accept its
|
||||
objects, in one place:
|
||||
|
||||
1. Write the container exactly: the `go object` line matching the target
|
||||
toolchain's configuration string, the `!\n` terminator, then the blob.
|
||||
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
|
||||
3. Deduplicate the string table, keep the empty string at offset 96, and
|
||||
reference it everywhere a name appears.
|
||||
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
|
||||
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
|
||||
5. Sort relocations by offset within each symbol.
|
||||
6. Fill Siz with the true payload length, set Align for every
|
||||
content-addressable symbol, and keep symbols under 2 GB.
|
||||
7. Reference symbols by the package-index rules. An assembly producer
|
||||
references everything outside the object by name, PkgIdxNone,
|
||||
except its own file-local statics and the builtins; PkgIdxSelf is
|
||||
reserved for definitions in this object. Assembly TEXT symbols
|
||||
carry ABI 0.
|
||||
8. Compute the content hashes exactly as the toolchain does, or emit no
|
||||
hashed definitions at all.
|
||||
|
||||
## How gasm-sdk implements and verifies it
|
||||
|
||||
The writer lives in `asm/goobj.go`, which carries the shared container and the
|
||||
amd64 relocation emission, with per-architecture relocation emitters in
|
||||
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
|
||||
resolution in `asm/goobj_resolve.go` and DWARF generation in
|
||||
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
|
||||
that `go build` consumes in place of the toolchain's own.
|
||||
|
||||
Verification is differential and continuous:
|
||||
|
||||
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
|
||||
output for the same source, byte for byte;
|
||||
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
|
||||
objects and runs them;
|
||||
- `gasm verify` keeps the machine code itself identical to the toolchain's,
|
||||
which is the precondition for the object comparison to be meaningful.
|
||||
|
||||
## Versioning and drift
|
||||
|
||||
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
|
||||
a toolchain release changes the format, it changes that string first, and the
|
||||
linker refuses blobs whose magic it does not know. The watch points for a new
|
||||
release are, in order: the magic, the block index list, the Aux type list,
|
||||
the tail of the relocation table, the FuncInfo layout, and the builtin table
|
||||
count. gasm's tests fail against any of these changes, which is the mechanism
|
||||
that keeps this document and the writer current.
|
||||
|
||||
The authoritative sources, for the release this document covers:
|
||||
|
||||
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
|
||||
document;
|
||||
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
|
||||
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
|
||||
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
|
||||
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
|
||||
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
|
||||
- `cmd/internal/objabi/reloctype.go`: relocation types;
|
||||
- `cmd/internal/objabi/symkind.go`: symbol kinds;
|
||||
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user