diff --git a/CHANGELOG.md b/CHANGELOG.md index a2d278e..ad013f8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -25,8 +25,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **`gasm audit-instructions --corpus [dir]`.** Assembles every `.s` file under a directory (default GOROOT/src) with the gasm encoder only: suffixed files for their architecture, suffix-less files for - all four, as a GOARCH build would. Reports the headline number (108 - of 627 GOROOT files, 17.2 %, assemble for every target architecture, + all four, as a GOARCH build would. Reports the headline number (127 + of 627 GOROOT files, 20.3 %, assemble for every target architecture, against 23 in the previous release), the per-architecture pass rates and the most common failure reasons with a representative file each, which drive the encodability backlog by frequency. @@ -263,6 +263,35 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 exit-2 contract now holds, `asm -o` no longer prints the hex dump it claimed to replace, and `verify --ground-truth` works for amd64 kernels on non-amd64 hosts instead of refusing with JIT advice. +- **arm64 store-exclusive instructions read their operands in the + toolchain's order.** `STXR` treated the first register as the status + register where `go tool asm` reads it as the data register, so the + same source assembled to different code in the two assemblers; the + pair forms (`STXP`, `LDXP` and their acquire/release variants) are + accepted now, in the toolchain spelling. +- **Large arm64 frames matched the toolchain's sequences.** A frame + beyond the immediate range that is not a movcon constant (roughly + 64 KiB and up) made `gasm verify` report a false mismatch: the + toolchain splits the prologue subtraction into two 12-bit immediates + and materialises the non-leaf epilogue addition through the temporary + register; gasm emits the same sequences and the spadj boundaries + follow the real word counts. +- **The width spellings GOROOT uses assemble.** `MOVLQZX` (four uses in + `runtime/asm_amd64.s`), `MOVBQSX`, `MOVWQSX`, `MOVBLSX`, `MOVBWSX`, + `MOVBWZX` and `PMOVMSKB` (the bytealg kernels) encode byte-identically + with `go tool asm`, and the linter reports them encodable; a + `MOVLQZX` is the plain 32-bit move, exactly as the toolchain lowers + it. +- **`verify --ground-truth` no longer reports a mismatch for functions + whose size is not a multiple of 16.** The toolchain pads text symbols + to 16-byte boundaries; the comparison now checks the padding is zero + instead of comparing it, the same rule the test suite applies. +- **riscv64 accepts the `g` spelling of the goroutine register**, like + the other architectures, and the abi kernels use it; every verify + kernel is now ground-truth checkable (the numeric `X27` spelling the + kernels used is one `go tool asm` rejects). +- The GOROOT corpus number rose to 127 of 627 files (20.3 %) assembling + for every target architecture, from 108. ## [0.33.0] - 2026-09-14 @@ -481,7 +510,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [0.31.0] - 2026-08-20 -The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission, +The arm64 encoder ships with ELF64 and GOOBJ emission, verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a real `go build`. The encoder covers the full integer instruction set, FP arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with @@ -489,7 +518,7 @@ bitmask immediate encoding. The project now requires Go 1.27. ### Added -- **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s` +- **arm64 encoder.** `gasm asm` can now assemble `_arm64.s` files: the AArch64 integer instruction set with the MOV pseudo-instruction and its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with logical bitmask encoding for values like `$1`), data-processing (shifted @@ -497,8 +526,9 @@ bitmask immediate encoding. The project now requires Go 1.27. immediate), conditional and unconditional branches, FP/SP frame mapping, SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth - verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 - (the other architectures; RISC-V, LoongArch, arm64) is now complete. + verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. The + encoder set for the remaining architectures (RISC-V, LoongArch, arm64) is + complete. ### Changed @@ -508,7 +538,7 @@ bitmask immediate encoding. The project now requires Go 1.27. ## [0.30.0] - 2026-08-13 -The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified +The LoongArch encoder ships with ELF64 and GOOBJ emission, verified byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real `go build`; the shared GOOBJ emitter now writes the per-function DWARF symbols the linker's DWARF pass reads. The RISC-V encoder reaches byte-for-byte parity @@ -519,7 +549,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only. ### Added -- **LoongArch encoder (Phase 5).** `gasm asm` can now assemble `_loong64.s` +- **LoongArch encoder.** `gasm asm` can now assemble `_loong64.s` files: the full LoongArch64 instruction set with the dual-form arithmetic mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and its immediate-constant expansions, FP/SP frame mapping, SB/global symbol @@ -608,7 +638,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only. - **Linux only.** The toolkit, its CI and the released binaries are now Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}. -- **Phase 4 closed.** README's "Remaining" list for the debugger is gone; +- **Debugger complete.** README's "Remaining" list for the debugger is gone; disassembly at PC, memory-write, watchpoints, and source-line mapping are all shipped. @@ -818,7 +848,7 @@ exposes the full dynamic-analysis toolkit. ## [0.20.0] - 2026-07-25 -Coverage profiling: the third pillar of Phase 3. Static basic-block +Coverage profiling. Static basic-block enumeration from the assembler's label map, combined with multi-input path diversity measurement; how many observationally distinct execution paths a test corpus exercises. @@ -843,7 +873,7 @@ execute) without fighting the runtime. ## [0.19.0] - 2026-07-24 -Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now +Runtime ABI checks. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers (BP, R14) before entering the assembled function and verifies they survive on return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that @@ -878,7 +908,7 @@ codes. This is the automated form of the project's bit-identical contract. ## [0.17.0] - 2026-07-22 -Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles +Dynamic analysis. A JIT execution substrate that assembles Plan 9 amd64 kernels into executable memory and calls them directly; pure Go (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external toolchain. @@ -1285,8 +1315,8 @@ support). ## [0.2.0] - 2026-07-07 -The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute -and the moves, on top of the Phase 1 VEX forms. +The assembler grows the SIMD set: shuffles, extract/insert, permute +and the moves, on top of the VEX forms of the first release. ### Added @@ -1322,7 +1352,7 @@ and the moves, on top of the Phase 1 VEX forms. ## [0.1.0] - 2026-07-06 -Initial release; the Phase 1 foundation. +Initial release: the foundation. ### Added diff --git a/README.md b/README.md index e98161d..ae9084c 100644 --- a/README.md +++ b/README.md @@ -119,7 +119,7 @@ can emit today is narrower, and a recognised but unencodable instruction is reported as an explicit error, never as a wrong byte. The same measurement runs over GOROOT's whole assembly corpus: -`gasm audit-instructions --corpus` reports 108 of 627 files (17.2 %) +`gasm audit-instructions --corpus` reports 127 of 627 files (20.3 %) assembling for every target architecture today, with the top failure reasons per architecture; the number moves with every release.