Compare commits

...
8 Commits
Author SHA1 Message Date
petrbalvin e836d6150d docs: changelog entry for the loong64 JIT enablement
Test / test (push) Successful in 2m7s
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin 8a51b060da feat(cmd): enable loong64 JIT execution, all trampolines qemu-validated
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin d3d47db727 test(verify): seed the arm64 ABI kernel arguments
Assisted-by: GLM 5.3
2026-09-20 00:57:02 +02:00
petrbalvin 0758556b7d docs: changelog entries for the parity round and corpus number
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin ddb8440340 fix(cmd): padding-aware ground-truth comparison
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin f15ff66fb1 fix(riscv64): accept the g spelling of the goroutine register
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin 187e4856d3 feat(amd64): encode the mixed-width extend family and PMOVMSKB
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
petrbalvin d315a998ce fix(arm64): store-exclusive operand order and large-frame parity
Assisted-by: GLM 5.3
2026-09-20 00:38:24 +02:00
24 changed files with 635 additions and 111 deletions
+54 -15
View File
@@ -25,8 +25,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **`gasm audit-instructions --corpus [dir]`.** Assembles every `.s` - **`gasm audit-instructions --corpus [dir]`.** Assembles every `.s`
file under a directory (default GOROOT/src) with the gasm encoder file under a directory (default GOROOT/src) with the gasm encoder
only: suffixed files for their architecture, suffix-less files for only: suffixed files for their architecture, suffix-less files for
all four, as a GOARCH build would. Reports the headline number (108 all four, as a GOARCH build would. Reports the headline number (127
of 627 GOROOT files, 17.2 %, assemble for every target architecture, of 627 GOROOT files, 20.3 %, assemble for every target architecture,
against 23 in the previous release), the per-architecture pass rates against 23 in the previous release), the per-architecture pass rates
and the most common failure reasons with a representative file each, and the most common failure reasons with a representative file each,
which drive the encodability backlog by frequency. which drive the encodability backlog by frequency.
@@ -263,6 +263,44 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
exit-2 contract now holds, `asm -o` no longer prints the hex dump it exit-2 contract now holds, `asm -o` no longer prints the hex dump it
claimed to replace, and `verify --ground-truth` works for amd64 claimed to replace, and `verify --ground-truth` works for amd64
kernels on non-amd64 hosts instead of refusing with JIT advice. kernels on non-amd64 hosts instead of refusing with JIT advice.
- **arm64 store-exclusive instructions read their operands in the
toolchain's order.** `STXR` treated the first register as the status
register where `go tool asm` reads it as the data register, so the
same source assembled to different code in the two assemblers; the
pair forms (`STXP`, `LDXP` and their acquire/release variants) are
accepted now, in the toolchain spelling.
- **Large arm64 frames matched the toolchain's sequences.** A frame
beyond the immediate range that is not a movcon constant (roughly
64 KiB and up) made `gasm verify` report a false mismatch: the
toolchain splits the prologue subtraction into two 12-bit immediates
and materialises the non-leaf epilogue addition through the temporary
register; gasm emits the same sequences and the spadj boundaries
follow the real word counts.
- **The width spellings GOROOT uses assemble.** `MOVLQZX` (four uses in
`runtime/asm_amd64.s`), `MOVBQSX`, `MOVWQSX`, `MOVBLSX`, `MOVBWSX`,
`MOVBWZX` and `PMOVMSKB` (the bytealg kernels) encode byte-identically
with `go tool asm`, and the linter reports them encodable; a
`MOVLQZX` is the plain 32-bit move, exactly as the toolchain lowers
it.
- **`verify --ground-truth` no longer reports a mismatch for functions
whose size is not a multiple of 16.** The toolchain pads text symbols
to 16-byte boundaries; the comparison now checks the padding is zero
instead of comparing it, the same rule the test suite applies.
- **riscv64 accepts the `g` spelling of the goroutine register**, like
the other architectures, and the abi kernels use it; every verify
kernel is now ground-truth checkable (the numeric `X27` spelling the
kernels used is one `go tool asm` rejects).
- The GOROOT corpus number rose to 127 of 627 files (20.3 %) assembling
for every target architecture, from 108.
- **loong64 JIT execution enabled.** The loong64 trampoline is now
validated end to end under qemu-user emulation (plain and ABI-checked
calls, goroutine-clobber detection), so `gasm verify` runs the JIT
checks on loong64 hosts instead of forcing every loong64 kernel down
the ground-truth path. The arm64 and riscv64 trampolines carry the
same validation; the arm64 ABI test now seeds its kernel arguments
(a zeroed block made the passthrough check meaningless), and the
loong64 basic kernel's branch maze terminates on every path so the
smoke sweep cannot spin on leftover register values.
## [0.33.0] - 2026-09-14 ## [0.33.0] - 2026-09-14
@@ -481,7 +519,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [0.31.0] - 2026-08-20 ## [0.31.0] - 2026-08-20
The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission, The arm64 encoder ships with ELF64 and GOOBJ emission,
verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a
real `go build`. The encoder covers the full integer instruction set, FP real `go build`. The encoder covers the full integer instruction set, FP
arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with
@@ -489,7 +527,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
### Added ### Added
- **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s` - **arm64 encoder.** `gasm asm` can now assemble `_arm64.s`
files: the AArch64 integer instruction set with the MOV pseudo-instruction and files: the AArch64 integer instruction set with the MOV pseudo-instruction and
its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with
logical bitmask encoding for values like `$1`), data-processing (shifted logical bitmask encoding for values like `$1`), data-processing (shifted
@@ -497,8 +535,9 @@ bitmask immediate encoding. The project now requires Go 1.27.
immediate), conditional and unconditional branches, FP/SP frame mapping, immediate), conditional and unconditional branches, FP/SP frame mapping,
SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations),
jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth
verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. The
(the other architectures; RISC-V, LoongArch, arm64) is now complete. encoder set for the remaining architectures (RISC-V, LoongArch, arm64) is
complete.
### Changed ### Changed
@@ -508,7 +547,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
## [0.30.0] - 2026-08-13 ## [0.30.0] - 2026-08-13
The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified The LoongArch encoder ships with ELF64 and GOOBJ emission, verified
byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real
`go build`; the shared GOOBJ emitter now writes the per-function DWARF symbols `go build`; the shared GOOBJ emitter now writes the per-function DWARF symbols
the linker's DWARF pass reads. The RISC-V encoder reaches byte-for-byte parity the linker's DWARF pass reads. The RISC-V encoder reaches byte-for-byte parity
@@ -519,7 +558,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
### Added ### Added
- **LoongArch encoder (Phase 5).** `gasm asm` can now assemble `_loong64.s` - **LoongArch encoder.** `gasm asm` can now assemble `_loong64.s`
files: the full LoongArch64 instruction set with the dual-form arithmetic files: the full LoongArch64 instruction set with the dual-form arithmetic
mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and
its immediate-constant expansions, FP/SP frame mapping, SB/global symbol its immediate-constant expansions, FP/SP frame mapping, SB/global symbol
@@ -608,7 +647,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
- **Linux only.** The toolkit, its CI and the released binaries are now - **Linux only.** The toolkit, its CI and the released binaries are now
Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}. Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}.
- **Phase 4 closed.** README's "Remaining" list for the debugger is gone; - **Debugger complete.** README's "Remaining" list for the debugger is gone;
disassembly at PC, memory-write, watchpoints, and source-line mapping are disassembly at PC, memory-write, watchpoints, and source-line mapping are
all shipped. all shipped.
@@ -818,7 +857,7 @@ exposes the full dynamic-analysis toolkit.
## [0.20.0] - 2026-07-25 ## [0.20.0] - 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block Coverage profiling. Static basic-block
enumeration from the assembler's label map, combined with multi-input path enumeration from the assembler's label map, combined with multi-input path
diversity measurement; how many observationally distinct execution paths a diversity measurement; how many observationally distinct execution paths a
test corpus exercises. test corpus exercises.
@@ -843,7 +882,7 @@ execute) without fighting the runtime.
## [0.19.0] - 2026-07-24 ## [0.19.0] - 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now Runtime ABI checks. The JIT trampoline now
has an ABI-checking variant that sets sentinels in the callee-saved registers has an ABI-checking variant that sets sentinels in the callee-saved registers
(BP, R14) before entering the assembled function and verifies they survive on (BP, R14) before entering the assembled function and verifies they survive on
return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that
@@ -878,7 +917,7 @@ codes. This is the automated form of the project's bit-identical contract.
## [0.17.0] - 2026-07-22 ## [0.17.0] - 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles Dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly; pure Go Plan 9 amd64 kernels into executable memory and calls them directly; pure Go
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
toolchain. toolchain.
@@ -1285,8 +1324,8 @@ support).
## [0.2.0] - 2026-07-07 ## [0.2.0] - 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute The assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms. and the moves, on top of the VEX forms of the first release.
### Added ### Added
@@ -1322,7 +1361,7 @@ and the moves, on top of the Phase 1 VEX forms.
## [0.1.0] - 2026-07-06 ## [0.1.0] - 2026-07-06
Initial release; the Phase 1 foundation. Initial release: the foundation.
### Added ### Added
+1 -1
View File
@@ -119,7 +119,7 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte. reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus: The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 108 of 627 files (17.2 %) `gasm audit-instructions --corpus` reports 127 of 627 files (20.3 %)
assembling for every target architecture today, with the top failure assembling for every target architecture today, with the top failure
reasons per architecture; the number moves with every release. reasons per architecture; the number moves with every release.
+60 -18
View File
@@ -314,7 +314,8 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
return encodeARM64CRC32(mnem, enc.op, ops) return encodeARM64CRC32(mnem, enc.op, ops)
} }
// Exclusive load/store (LDXR, STXR, LDAXR, STLXR). // Exclusive load/store (LDXR, STXR, LDAXR, STLXR and the register-pair
// forms LDXP, STXP).
if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FExcl { if enc, ok := a64InstrTable[mnem]; ok && enc.format == a64FExcl {
return encodeARM64Excl(mnem, enc.op, ops) return encodeARM64Excl(mnem, enc.op, ops)
} }
@@ -771,14 +772,17 @@ func encodeARM64LoadImm(rd int, v int64, mnem string) ([]byte, error) {
return a64wordLE(op | 31<<16 | 31<<5 | uint32(rd)), nil return a64wordLE(op | 31<<16 | 31<<5 | uint32(rd)), nil
} }
// The Go toolchain classifies immediates: // The Go toolchain classifies immediates (asm7.go conclass):
// - C_ABCON0 (0 < v ≤ 4095): bitmask first for positive values // - inside the imm12/shifted-imm12 "addcon" band (C_ABCON0/C_ABCON,
// - Negative values: MOVN first, then bitmask // 0 < v ≤ 4095 or a 4096 multiple up to 0xFFF000): bitmask first, so
// - C_MOVCON (movcon-eligible, outside ABCON range): MOVZ/MOVN first // `MOVD $4096, R27` is ORR $4096, not MOVZ $(1<<12)
tryBitmaskFirst := d > 0 && d <= 0xFFF // - outside that band: MOVZ/MOVN first (C_MOVCON before C_BITCON), and
// negative values reach MOVN before the bitmask test
tryBitmaskFirst := d > 0 && (d <= 0xFFF || (d&0xFFF == 0 && d <= 0xFFF000))
if tryBitmaskFirst { if tryBitmaskFirst {
// Small immediate: try bitmask first (Go uses ORR for values like $1, $256). // Addcon-band immediate: try bitmask first (Go uses ORR for values
// like $1, $256 and $65536).
N, immr, imms, ok := arm64Bitmask(uint64(d), int(sf)) N, immr, imms, ok := arm64Bitmask(uint64(d), int(sf))
if ok { if ok {
return a64wordLE(sf<<31 | 1<<29 | 0x24<<23 | N<<22 | immr<<16 | imms<<10 | 31<<5 | uint32(rd)), nil return a64wordLE(sf<<31 | 1<<29 | 0x24<<23 | N<<22 | immr<<16 | imms<<10 | 31<<5 | uint32(rd)), nil
@@ -1354,14 +1358,43 @@ func arm64ExclMem(mnem string, op *ast.Operand) (int, error) {
return rn, nil return rn, nil
} }
// encodeARM64Excl encodes an exclusive load/store instruction. // arm64PairOf parses a register-pair operand `(R1, R2)`, reporting false
// LDXR (Rn), Rt → LDXR Rt, [Rn] (2 operands: mem, reg) // when the operand is not a pair. The toolchain takes the second register of
// STXR Rs, (Rn), Rt → STXR Rs, Rt, [Rn] (3 operands: Rs, mem, Rt-status) // the pair from the operand's Offset (its C_PAIR class,
// cmd/internal/obj/arm64/asm7.go cases 58/59).
func arm64PairOf(op *ast.Operand) (int, int, bool) {
raw := strings.TrimSpace(op.Raw)
if !strings.HasPrefix(raw, "(") || !strings.HasSuffix(raw, ")") {
return -1, -1, false
}
parts := strings.Split(raw[1:len(raw)-1], ",")
if len(parts) != 2 {
return -1, -1, false
}
r1 := arm64RegNum(strings.TrimSpace(parts[0]))
r2 := arm64RegNum(strings.TrimSpace(parts[1]))
if r1 < 0 || r2 < 0 {
return -1, -1, false
}
return r1, r2, true
}
// encodeARM64Excl encodes the exclusive load/store family with the operand
// order the toolchain parses (cmd/internal/obj/arm64/asm7.go cases 58 and 59,
// and its own spellings in arm64enc.s):
//
// STXR Rt, (Rn), Rs store, single register
// STXP (Rt1, Rt2), (Rn), Rs store, register pair
// LDXR (Rn), Rt load, single register
// LDXP (Rn), (Rt1, Rt2) load, register pair
//
// Decoded toolchain evidence: `STXR R1, (R2), R3` assembles to 0xc8037c41,
// whose fields are Rs=3, Rn=2, Rt=1: the FIRST register operand is the data
// register and the LAST the status register.
func encodeARM64Excl(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) { func encodeARM64Excl(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, error) {
// LDXR/STXR have different operand forms.
isLoad := strings.HasPrefix(mnem, "LD") isLoad := strings.HasPrefix(mnem, "LD")
if isLoad { if isLoad {
// LDXR (Rn), Rt → 2 operands: mem, reg // LDXR (Rn), Rt / LDXP (Rn), (Rt1, Rt2): 2 operands.
if len(ops) != 2 { if len(ops) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
} }
@@ -1369,25 +1402,34 @@ func encodeARM64Excl(mnem string, baseOp uint32, ops []*ast.Operand) ([]byte, er
if err != nil { if err != nil {
return nil, err return nil, err
} }
if rt1, rt2, ok := arm64PairOf(ops[1]); ok {
// The single-register opcodes pre-set the unused Rs (bits 20:16)
// and Rt2 (bits 14:10) fields to 31; the pair forms carry a real
// Rt2 and keep Rs at 31.
return a64wordLE(baseOp | 0x1F<<16 | uint32(rt2)<<10 | uint32(rn)<<5 | uint32(rt1)), nil
}
rt := arm64RegNum(operandRegName(ops[1])) rt := arm64RegNum(operandRegName(ops[1]))
if rt < 0 { if rt < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem) return nil, fmt.Errorf("invalid operand in %s", mnem)
} }
return a64wordLE(baseOp | uint32(rn)<<5 | uint32(rt)), nil return a64wordLE(baseOp | uint32(rn)<<5 | uint32(rt)), nil
} }
// STXR Rs, (Rn), Rt → 3 operands: Rs, mem, Rt // STXR Rt, (Rn), Rs / STXP (Rt1, Rt2), (Rn), Rs: 3 operands.
if len(ops) != 3 { if len(ops) != 3 {
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops)) return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
} }
rs := arm64RegNum(operandRegName(ops[0]))
if rs < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem)
}
rn, err := arm64ExclMem(mnem, ops[1]) rn, err := arm64ExclMem(mnem, ops[1])
if err != nil { if err != nil {
return nil, err return nil, err
} }
rt := arm64RegNum(operandRegName(ops[2])) rs := arm64RegNum(operandRegName(ops[2]))
if rs < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem)
}
if rt1, rt2, ok := arm64PairOf(ops[0]); ok {
return a64wordLE(baseOp | uint32(rs)<<16 | uint32(rt2)<<10 | uint32(rn)<<5 | uint32(rt1)), nil
}
rt := arm64RegNum(operandRegName(ops[0]))
if rt < 0 { if rt < 0 {
return nil, fmt.Errorf("invalid operand in %s", mnem) return nil, fmt.Errorf("invalid operand in %s", mnem)
} }
+23 -3
View File
@@ -98,8 +98,12 @@ func arm64RegNum(name string) int {
return 30 return 30
case "R31", "ZR": case "R31", "ZR":
return 31 return 31
case "SP": case "SP", "RSP":
return 31 // SP and ZR share encoding 31; context determines meaning // RSP is the toolchain's spelling for register 31 (it rejects
// R31 in an operand); SP stays for sources that spell it the
// amd64 way. SP and ZR share encoding 31; context determines
// the meaning.
return 31
} }
// F0-F31. // F0-F31.
if len(name) >= 1 && name[0] == 'F' { if len(name) >= 1 && name[0] == 'F' {
@@ -282,7 +286,7 @@ const (
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FCRC32 // CRC32 a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FSIMD3 // SIMD 3-operand: VADD, VSUB, VMUL a64FSIMD3 // SIMD 3-operand: VADD, VSUB, VMUL
) )
@@ -575,6 +579,10 @@ func init() {
} }
// ---- exclusive load/store ---- // ---- exclusive load/store ----
// Single-register forms pre-set the unused Rs and Rt2 fields to 31 (the
// 0x7c00/0x1f0000 halves of the constants below); the register-pair
// forms carry a real Rt2 in bits 14:10, so their opcodes pre-set
// neither field.
a64InstrTable["LDXR"] = a64Enc{format: a64FExcl, op: 0xc85f7c00} a64InstrTable["LDXR"] = a64Enc{format: a64FExcl, op: 0xc85f7c00}
a64InstrTable["LDXRB"] = a64Enc{format: a64FExcl, op: 0x085f7c00} a64InstrTable["LDXRB"] = a64Enc{format: a64FExcl, op: 0x085f7c00}
a64InstrTable["LDXRH"] = a64Enc{format: a64FExcl, op: 0x485f7c00} a64InstrTable["LDXRH"] = a64Enc{format: a64FExcl, op: 0x485f7c00}
@@ -583,6 +591,12 @@ func init() {
a64InstrTable["LDAXRB"] = a64Enc{format: a64FExcl, op: 0x085ffc00} a64InstrTable["LDAXRB"] = a64Enc{format: a64FExcl, op: 0x085ffc00}
a64InstrTable["LDAXRH"] = a64Enc{format: a64FExcl, op: 0x485ffc00} a64InstrTable["LDAXRH"] = a64Enc{format: a64FExcl, op: 0x485ffc00}
a64InstrTable["LDAXRW"] = a64Enc{format: a64FExcl, op: 0x885ffc00} a64InstrTable["LDAXRW"] = a64Enc{format: a64FExcl, op: 0x885ffc00}
// Pair loads, LDSTX(sz, 0, l=1, o1=1, o0) in asm7.go: LDXP/ LDXPW have
// o0=0, LDAXP/LDAXPW o0=1 (bit 15). Rs (bits 20:16) stays 31.
a64InstrTable["LDXP"] = a64Enc{format: a64FExcl, op: 0xc8600000}
a64InstrTable["LDXPW"] = a64Enc{format: a64FExcl, op: 0x88600000}
a64InstrTable["LDAXP"] = a64Enc{format: a64FExcl, op: 0xc8608000}
a64InstrTable["LDAXPW"] = a64Enc{format: a64FExcl, op: 0x88608000}
a64InstrTable["STXR"] = a64Enc{format: a64FExcl, op: 0xc8007c00} a64InstrTable["STXR"] = a64Enc{format: a64FExcl, op: 0xc8007c00}
a64InstrTable["STXRB"] = a64Enc{format: a64FExcl, op: 0x08007c00} a64InstrTable["STXRB"] = a64Enc{format: a64FExcl, op: 0x08007c00}
a64InstrTable["STXRH"] = a64Enc{format: a64FExcl, op: 0x48007c00} a64InstrTable["STXRH"] = a64Enc{format: a64FExcl, op: 0x48007c00}
@@ -591,6 +605,12 @@ func init() {
a64InstrTable["STLXRB"] = a64Enc{format: a64FExcl, op: 0x0800fc00} a64InstrTable["STLXRB"] = a64Enc{format: a64FExcl, op: 0x0800fc00}
a64InstrTable["STLXRH"] = a64Enc{format: a64FExcl, op: 0x4800fc00} a64InstrTable["STLXRH"] = a64Enc{format: a64FExcl, op: 0x4800fc00}
a64InstrTable["STLXRW"] = a64Enc{format: a64FExcl, op: 0x8800fc00} a64InstrTable["STLXRW"] = a64Enc{format: a64FExcl, op: 0x8800fc00}
// Pair stores, LDSTX(sz, 0, l=0, o1=1, o0): STXP/STXPW have o0=0,
// STLXP/STLXPW o0=1 (bit 15). Both Rs and Rt2 are real fields.
a64InstrTable["STXP"] = a64Enc{format: a64FExcl, op: 0xc8200000}
a64InstrTable["STXPW"] = a64Enc{format: a64FExcl, op: 0x88200000}
a64InstrTable["STLXP"] = a64Enc{format: a64FExcl, op: 0xc8208000}
a64InstrTable["STLXPW"] = a64Enc{format: a64FExcl, op: 0x88208000}
// ---- LSE atomics ---- // ---- LSE atomics ----
a64InstrTable["LDADDD"] = a64Enc{format: a64FLSE, op: 3<<30 | 0x1c1<<21 | 0x00<<10} a64InstrTable["LDADDD"] = a64Enc{format: a64FLSE, op: 3<<30 | 0x1c1<<21 | 0x00<<10}
+74 -6
View File
@@ -820,15 +820,28 @@ func TestArm64ExclOffsetErrors(t *testing.T) {
} }
} }
// TestArm64ExclNoOffset pins the plain (Rn) forms. gasm parses the store // TestArm64ExclNoOffset pins the plain (Rn) forms, byte-for-byte against
// with the status register first (ARM ARM order); go tool asm parses the // go tool asm. The toolchain parses the FIRST register of a store as the
// same text with the data register first, so the two spellings differ and // data register and the LAST as the status register (asm7.go case 59), and
// the store word below is gasm's own. // the pair forms as (Rt1, Rt2) (case 58/59):
//
// STXR R3, (R1), R4 → c8047c23 (Rt=3, Rn=1, Rs=4)
// STXP (R3, R4), (R1), R5 → c8251023 (Rt=3, Rt2=4, Rn=1, Rs=5)
// LDXP (R1), (R3, R4) → c87f1023 (Rn=1, Rt=3, Rt2=4)
func TestArm64ExclNoOffset(t *testing.T) { func TestArm64ExclNoOffset(t *testing.T) {
got := arm64Words(t, "\tLDXR (R1), R2\n\tSTXR R3, (R1), R4\n") got := arm64Words(t, "\tLDXR (R1), R2\n\tSTXR R3, (R1), R4\n"+
"\tSTXP (R3, R4), (R1), R5\n\tSTXPW (R3, R4), (R1), R5\n"+
"\tLDXP (R1), (R3, R4)\n\tLDXPW (R1), (R3, R4)\n"+
"\tSTXR R3, (RSP), R4\n\tLDXR (RSP), R2\n")
want := []uint32{ want := []uint32{
0xc85f7c22, // LDXR X2, [X1] 0xc85f7c22, // LDXR X2, [X1]
0xc8037c24, // STXR W3, X4, [X1] with Rs = R3, Rt = R4 0xc8047c23, // STXR W3, [X1], W4 with Rt = R3, Rs = R4
0xc8251023, // STXP (R3, R4), [X1], R5
0x88251023, // STXPW (R3, R4), [X1], R5
0xc87f1023, // LDXP [X1], (R3, R4)
0x887f1023, // LDXPW [X1], (R3, R4)
0xc8047fe3, // STXR R3, [SP], R4
0xc85f7fe2, // LDXR [SP], R2
0xd65f03c0, // RET 0xd65f03c0, // RET
} }
for i := range want { for i := range want {
@@ -920,3 +933,58 @@ func TestArm64LargeFrameSpadj(t *testing.T) {
t.Errorf("final RET word at byte 60 = %08x, want d65f03c0", got) t.Errorf("final RET word at byte 60 = %08x, want d65f03c0", got)
} }
} }
// TestArm64SplitFrameSpadj pins the addcon2 band, where neither imm12 form
// nor a single MOVZ carries the autosize and the toolchain splits the
// prologue SUB into two imm12 instructions (asm7.go case 48) while the
// non-leaf RET still materialises the value into REGTMP (obj7.go ARET,
// issue 73259). $65664 rounds the autosize to 65680 = 144 + 16<<12:
//
// [SUB $144, RSP, R20][SUB $(16<<12), R20, R20][STP][MOVD R20, SP][SUB $8]
// [CALL]
// [LDP][MOVD $144, R27][MOVK $(1<<16), R27][ADD R27, RSP, RSP][RET]
//
// SP moves at the fourth word (byte 12) and returns to zero at the final
// RET (byte 40); the words are go tool asm's own for the same source.
func TestArm64SplitFrameSpadj(t *testing.T) {
f, errs := parser.Parse("frame_arm64.s", "#include \"textflag.h\"\n\nTEXT ·framed(SB), NOSPLIT, $65664-0\n\tCALL ·other(SB)\n\tRET\n\nTEXT ·other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
fn := img.Funcs[0]
wantSpadj := []SpadjStep{{PC: 12, Value: 65680}, {PC: 40, Value: 0}}
if len(fn.Spadj) != len(wantSpadj) {
t.Fatalf("spadj = %v, want %v", fn.Spadj, wantSpadj)
}
for i := range wantSpadj {
if fn.Spadj[i] != wantSpadj[i] {
t.Errorf("spadj[%d] = %v, want %v", i, fn.Spadj[i], wantSpadj[i])
}
}
want := []uint32{
0xd10243f4, // SUB $144, RSP, R20
0xd1404294, // SUB $(16<<12), R20, R20
0xa93ffa9d, // STP (R29, R30), -8(R20)
0x9100029f, // MOVD R20, RSP
0xd10023fd, // SUB $8, RSP, R29
0x94000000, // CALL (relocation masked at link time)
0xa97ffbfd, // LDP -8(RSP), (R29, R30)
0xd280121b, // MOVD $144, R27
0xf2a0003b, // MOVK $(1<<16), R27
0x8b3b63ff, // ADD R27, RSP, RSP
0xd65f03c0, // RET
}
words := leWords(img.Code[fn.Offset : fn.Offset+fn.Size])
if len(words) != len(want) {
t.Fatalf("framed = %d words, want %d", len(words), len(want))
}
for i, w := range want {
if words[i] != w {
t.Errorf("word %d = %08x, want %08x", i, words[i], w)
}
}
}
+75 -21
View File
@@ -187,10 +187,32 @@ func arm64Prologue(fi arm64FrameInfo) []byte {
return a64WordsLE(ws...) return a64WordsLE(ws...)
} }
// arm64SubImmWords emits SUB $imm, SP, Rd: the immediate form when the value // arm64SplitImm12 reports whether the toolchain decomposes ADD/SUB $imm into
// fits the imm12 field (plain, or shifted left by 12 when it is a multiple // two imm12 instructions instead of materialising it into REGTMP
// of 4096); otherwise the toolchain materialises it into REGTMP (R27) and // (asm7.go case 48, the C_ADDCON2 class): the value must fit 24 bits
// subtracts the register in the extended-register form. // unsigned and be neither encodable as one imm12 (checked by the callers
// first), nor loadable into a register in a single MOVZ/MOVN word, nor a
// logical immediate, because conclass tests all three before C_ADDCON2.
func arm64SplitImm12(imm uint32) bool {
if imm > 0xFFFFFF {
return false
}
if _, _, _, ok := arm64Bitmask(uint64(imm), 1); ok {
return false
}
return arm64Movcon(int64(imm)) < 0 && arm64Movcon(^int64(imm)) < 0
}
// arm64SubImmWords emits SUB $imm, SP, Rd with the toolchain's ladder for an
// ADD/SUB constant (asm7.go conclass and cases 2, 48, 62 and 13): the
// immediate form when the value fits imm12 (plain, or shifted left by 12
// when it is a multiple of 4096); a value with a single 16-bit chunk, a
// logical immediate, or one wider than 24 bits is materialised into REGTMP
// (R27) and subtracted in the extended-register form; everything else up to
// 0xFFFFFF is split into two imm12 instructions:
//
// SUB $(imm&0xfff), SP, Rd
// SUB $((imm&0xfff000)>>12)<<12, Rd, Rd
func arm64SubImmWords(imm uint32, rd uint32) []uint32 { func arm64SubImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 0xFFF { if imm <= 0xFFF {
return []uint32{a64AddSub(1, 1, 0, 0, imm, 31, rd)} return []uint32{a64AddSub(1, 1, 0, 0, imm, 31, rd)}
@@ -198,15 +220,21 @@ func arm64SubImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 4095<<12 && imm&0xFFF == 0 { if imm <= 4095<<12 && imm&0xFFF == 0 {
return []uint32{a64AddSub(1, 1, 0, 1, imm>>12, 31, rd)} return []uint32{a64AddSub(1, 1, 0, 1, imm>>12, 31, rd)}
} }
mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD") if !arm64SplitImm12(imm) {
if err != nil { mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD")
mov = nil if err != nil {
mov = nil
}
return append(wordsOf(mov), arm64DPExtWords(arm64OpSub, 27, 31, rd))
}
return []uint32{
a64AddSub(1, 1, 0, 0, imm&0xFFF, 31, rd),
a64AddSub(1, 1, 0, 1, (imm&0xFFF000)>>12, rd, rd),
} }
return append(wordsOf(mov), arm64DPExtWords(arm64OpSub, 27, 31, rd))
} }
// arm64AddImmWords emits ADD $imm, SP, Rd with the same imm12, shifted-imm12 // arm64AddImmWords emits ADD $imm, SP, Rd with the same imm12, shifted-imm12,
// and REGTMP fallback ladder. // split and REGTMP ladder as arm64SubImmWords.
func arm64AddImmWords(imm uint32, rd uint32) []uint32 { func arm64AddImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 0xFFF { if imm <= 0xFFF {
return []uint32{a64AddSub(1, 0, 0, 0, imm, 31, rd)} return []uint32{a64AddSub(1, 0, 0, 0, imm, 31, rd)}
@@ -214,11 +242,35 @@ func arm64AddImmWords(imm uint32, rd uint32) []uint32 {
if imm <= 4095<<12 && imm&0xFFF == 0 { if imm <= 4095<<12 && imm&0xFFF == 0 {
return []uint32{a64AddSub(1, 0, 0, 1, imm>>12, 31, rd)} return []uint32{a64AddSub(1, 0, 0, 1, imm>>12, 31, rd)}
} }
mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD") if !arm64SplitImm12(imm) {
mov, err := encodeARM64LoadImm(27, int64(imm), "MOVD")
if err != nil {
mov = nil
}
return append(wordsOf(mov), arm64DPExtWords(arm64OpAdd, 27, 31, rd))
}
return []uint32{
a64AddSub(1, 0, 0, 0, imm&0xFFF, 31, rd),
a64AddSub(1, 0, 0, 1, (imm&0xFFF000)>>12, rd, rd),
}
}
// arm64RetAddWords emits the frame deallocation of a non-leaf RET with a
// large frame. The toolchain adds the frame back with a single instruction:
// a plain imm12 ADD when autosize fits 12 bits, otherwise the value is
// materialised into REGTMP and added as a register, so the epilogue never
// leaves a partially deallocated frame (obj7.go ARET, issue 73259). The
// shifted-imm12 and split-imm12 forms are therefore never used here, unlike
// the leaf epilogue's plain ADD instructions.
func arm64RetAddWords(autosize uint32) []uint32 {
if autosize < 1<<12 {
return []uint32{a64AddSub(1, 0, 0, 0, autosize, 31, 31)}
}
mov, err := encodeARM64LoadImm(27, int64(autosize), "MOVD")
if err != nil { if err != nil {
mov = nil mov = nil
} }
return append(wordsOf(mov), arm64DPExtWords(arm64OpAdd, 27, 31, rd)) return append(wordsOf(mov), arm64DPExtWords(arm64OpAdd, 27, 31, 31))
} }
// arm64Return returns the bytes for a RET: the epilogue (restore FP/LR and // arm64Return returns the bytes for a RET: the epilogue (restore FP/LR and
@@ -237,11 +289,11 @@ func arm64Return(fi arm64FrameInfo) []byte {
arm64PostLoad(3, 0, int32(fi.autosize), 31, 30), // LDR.P LR, [SP], #autosize arm64PostLoad(3, 0, int32(fi.autosize), 31, 30), // LDR.P LR, [SP], #autosize
) )
} else { } else {
// Large frame: LDP -8(SP), (FP, LR); ADD $autosize, SP, SP // Large frame: LDP -8(SP), (FP, LR), then deallocate.
ws = append(ws, ws = append(ws,
a64LSP(2, 0, 1, -1, 30, 31, 29), // LDP FP, LR, [SP, #-8] (opc=2 for 64-bit pair) a64LSP(2, 0, 1, -1, 30, 31, 29), // LDP FP, LR, [SP, #-8] (opc=2 for 64-bit pair)
) )
ws = append(ws, arm64AddImmWords(uint32(fi.autosize), 31)...) ws = append(ws, arm64RetAddWords(uint32(fi.autosize))...)
} }
} }
// RET: BR LR (0xd65f03c0) // RET: BR LR (0xd65f03c0)
@@ -260,15 +312,16 @@ func arm64PrologueSpadjPC(fi arm64FrameInfo) int {
} }
// Large frame: [SUB words][STP][ADD R20, SP]; SP moves at the ADD, whose // Large frame: [SUB words][STP][ADD R20, SP]; SP moves at the ADD, whose
// position depends on how many words the SUB itself took (immediate, // position depends on how many words the SUB itself took (immediate,
// shifted immediate, or a materialised REGTMP sequence). // shifted immediate, the two-word imm12 split, or a materialised REGTMP
// sequence).
return 4 * (len(arm64SubImmWords(uint32(fi.autosize), 20)) + 1) return 4 * (len(arm64SubImmWords(uint32(fi.autosize), 20)) + 1)
} }
// arm64ReturnEpilogueLen returns the byte length of the RET's epilogue up to // arm64ReturnEpilogueLen returns the byte length of the RET's epilogue up to
// (but not including) the final RET instruction. The ADD sequences share the // (but not including) the final RET instruction. The lengths are read from
// prologue's immediate ladder, so their length is read from the same helper // the same word-emitting helpers the epilogue uses rather than assumed: the
// rather than assumed: a materialised autosize costs its MOV words plus the // leaf path shares the prologue's immediate ladder, and a materialised
// ADD itself. // autosize costs its MOV words plus the ADD itself.
func arm64ReturnEpilogueLen(fi arm64FrameInfo) int { func arm64ReturnEpilogueLen(fi arm64FrameInfo) int {
if fi.autosize == 0 { if fi.autosize == 0 {
return 0 return 0
@@ -280,8 +333,9 @@ func arm64ReturnEpilogueLen(fi arm64FrameInfo) int {
if fi.autosize <= 0xf0 { if fi.autosize <= 0xf0 {
return 8 // LDR + LDR.P return 8 // LDR + LDR.P
} }
// LDP + the ADD ladder that deallocates the frame. // LDP + the deallocation emitted by arm64RetAddWords, so the length
return 4 + 4*len(arm64AddImmWords(uint32(fi.autosize), 31)) // tracks whatever the MOVD ladder needs.
return 4 + 4*len(arm64RetAddWords(uint32(fi.autosize)))
} }
// arm64ResolvePseudo translates a pseudo-register memory reference into a // arm64ResolvePseudo translates a pseudo-register memory reference into a
+7
View File
@@ -86,9 +86,16 @@ func Encodable(mnemonic string) bool {
"BSWAP", "BSWAP",
"PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2", "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2",
"MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX", "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX",
"MOVBWZX", "MOVBWSX", "MOVBLSX", "MOVBQSX", "MOVWQSX", "MOVLQZX",
"CVTSL2SD", "CVTSQ2SD", "CVTSL2SD", "CVTSQ2SD",
"MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS": "MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
return true return true
} }
// Full-name dispatches the size split would eat (a trailing width
// letter that is part of the mnemonic).
switch upper {
case "PMOVMSKB":
return true
}
return false return false
} }
+7 -1
View File
@@ -101,6 +101,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
if m, ok := sseBinTable[base]; ok { if m, ok := sseBinTable[base]; ok {
return e.encodeSSEBin(m, ops) return e.encodeSSEBin(m, ops)
} }
// PMOVMSKB ends in a width letter the size split would eat, so it
// dispatches on the full name like the packed binaries above.
if upper == "PMOVMSKB" {
return e.encodePmovmskb(upper, ops)
}
switch base { switch base {
case "MOV": case "MOV":
return e.encodeMov(ops, size) return e.encodeMov(ops, size)
@@ -126,7 +131,8 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeBswap(ops, size) return e.encodeBswap(ops, size)
case "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2": case "PREFETCHNTA", "PREFETCHT0", "PREFETCHT1", "PREFETCHT2":
return e.encodePrefetch(base, ops) return e.encodePrefetch(base, ops)
case "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX": case "MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX",
"MOVBWZX", "MOVBWSX", "MOVBLSX", "MOVBQSX", "MOVWQSX", "MOVLQZX":
return e.encodeMovExtend(base, ops) return e.encodeMovExtend(base, ops)
case "CVTSL2SD", "CVTSQ2SD": case "CVTSL2SD", "CVTSQ2SD":
return e.encodeCvtsi2sd(base == "CVTSQ2SD", ops) return e.encodeCvtsi2sd(base == "CVTSQ2SD", ops)
+12
View File
@@ -357,6 +357,18 @@ func TestScalarGroundTruth(t *testing.T) {
{"MOVBQZX AL,R8", "MOVBQZX", []Operand{AL, r8}, "4c0fb6c0", "MOVZX"}, {"MOVBQZX AL,R8", "MOVBQZX", []Operand{AL, r8}, "4c0fb6c0", "MOVZX"},
{"MOVWLZX AX,CX", "MOVWLZX", []Operand{AX, CX}, "0fb7c8", "MOVZX"}, {"MOVWLZX AX,CX", "MOVWLZX", []Operand{AX, CX}, "0fb7c8", "MOVZX"},
{"MOVWQZX AX,R8", "MOVWQZX", []Operand{AX, r8}, "4c0fb7c0", "MOVZX"}, {"MOVWQZX AX,R8", "MOVWQZX", []Operand{AX, r8}, "4c0fb7c0", "MOVZX"},
// The width pairs the toolchain accepts and GOROOT uses; bytes
// pinned from go tool asm (see testdata/verify/widen_amd64.s).
{"MOVBWZX (BX),R11W", "MOVBWZX", []Operand{Ptr(BX, 0, 1), Reg{idx: 11, size: 2}}, "66440fb61b", "MOVZX"},
{"MOVBWSX (BX),R11W", "MOVBWSX", []Operand{Ptr(BX, 0, 1), Reg{idx: 11, size: 2}}, "66440fbe1b", "MOVSX"},
{"MOVBLSX (BX),AX", "MOVBLSX", []Operand{Ptr(BX, 0, 1), AX}, "0fbe03", "MOVSX"},
{"MOVBQSX (BX),R8", "MOVBQSX", []Operand{Ptr(BX, 0, 1), r8}, "4c0fbe03", "MOVSX"},
{"MOVWQSX (BX),R9", "MOVWQSX", []Operand{Ptr(BX, 0, 2), r9}, "4c0fbf0b", "MOVSX"},
// A long to quad zero-extend is a plain 32-bit move.
{"MOVLQZX (BX),DX", "MOVLQZX", []Operand{Ptr(BX, 0, 4), DX}, "8b13", "MOV"},
{"MOVLQZX AX,DX", "MOVLQZX", []Operand{AX, DX}, "8bd0", "MOV"},
{"PMOVMSKB X1,AX", "PMOVMSKB", []Operand{vreg(t, "X1"), AX}, "660fd7c1", "PMOVMSKB"},
{"PMOVMSKB X11,CX", "PMOVMSKB", []Operand{vreg(t, "X11"), CX}, "66410fd7cb", "PMOVMSKB"},
{"CVTSL2SD R8,X13", "CVTSL2SD", []Operand{r8, vreg(t, "X13")}, "f2450f2ae8", "CVTSI2SD"}, {"CVTSL2SD R8,X13", "CVTSL2SD", []Operand{r8, vreg(t, "X13")}, "f2450f2ae8", "CVTSI2SD"},
{"CVTSL2SD AX,X0", "CVTSL2SD", []Operand{AX, vreg(t, "X0")}, "f20f2ac0", "CVTSI2SD"}, {"CVTSL2SD AX,X0", "CVTSL2SD", []Operand{AX, vreg(t, "X0")}, "f20f2ac0", "CVTSI2SD"},
{"CVTSQ2SD R8,X13", "CVTSQ2SD", []Operand{r8, vreg(t, "X13")}, "f24d0f2ae8", "CVTSI2SD"}, {"CVTSQ2SD R8,X13", "CVTSQ2SD", []Operand{r8, vreg(t, "X13")}, "f24d0f2ae8", "CVTSI2SD"},
+40 -13
View File
@@ -838,15 +838,24 @@ func (e *enc) encodeBswap(ops []Operand, size int) error {
// width. The source is narrower than the destination, so the plain size-suffix // width. The source is narrower than the destination, so the plain size-suffix
// convention does not apply to these names. // convention does not apply to these names.
var movExtendOp = map[string]struct { var movExtendOp = map[string]struct {
op []byte op []byte
dst64 bool dstSize int
}{ }{
"MOVBLZX": {[]byte{0x0F, 0xB6}, false}, // byte → long, zero-extend "MOVBLZX": {[]byte{0x0F, 0xB6}, 4}, // byte → long, zero-extend
"MOVBQZX": {[]byte{0x0F, 0xB6}, true}, // byte → quad, zero-extend "MOVBQZX": {[]byte{0x0F, 0xB6}, 8}, // byte → quad, zero-extend
"MOVWLZX": {[]byte{0x0F, 0xB7}, false}, // word → long, zero-extend "MOVWLZX": {[]byte{0x0F, 0xB7}, 4}, // word → long, zero-extend
"MOVWQZX": {[]byte{0x0F, 0xB7}, true}, // word → quad, zero-extend "MOVWQZX": {[]byte{0x0F, 0xB7}, 8}, // word → quad, zero-extend
"MOVWLSX": {[]byte{0x0F, 0xBF}, false}, // word → long, sign-extend "MOVWLSX": {[]byte{0x0F, 0xBF}, 4}, // word → long, sign-extend
"MOVLQSX": {[]byte{0x63}, true}, // long → quad, sign-extend (MOVSXD) "MOVLQSX": {[]byte{0x63}, 8}, // long → quad, sign-extend (MOVSXD)
"MOVBWZX": {[]byte{0x0F, 0xB6}, 2}, // byte → word, zero-extend
"MOVBWSX": {[]byte{0x0F, 0xBE}, 2}, // byte → word, sign-extend
"MOVBLSX": {[]byte{0x0F, 0xBE}, 4}, // byte → long, sign-extend
"MOVBQSX": {[]byte{0x0F, 0xBE}, 8}, // byte → quad, sign-extend
"MOVWQSX": {[]byte{0x0F, 0xBF}, 8}, // word → quad, sign-extend
// A long → quad zero-extend is a plain 32-bit move: every 32-bit
// operation zero-extends its result into the full register, so the
// toolchain lowers MOVLQZX to the plain MOVL encoding.
"MOVLQZX": {[]byte{0x8B}, 4},
} }
// encodeMovExtend encodes a mixed-width extending move: reg = dst (the wider // encodeMovExtend encodes a mixed-width extending move: reg = dst (the wider
@@ -860,12 +869,30 @@ func (e *enc) encodeMovExtend(base string, ops []Operand) error {
if !ok { if !ok {
return fmt.Errorf("%s destination must be a register", base) return fmt.Errorf("%s destination must be a register", base)
} }
size := 4 i := newInstr(spec.dstSize, spec.op)
if spec.dst64 { if err := setRM(i, dstReg, ops[0], spec.dstSize); err != nil {
size = 8 return err
} }
i := newInstr(size, spec.op) return e.emit(i)
if err := setRM(i, dstReg, ops[0], size); err != nil { }
// encodePmovmskb encodes PMOVMSKB, the legacy SSE2 byte mask extract: the
// XMM source's sign bytes pack into a GP destination, 66 0F D7 /r.
func (e *enc) encodePmovmskb(base string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", base, len(ops))
}
srcReg, srcVec := vecReg(ops[0])
if !srcVec {
return fmt.Errorf("%s source must be an XMM register", base)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("%s destination must be a register", base)
}
i := newInstr(4, []byte{0x0F, 0xD7})
i.prefix = 0x66
if err := setRM(i, dstReg, srcReg, 4); err != nil {
return err return err
} }
return e.emit(i) return e.emit(i)
+1 -1
View File
@@ -65,7 +65,7 @@ func riscvRegNum(name string) int {
return 25 return 25
case "X26", "S10": case "X26", "S10":
return 26 return 26
case "X27", "S11": case "X27", "S11", "g":
return 27 return 27
case "X28", "T3": case "X28", "T3":
return 28 return 28
+39 -13
View File
@@ -938,11 +938,40 @@ func compareGroundTruth(img *asm.Image, gt map[string][]byte) (matched, total, d
goCmp[j] = 0 goCmp[j] = 0
} }
} }
if bytes.Equal(gasmCmp, goCmp) { // The toolchain pads text symbols to 16-byte boundaries with
// zeros, so a function whose size is not a multiple of 16
// carries trailing zeros in the ground truth that are not part
// of the encoding. Compare up to the shorter side and require
// the remainder of whichever is longer to be zero, so padding
// never masks a real difference.
cmpLen := min(len(gasmCmp), len(goCmp))
equal := bytes.Equal(gasmCmp[:cmpLen], goCmp[:cmpLen])
if equal {
for _, b := range gasmCmp[cmpLen:] {
if b != 0 {
equal = false
break
}
}
}
if equal {
for _, b := range goCmp[cmpLen:] {
if b != 0 {
equal = false
break
}
}
}
if equal {
matched++ matched++
if len(fn.Relocs) > 0 { switch {
case len(fn.Relocs) > 0 && len(goCmp) > cmpLen:
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked, %d padding)\n", fn.Name, fn.Size, len(fn.Relocs), len(goCmp)-cmpLen)
case len(fn.Relocs) > 0:
fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs)) fmt.Printf(" %s: MATCH (%d bytes, %d relocs masked)\n", fn.Name, fn.Size, len(fn.Relocs))
} else { case len(goCmp) > cmpLen:
fmt.Printf(" %s: MATCH (%d bytes, %d padding)\n", fn.Name, fn.Size, len(goCmp)-cmpLen)
default:
fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size) fmt.Printf(" %s: MATCH (%d bytes)\n", fn.Name, fn.Size)
} }
} else { } else {
@@ -990,8 +1019,8 @@ that tolerate nil pointers and zero lengths in their arguments.
With -abi, each function is called with sentinel values in the registers With -abi, each function is called with sentinel values in the registers
the Go ABI fixes across calls (the frame pointer and the goroutine the Go ABI fixes across calls (the frame pointer and the goroutine
pointer) plus a canary below SP; violations are reported. JIT-based pointer) plus a canary below SP; violations are reported. JIT-based
checks run when the host matches the file's architecture (all but checks run when the host matches the file's architecture, on all four
loong64, which is ground-truth only for now). architectures.
With -fuzz, each function with a // func signature is differentially fuzzed With -fuzz, each function with a // func signature is differentially fuzzed
against the go-tool-asm version in a subprocess (so a crash on a partial against the go-tool-asm version in a subprocess (so a crash on a partial
@@ -1033,16 +1062,13 @@ each entry reproduces.
path := set.Arg(0) path := set.Arg(0)
targetArch := arch.FromFilename(path) targetArch := arch.FromFilename(path)
// JIT execution runs when the host CPU matches the kernel's // JIT execution runs when the host CPU matches the kernel's
// architecture, except loong64: its trampoline is implemented but not // architecture; every trampoline is validated end to end under
// yet validated against real hardware (the Go runtime cannot start // qemu-user emulation (the loong64 one included, via the raw-address
// under the available loong64 emulators), so those kernels take the // leave handoff).
// toolchain-comparison path. if targetArch != hostArch() {
if targetArch != hostArch() || targetArch == arch.LOONG64 {
// No JIT on this host: ground truth and profile remain available for // No JIT on this host: ground truth and profile remain available for
// every architecture, because cmdVerifyNonJIT assembles and compares // every architecture, because cmdVerifyNonJIT assembles and compares
// against the toolchain without executing anything. (loong64 is // against the toolchain without executing anything.
// ground-truth-only everywhere for now: its trampoline is implemented
// but not yet validated against real hardware.)
switch targetArch { switch targetArch {
case arch.AMD64, arch.RISCV, arch.LOONG64, arch.ARM64: case arch.AMD64, arch.RISCV, arch.LOONG64, arch.ARM64:
return cmdVerifyNonJIT(path, targetArch, *groundTruth, *profile) return cmdVerifyNonJIT(path, targetArch, *groundTruth, *profile)
+20
View File
@@ -15,6 +15,7 @@ import (
"testing" "testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch" "sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
) )
const clean = "#include \"textflag.h\"\n" + const clean = "#include \"textflag.h\"\n" +
@@ -440,3 +441,22 @@ func TestRunCorpusAudit(t *testing.T) {
t.Errorf("arm64 unencodable reasons = %d, want 1", r) t.Errorf("arm64 unencodable reasons = %d, want 1", r)
} }
} }
// TestCompareGroundTruthPadding pins the padding-aware ground-truth
// comparison: the toolchain pads text symbols to 16-byte boundaries, so
// trailing zeros in the reference must not read as a mismatch, while any
// non-zero tail still must.
func TestCompareGroundTruthPadding(t *testing.T) {
code := []byte{0x48, 0x8b, 0x07, 0xc3} // 4 bytes, not a multiple of 16
img := &asm.Image{Code: code, Funcs: []asm.FuncLayout{{Name: "f", Offset: 0, Size: len(code)}}}
padded := append(append([]byte(nil), code...), 0, 0, 0)
matched, total, diffs := compareGroundTruth(img, map[string][]byte{"f": padded})
if matched != 1 || total != 1 || diffs != 0 {
t.Fatalf("zero padding should match: matched=%d total=%d diffs=%d", matched, total, diffs)
}
dirty := append(append([]byte(nil), code...), 0, 0x90, 0)
matched, _, diffs = compareGroundTruth(img, map[string][]byte{"f": dirty})
if matched != 0 || diffs != 1 {
t.Fatalf("non-zero padding must mismatch: matched=%d diffs=%d", matched, diffs)
}
}
+3 -3
View File
@@ -410,9 +410,9 @@ every architecture too: `enterJITChecked` plants sentinels in the registers
the Go ABI fixes across calls (amd64 `BP`/`R14`, arm64 `R29`/`R28`, riscv64 the Go ABI fixes across calls (amd64 `BP`/`R14`, arm64 `R29`/`R28`, riscv64
`X27`, loong64 `R22`; the latter two keep no hardware frame pointer) and the `X27`, loong64 `R22`; the latter two keep no hardware frame pointer) and the
raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the raw return trampoline `leaveJITCheckedRaw` verifies them, restoring the
saved registers before Go code resumes. riscv64 is validated end to saved registers before Go code resumes. All three non-amd64 trampolines
end under qemu-user emulation; arm64 shares the same stack convention and are validated end to end under qemu-user emulation, the loong64 one
fix; loong64 stays ground-truth-only until hardware validation. through its raw-address leave handoff.
`gasm verify` runs the JIT checks when the host `gasm verify` runs the JIT checks when the host
matches the kernel's architecture and the toolchain comparisons matches the kernel's architecture and the toolchain comparisons
elsewhere. elsewhere.
+1 -2
View File
@@ -219,8 +219,7 @@ The JIT checks run when the host matches the file's architecture; the
toolchain comparison works everywhere. `--fuzz`, `--smoke` and `--abi` run each toolchain comparison works everywhere. `--fuzz`, `--smoke` and `--abi` run each
function in its own child process, so a partial function that faults on random function in its own child process, so a partial function that faults on random
input is reported as `CRASH` instead of ending the sweep; `--call` with `--buf` input is reported as `CRASH` instead of ending the sweep; `--call` with `--buf`
invokes such a function with valid data. loong64 stays on the ground-truth path invokes such a function with valid data.
until hardware validation.
```sh ```sh
gasm verify --ground-truth hello_amd64.s gasm verify --ground-truth hello_amd64.s
+1 -2
View File
@@ -20,8 +20,7 @@ With
each function is called with sentinel values in the registers the Go each function is called with sentinel values in the registers the Go
ABI fixes across calls (the frame pointer and the goroutine pointer) ABI fixes across calls (the frame pointer and the goroutine pointer)
plus a canary below SP; violations are reported. JIT-based checks run plus a canary below SP; violations are reported. JIT-based checks run
when the host matches the file's architecture (all but loong64, which when the host matches the file's architecture, on all four architectures.
is ground-truth only for now).
.PP .PP
With With
.BR \-fuzz , .BR \-fuzz ,
+2 -2
View File
@@ -22,9 +22,9 @@ TEXT ·dirtyFP(SB), NOSPLIT, $0-16
RET RET
// func dirtyG(a int64) int64 // func dirtyG(a int64) int64
// Deliberately clobbers R28, the goroutine pointer (a serious ABI violation). // Deliberately clobbers g, the goroutine pointer (R28; a serious ABI violation).
TEXT ·dirtyG(SB), NOSPLIT, $0-16 TEXT ·dirtyG(SB), NOSPLIT, $0-16
MOVD $0x5678, R28 MOVD $0x5678, g
MOVD a+0(FP), R0 MOVD a+0(FP), R0
MOVD R0, ret+8(FP) MOVD R0, ret+8(FP)
RET RET
+2 -2
View File
@@ -13,9 +13,9 @@ TEXT ·cleanAdd(SB), NOSPLIT, $0-24
RET RET
// func dirtyG(a int64) int64 // func dirtyG(a int64) int64
// Deliberately clobbers R22, the goroutine pointer (a serious ABI violation). // Deliberately clobbers g, the goroutine pointer (R22; a serious ABI violation).
TEXT ·dirtyG(SB), NOSPLIT, $0-16 TEXT ·dirtyG(SB), NOSPLIT, $0-16
MOVV $0x5678, R22 MOVV $0x5678, g
MOVV a+0(FP), R4 MOVV a+0(FP), R4
MOVV R4, ret+8(FP) MOVV R4, ret+8(FP)
RET RET
+2 -2
View File
@@ -13,9 +13,9 @@ TEXT ·cleanAdd(SB), NOSPLIT, $0-24
RET RET
// func dirtyG(a int64) int64 // func dirtyG(a int64) int64
// Deliberately clobbers X27, the goroutine pointer (a serious ABI violation). // Deliberately clobbers g, the goroutine pointer (X27; a serious ABI violation).
TEXT ·dirtyG(SB), NOSPLIT, $0-16 TEXT ·dirtyG(SB), NOSPLIT, $0-16
MOV $0x5678, X27 MOV $0x5678, g
MOV a+0(FP), X5 MOV a+0(FP), X5
MOV X5, ret+8(FP) MOV X5, ret+8(FP)
RET RET
+8 -2
View File
@@ -37,7 +37,10 @@ TEXT ·imm(SB), NOSPLIT, $0-0
MOVV $0x12345, R17 MOVV $0x12345, R17
RET RET
// branch exercises conditional and unconditional control flow. // branch exercises conditional and unconditional control flow. Every
// path must terminate: the smoke harness calls functions with a zeroed
// argument block, and a $0-0 function's registers carry whatever the
// caller left, so a branch maze can reach any label.
TEXT ·branch(SB), NOSPLIT, $0-0 TEXT ·branch(SB), NOSPLIT, $0-0
BEQ R4, R5, done BEQ R4, R5, done
BNE R6, R7, skip BNE R6, R7, skip
@@ -50,7 +53,10 @@ skip:
JMP loop JMP loop
loop: loop:
JAL skip JAL fin
RET
fin:
RET RET
done: done:
+69
View File
@@ -0,0 +1,69 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Large frames across every immediate band of the prologue SUB and the RET
// epilogue, byte-parity-checked against go tool asm:
//
// $5000 autosize 5024 one 16-bit chunk, materialised into REGTMP
// $65664 autosize 65680 split into two imm12 instructions
// $70000 autosize 70016 split into two imm12 instructions
// $65520 autosize 65536 one shifted imm12 in the prologue, a logical
// immediate (ORR) in the non-leaf epilogue
// $16777232 autosize 16777248 wider than 24 bits, MOVZ/MOVK into REGTMP
#include "textflag.h"
TEXT ·leaf5000(SB), NOSPLIT, $5000-0
MOVD R0, R1
MOVD R1, R2
RET
TEXT ·leaf65664(SB), NOSPLIT, $65664-0
MOVD R0, R1
MOVD R1, R2
RET
TEXT ·leaf70000(SB), NOSPLIT, $70000-0
MOVD R0, R1
MOVD R1, R2
RET
TEXT ·leaf65520(SB), NOSPLIT, $65520-0
MOVD R0, R1
MOVD R1, R2
MOVD R2, R3
MOVD R3, R4
RET
TEXT ·nl5000(SB), NOSPLIT, $5000-0
MOVD R0, R1
MOVD R1, R2
CALL ·other(SB)
RET
TEXT ·nl65664(SB), NOSPLIT, $65664-0
MOVD R0, R1
CALL ·other(SB)
RET
TEXT ·nl70000(SB), NOSPLIT, $70000-0
MOVD R0, R1
CALL ·other(SB)
RET
TEXT ·nl65520(SB), NOSPLIT, $65520-0
MOVD R0, R1
MOVD R1, R2
MOVD R2, R3
CALL ·other(SB)
RET
TEXT ·nlhuge(SB), NOSPLIT, $16777232-0
CALL ·other(SB)
RET
TEXT ·other(SB), NOSPLIT, $0-0
MOVD R0, R1
MOVD R1, R2
MOVD R2, R3
RET
+81
View File
@@ -0,0 +1,81 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Exclusive load/store family, LSE atomics and their memory operands,
// byte-parity-checked against go tool asm. The toolchain parses the FIRST
// register of an exclusive store as the data register and the LAST as the
// status register, and takes register pairs as (Rt1, Rt2) operands.
#include "textflag.h"
TEXT ·loads(SB), NOSPLIT, $0-0
LDXR (R1), R2
LDXRB (R1), R3
LDXRH (R1), R4
LDXRW (R1), R5
LDAXR (R1), R2
LDAXRB (R1), R3
LDAXRH (R1), R4
LDAXRW (R1), R5
LDXR (RSP), R2
LDXRB (RSP), R3
LDXRH (RSP), R4
LDXRW (RSP), R5
LDAXR (RSP), R2
LDAXRB (RSP), R3
LDAXRH (RSP), R4
RET
TEXT ·stores(SB), NOSPLIT, $0-0
STXR R2, (R1), R6
STXRB R3, (R1), R6
STXRH R4, (R1), R6
STXRW R5, (R1), R6
STLXR R2, (R1), R6
STLXRB R3, (R1), R6
STLXRH R4, (R1), R6
STLXRW R5, (R1), R6
STXR R2, (RSP), R6
STXRB R3, (RSP), R6
STXRH R4, (RSP), R6
STXRW R5, (RSP), R6
STLXR R2, (RSP), R6
STLXRB R3, (RSP), R6
STLXRH R4, (RSP), R6
RET
TEXT ·pairs(SB), NOSPLIT, $0-0
LDXP (R1), (R2, R3)
LDXPW (R1), (R2, R3)
LDAXP (R1), (R2, R3)
LDAXPW (R1), (R2, R3)
STXP (R2, R3), (R1), R6
STXPW (R2, R3), (R1), R6
STLXP (R2, R3), (R1), R6
STLXPW (R2, R3), (R1), R6
LDXP (RSP), (R2, R3)
LDXPW (RSP), (R2, R3)
LDAXP (RSP), (R4, R5)
LDAXPW (RSP), (R4, R5)
STXP (R2, R3), (RSP), R6
STXPW (R2, R3), (RSP), R6
STLXP (R4, R5), (RSP), R7
RET
TEXT ·atomics(SB), NOSPLIT, $0-0
LDADDB R2, (R1), R3
LDADDH R2, (R1), R3
LDADDW R2, (R1), R3
LDADDD R2, (R1), R3
LDADDB R2, (R1), ZR
LDADDH R2, (R1), ZR
LDADDW R2, (R1), ZR
LDADDD R2, (R1), ZR
CASW R2, (R1), R3
CASD R2, (R1), R3
CASW R2, (R1), ZR
CASD R2, (R1), ZR
SWPW R2, (R1), R3
SWPD R2, (R1), R3
SWPD R2, (R1), ZR
RET
+39
View File
@@ -0,0 +1,39 @@
// Mixed-width sign- and zero-extending moves plus PMOVMSKB, the spellings
// GOROOT's runtime and bytealg kernels use. Every result is folded back so
// no instruction is dead.
#include "textflag.h"
// func widen(p *byte) uint64
TEXT ·widen(SB), NOSPLIT, $0-16
MOVBQZX 0(DI), AX
MOVWQZX 2(DI), CX
ADDQ CX, AX
MOVLQZX 4(DI), DX
ADDQ DX, AX
MOVBQSX 8(DI), R8
ADDQ R8, AX
MOVWQSX 12(DI), R9
ADDQ R9, AX
MOVBLSX 16(DI), R10
ADDL R10, AX
MOVLQSX 20(DI), R11
ADDQ R11, AX
MOVQ AX, ret+8(FP)
RET
// func widenw(p *byte) int32
TEXT ·widenw(SB), NOSPLIT, $0-16
MOVBWZX 0(DI), AX
MOVBWSX 1(DI), CX
ADDL CX, AX
MOVLQZX AX, DX
MOVL DX, ret+8(FP)
RET
// func mask(x *XMM) int
TEXT ·mask(SB), NOSPLIT, $0-16
MOVOU 0(DI), X1
PMOVMSKB X1, AX
MOVQ AX, ret+8(FP)
RET
+14 -4
View File
@@ -42,8 +42,12 @@ func TestABIArm64(t *testing.T) {
t.Errorf("cleanAdd: %s", report) t.Errorf("cleanAdd: %s", report)
} }
// dirtyFP clobbers the frame pointer (R29). // dirtyFP clobbers the frame pointer (R29). The kernel passes its
out, report, err = k.CallFuncChecked("dirtyFP", make([]byte, 16)) // argument through, so the argument must carry the expected value the
// way the amd64 twin test seeds it.
args = make([]byte, 16)
PutUint64(args, 0, 0x1234)
out, report, err = k.CallFuncChecked("dirtyFP", args)
if err != nil { if err != nil {
t.Fatalf("CallFuncChecked: %v", err) t.Fatalf("CallFuncChecked: %v", err)
} }
@@ -57,11 +61,17 @@ func TestABIArm64(t *testing.T) {
t.Errorf("dirtyFP: only R29 should be clobbered: %s", report) t.Errorf("dirtyFP: only R29 should be clobbered: %s", report)
} }
// dirtyG clobbers the goroutine pointer (R28). // dirtyG clobbers the goroutine pointer (R28) and still returns its
_, report, err = k.CallFuncChecked("dirtyG", make([]byte, 16)) // argument.
args = make([]byte, 16)
PutUint64(args, 0, 0x5678)
out, report, err = k.CallFuncChecked("dirtyG", args)
if err != nil { if err != nil {
t.Fatalf("CallFuncChecked: %v", err) t.Fatalf("CallFuncChecked: %v", err)
} }
if got := int64(GetUint64(out, 8)); got != 0x5678 {
t.Errorf("dirtyG returned %d, want %d", got, int64(0x5678))
}
if !report.GClobbered { if !report.GClobbered {
t.Error("dirtyG: expected g clobbered, but report says clean") t.Error("dirtyG: expected g clobbered, but report says clean")
} }