Commit Graph
180 Commits
Author SHA1 Message Date
petrbalvin 36d1ac804a fix(asm): route the riscv64 register moves through the toolchain forms
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin da35883633 fix(asm): compress the riscv64 two-operand arithmetic and immediate tail
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 8bded3ea39 test(asm): pin the byte-form width reconciliation and its oracle
Table-driven rows for the renderer's spellings (the L suffix or none with
a byte register encodes the byte form, every register joining at its low
byte), the refusals (W, Q and the MOVD alias take no byte register) and a
differential kernel of the B-suffixed spellings assembled through both
gasm and go tool asm, byte for byte.

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 98a562d8b3 fix(asm): settle the byte-form width from the register operands
The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode.  operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX).  The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 33e7fdac98 fix(asm): enforce the arm64 TLBI, RPRFM, FCVT and integer-pair arities
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 7a69be8b59 fix(asm): reject the arm64 REGTMP spellings the toolchain refuses
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 2385bb7069 fix(asm): enforce the arm64 VLD/VST post-index contract
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 7bc80ccb54 fix(asm): emit nothing for the arm64 NOP pseudo-instruction
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 405e2427ed fix(asm): place the arm64 literal pool the way the toolchain flushes it
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 4fc96decc4 fix(asm): size the arm64 logical-immediate materialisation exactly
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 409c8b348d fix(asm): bound the arm64 VTBL table list before the destination read
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 2b2a72d54e fix(asm): encode the arm64 bitfield aliases with their own opc
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 8ab99c9b0c fix(asm): tighten the arm64 acceptance toward the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 82d741514a fix(asm): key the arm64 immediate class order on the ZR spelling
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin f20aa156d0 fix(asm): treat the arm64 $-8 frame as frameless and encode the RET forms
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin d853432dba fix(asm): route the arm64 logical immediates to ZR through REGTMP
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin fbdad8424f feat(asm): encode the arm64 FP immediate moves
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a9b54b6868 fix(asm): carry the arm64 immediate to ZR through MOVZ
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin d786b90fa1 feat(asm): lower the arm64 con(register) form to the ADD chain
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 9ef14bdb71 feat(asm): encode the arm64 SIMD arrangement bits
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 458cdd2066 feat(asm): encode the arm64 register-offset addressing forms
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a344399b81 fix(asm): set the CASALH opcode bit fifteen
The CASALH entry carried the CASB/CASH opcode pattern where the acquire
forms take the full fixed field, so the word differed from the
toolchain's in one opcode bit.  Pinned against the oracle word.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a0fa7e802c feat(asm): pool the arm64 offsets the split bands cannot carry
Offsets beyond the split bands ride a per-function literal pool the way
the toolchain lays one out: a PC-relative literal load into REGTMP, then
the register-offset access (the pair family adds the base addition), the
pooled words appended after the last instruction behind the UNDEF guard,
deduplicated by value with the sign- and width-aware load selection.

The same differential pass against the corpus exposed three wrong-code
bugs and fixes them: the logical-immediate period marker rode the wrong
position for every element below 64 bits, so the 32-bit forms encoded a
different constant than written; the plain register operand of an
ADD/SUB against SP took the shifted-register form where the toolchain
uses the extended one with the identity extend, silently truncating
through UXTB; and the AUTIA1716 and AUTIB1716 hint constants were the
PACIA and PACIB encodings.  An offset sweep across every band boundary
now pins all three against the live oracle.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 71e8dd550d feat(asm): split wide arm64 load and store offsets into REGTMP
Offsets the single-instruction forms cannot carry lower the way the
toolchain lowers them: ADD or SUB moves the whole distance into REGTMP
within the ±4095 band, and the 24-bit band above it splits into an ADD of
the high half and an access of the low half, with the pair family taking
the two-ADD sequence.  The split band follows loadStoreClass per width,
byte accesses taking the full 24 bits and the Q width the widest, so an
offset the toolchain pools is never split instead.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 2bd52eb7ad fix(asm): scale the FLDPQ and FSTPQ pair offsets by sixteen
The pair encoder derived the imm7 divisor from the width suffix alone, so
the 128-bit FP pairs divided their offsets by eight and encoded twice the
distance.  The Q spellings scale by sixteen like every other 128-bit
access; the differential kernel carries them now.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin daf7fad5b9 feat(asm): encode the arm64 Q-width FP load and store
FMOVQ routes through the MOV load/store machinery in the plain, post-index,
pre-index and static-symbol forms.  The Q width carries its size in the opc
field, so the store spelling is opc=10 and the access scales by sixteen;
both come from helpers now instead of the size exponent.  The static-symbol
form takes the toolchain's twelve-byte ADRP + ADD + access fallback with the
R_ADDRARM64 pair.  The register-to-register and immediate forms stay
rejected, matching the toolchain's own table.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a6bd9c1ebe feat(asm): encode the LDEORAL acquire variants
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 6c4932c4ec feat(arch): the SVE2.1 Z-alias permutations and copies
Assisted-by: GLM 5.3 Flash
2026-10-07 02:27:35 +02:00
petrbalvin 8231302bca feat(arch): the SVE predicate family in the extended layer
Assisted-by: GLM 5.3 Flash
2026-10-07 02:26:19 +02:00
petrbalvin 9c951c232e feat(asm): assemble the extended instruction layer on arm64
Assisted-by: GLM 5.3 Flash
2026-10-07 02:23:37 +02:00
petrbalvin 470639cd68 test(asm): carry the fuzz pipeline to arm64
Assisted-by: GLM 5.3 Flash
2026-10-07 02:14:37 +02:00
petrbalvin e7a1c20467 fix(asm): prefix MOVQ2DQ with F3
The two bank-crossing quadword moves take the mandatory prefix by
direction: the toolchain renders F3 0F D6 as MOVQ2DQ with the MMX
source and F2 0F D6 as MOVDQ2Q with the XMM source, and the encoder
emitted F2 for both, so MOVQ2DQ encoded MOVDQ2Q.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:12:21 +02:00
petrbalvin 0cfed5516c fix(asm): carry the riscv64 U-type immediate raw
The toolchain writes the source immediate straight into imm[31:12]
(riscv64.s: AUIPC 24287, X10 encodes 7ffff517), and rejects values
beyond the signed 20-bit span; the encoder divided by 4096 instead and
truncated silently, so the high bits of every large AUIPC and LUI were
lost.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:12:21 +02:00
petrbalvin a829b8f321 fix(parser): read the segment-absolute rendering FS:0
The toolchain's disassembler prints the segment-prefixed disp32
absolute as FS:0, but the bare-name branch read the segment register
alone and dropped the offset, so MOVQ FS:0, DX silently encoded a
register move.  The colon-offset form now lowers to the same
segment-absolute operand the 0(FS) spelling takes.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:12:21 +02:00
petrbalvin abc2d32b83 fix(asm): encode the CMOV condition the renderer prints
The renderer spells a conditional move CMOV plus the condition alone
(CMOVLE, CMOVG), the width carried by the operand registers, so
CMOVLE parsed as the size L and the condition E and encoded CMOVE.
A suffix that is itself a condition name now reads as that condition
with the width from the destination register, and the Plan 9
size-prefixed spellings keep their parse.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:12:21 +02:00
petrbalvin d275dee3ae feat(arch): add the scalar FP16 memory forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 01:35:48 +02:00
petrbalvin 380314a05f test(asm): keep the data-section ceiling out of the seed corpus
Assisted-by: GLM 5.3 Flash
2026-10-07 01:14:58 +02:00
petrbalvin de2f28504a test(asm): parse once and assemble twice in the fuzz body
Assisted-by: GLM 5.3 Flash
2026-10-07 01:14:58 +02:00
petrbalvin 7878112b93 fix(asm): bound the data section to what the image can materialise
Assisted-by: GLM 5.3 Flash
2026-10-07 01:14:58 +02:00
petrbalvin dd074dbafd fix(asm): reject an out-of-range GLOBL size with a diagnostic
Assisted-by: GLM 5.3 Flash
2026-10-07 01:13:40 +02:00
petrbalvin 14de1dd287 test(asm): fuzz the parse-and-assemble pipeline for amd64
Assisted-by: GLM 5.3 Flash
2026-10-07 01:12:27 +02:00
petrbalvin 9b751f6e58 feat(asm): encode the riscv64 quad-precision family and fix the fp cvt paths
Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin 955bc6643e fix(asm): carry the riscv64 scalar pseudos and the swapped branches
Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin 82e8919208 feat(asm): encode the riscv64 privileged instructions
Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin 778c297214 feat(asm): encode the riscv64 vector arithmetic families
Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin 5de9e985f8 feat(asm): encode the riscv64 vector load and store families
The vector memory section stopped at the three hand-written shapes the
GOROOT kernels use: every other spelling the toolchain accepts, the
width variants, the constant-stride and indexed accesses, the segment
families, the fault-only-first loads, the whole-register moves and the
bit-mask pair were names without an encoder.  The mnemonic now parses
into its own fields (direction, segment count, addressing mode, width,
fault-only-first and whole-register markers) and one encoder lays the
word down, with the optional V0 mask operand and the toolchain's
operand shapes.  VSETVL joins the configuration settings.  The
toolchain's whole vector memory section, six hundred and twenty-eight
statements of masked and unmasked forms, is a differential test against
the oracle, word for word.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin e43fa39dc9 feat(asm): encode the explicit riscv64 compressed instructions
The C extension's own spellings were names the table carried and the
encoder refused: CLWSP stopped the corpus audit's riscv64 file first.
Thirty-eight mnemonics now encode directly to their halfword, with the
toolchain's operand spellings and validation: the stack loads and stores
pin their base to SP, the register-based loads, stores and arithmetic
carry prime registers, CLUI refuses zero and SP, CADDI4SPN scales by
four, CADDI16SP by sixteen, and CJ, CBEQZ and CBNEZ resolve their N(PC)
targets against the final layout, taking a two-byte placeholder in the
early passes so the offsets stay honest.  CAND with an immediate is the
toolchain's C.ANDI spelling.  The toolchain's whole C extension testdata
block is a differential test, halfword for halfword, beside a range test
at the toolchain's own boundaries.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin ddfa33ccc1 feat(asm): encode the riscv64 bit-manipulation families
The Zba address generation, Zbb unary bit operations, Zbc carry-less
multiplication and Zbs single-bit families were names the table carried
and the encoder refused: thirty spellings plus RORI and XNOR fell over.
The register and immediate forms now encode as the toolchain does, the
unary operations carry their fixed rs2 constant, RORI lowers to ROR's
expansion (its reverse shift compressing like ROR's), XNOR XORs and
inverts in place, and ROL/ROLW rotate left through the same temporary
the toolchain uses, taking a register amount only as its own expansion
requires.  The toolchain's whole testdata block for these families is
now a differential test: every word must agree byte for byte.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin d82ef33fa9 fix(asm): bound the riscv64 shift immediate at the instruction width
SLLI $64 assembled with the amount silently masked into the six-bit
field where the toolchain rejects it, and the word forms took 0-63 where
they take 0-31.  Both families now validate against their own width and
the check reads the immediate at full width, so a value the source
spelled beyond int32 cannot wrap into the range; the boundary is pinned
in a test.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin c3540f0549 fix(asm): read the riscv64 raw-data immediates at full width
WORD and BYTE read their immediate through the truncating helper, so the
int32 wrap turned WORD $0xffffffff into -1 and rejected it, while WORD
$0x100000000 and BYTE $0x100000001 arrived pre-truncated and slipped
past the range check as small values.  Both statements now read the
immediate the source wrote and bound it at the toolchain's own limits:
[0, 0xffffffff] for WORD, [0, 0xff] for BYTE, with the bounds pinned in
a test.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00