Table-driven rows for the renderer's spellings (the L suffix or none with
a byte register encodes the byte form, every register joining at its low
byte), the refusals (W, Q and the MOVD alias take no byte register) and a
differential kernel of the B-suffixed spellings assembled through both
gasm and go tool asm, byte for byte.
Assisted-by: GLM 5.3
The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode. operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX). The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).
Assisted-by: GLM 5.3
The CASALH entry carried the CASB/CASH opcode pattern where the acquire
forms take the full fixed field, so the word differed from the
toolchain's in one opcode bit. Pinned against the oracle word.
Assisted-by: GLM 5.3 Flash
Offsets beyond the split bands ride a per-function literal pool the way
the toolchain lays one out: a PC-relative literal load into REGTMP, then
the register-offset access (the pair family adds the base addition), the
pooled words appended after the last instruction behind the UNDEF guard,
deduplicated by value with the sign- and width-aware load selection.
The same differential pass against the corpus exposed three wrong-code
bugs and fixes them: the logical-immediate period marker rode the wrong
position for every element below 64 bits, so the 32-bit forms encoded a
different constant than written; the plain register operand of an
ADD/SUB against SP took the shifted-register form where the toolchain
uses the extended one with the identity extend, silently truncating
through UXTB; and the AUTIA1716 and AUTIB1716 hint constants were the
PACIA and PACIB encodings. An offset sweep across every band boundary
now pins all three against the live oracle.
Assisted-by: GLM 5.3 Flash
Offsets the single-instruction forms cannot carry lower the way the
toolchain lowers them: ADD or SUB moves the whole distance into REGTMP
within the ±4095 band, and the 24-bit band above it splits into an ADD of
the high half and an access of the low half, with the pair family taking
the two-ADD sequence. The split band follows loadStoreClass per width,
byte accesses taking the full 24 bits and the Q width the widest, so an
offset the toolchain pools is never split instead.
Assisted-by: GLM 5.3 Flash
The pair encoder derived the imm7 divisor from the width suffix alone, so
the 128-bit FP pairs divided their offsets by eight and encoded twice the
distance. The Q spellings scale by sixteen like every other 128-bit
access; the differential kernel carries them now.
Assisted-by: GLM 5.3 Flash
FMOVQ routes through the MOV load/store machinery in the plain, post-index,
pre-index and static-symbol forms. The Q width carries its size in the opc
field, so the store spelling is opc=10 and the access scales by sixteen;
both come from helpers now instead of the size exponent. The static-symbol
form takes the toolchain's twelve-byte ADRP + ADD + access fallback with the
R_ADDRARM64 pair. The register-to-register and immediate forms stay
rejected, matching the toolchain's own table.
Assisted-by: GLM 5.3 Flash
The two bank-crossing quadword moves take the mandatory prefix by
direction: the toolchain renders F3 0F D6 as MOVQ2DQ with the MMX
source and F2 0F D6 as MOVDQ2Q with the XMM source, and the encoder
emitted F2 for both, so MOVQ2DQ encoded MOVDQ2Q.
Assisted-by: GLM 5.3 Flash
The toolchain writes the source immediate straight into imm[31:12]
(riscv64.s: AUIPC 24287, X10 encodes 7ffff517), and rejects values
beyond the signed 20-bit span; the encoder divided by 4096 instead and
truncated silently, so the high bits of every large AUIPC and LUI were
lost.
Assisted-by: GLM 5.3 Flash