The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode. operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX). The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).
Assisted-by: GLM 5.3
The two bank-crossing quadword moves take the mandatory prefix by
direction: the toolchain renders F3 0F D6 as MOVQ2DQ with the MMX
source and F2 0F D6 as MOVDQ2Q with the XMM source, and the encoder
emitted F2 for both, so MOVQ2DQ encoded MOVDQ2Q.
Assisted-by: GLM 5.3 Flash
The x87 stack controls, the D8/DC arithmetic pair, the conditional moves,
the register compares, FADDDP, the memory loads and the FXSAVE pair, each
pinned byte for byte against go tool asm through every corpus line the
toolchain's own amd64enc.s carries for the family (78 lines).
Assisted-by: GLM 5.3 Flash