The toolchain's assembler corpus carries 195 amd64 encodings the
x/arch decoder rejects or degenerates: the BMI1/BMI2 VEX families
(ANDN, BEXTR, BLSI, BLSMSK, BLSR, BZHI, MULX, PDEP, PEXT, RORX,
SARX, SHLX, SHRX), the 0F 01 quartet CLAC, STAC, RDPKRU and WRPKRU,
the bare and REX-only RDSEED forms, and UD1. The supplementary
naming table decodes the VEX prefix and the ModR/M shape and renders
the toolchain's own spellings; every corpus row is pinned in the
unlisted fixture and round-trips byte for byte through the encoder,
and the boundary test pins the prefix shapes no family carries.
Assisted-by: GLM 5.3
The operand-width reconciliation closes the byte-register fixture lines
byte for byte: XADDL, XCHGL, CMPXCHGL and CRC32 with byte registers, the
ALU and TEST immediates against AL and DL, and the unlisted accumulator
short forms. MOVL $0x7, DL stays mapped for the legal-encoding choice
alone: the toolchain's own table says "c6c207 or b207", go tool asm emits
b207, and the fixed-point invariant holds with it.
Assisted-by: GLM 5.3
Table-driven rows for the renderer's spellings (the L suffix or none with
a byte register encodes the byte form, every register joining at its low
byte), the refusals (W, Q and the MOVD alias take no byte register) and a
differential kernel of the B-suffixed spellings assembled through both
gasm and go tool asm, byte for byte.
Assisted-by: GLM 5.3
The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode. operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX). The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).
Assisted-by: GLM 5.3
The launch-failure audit raced the clock: it asserted the dead debuggee
surfaced within 1.5 seconds, a bound the loaded machine behind a ten-way
test storm regularly starved past even though the poll detects the dead
notice within milliseconds of its appearance. The audit now proves the
property itself: a stub debuggee that starts in single-digit milliseconds
marks itself dead, so the notice is always inside the poll's budget and
the error must come from the dead-file watch, while the real binary is
checked without any wall-clock bound. A companion audit drives Kill
through a parked, a doubly killed and a run-to-exit session, the states
whose cleanup used to hang the package, under a watchdog.
Assisted-by: GLM 5.3
The Go runtime can migrate the debuggee's target-mode goroutine off the
process leader before PTRACE_TRACEME, which left the trace relation on a
thread the session never addressed: its stops starved the waits on the
leader, and a kill sequence that resumed nothing and then blocked in
Wait4 hung the whole package. The debuggee now reports the traced thread
in the launch handshake and parks with a thread-directed stop, every
ptrace request and wait addresses that thread, a SIGURG arriving on a
single-step resumes it as a single-step again instead of letting the
tracee run uncontrolled, a resume rejected with ESRCH lifts a group-stop
with SIGCONT and retries once, and Kill resumes, kills and reaps through
non-blocking waits so it returns for a tracee in any state.
Assisted-by: GLM 5.3
The CASALH entry carried the CASB/CASH opcode pattern where the acquire
forms take the full fixed field, so the word differed from the
toolchain's in one opcode bit. Pinned against the oracle word.
Assisted-by: GLM 5.3 Flash
Offsets beyond the split bands ride a per-function literal pool the way
the toolchain lays one out: a PC-relative literal load into REGTMP, then
the register-offset access (the pair family adds the base addition), the
pooled words appended after the last instruction behind the UNDEF guard,
deduplicated by value with the sign- and width-aware load selection.
The same differential pass against the corpus exposed three wrong-code
bugs and fixes them: the logical-immediate period marker rode the wrong
position for every element below 64 bits, so the 32-bit forms encoded a
different constant than written; the plain register operand of an
ADD/SUB against SP took the shifted-register form where the toolchain
uses the extended one with the identity extend, silently truncating
through UXTB; and the AUTIA1716 and AUTIB1716 hint constants were the
PACIA and PACIB encodings. An offset sweep across every band boundary
now pins all three against the live oracle.
Assisted-by: GLM 5.3 Flash
Offsets the single-instruction forms cannot carry lower the way the
toolchain lowers them: ADD or SUB moves the whole distance into REGTMP
within the ±4095 band, and the 24-bit band above it splits into an ADD of
the high half and an access of the low half, with the pair family taking
the two-ADD sequence. The split band follows loadStoreClass per width,
byte accesses taking the full 24 bits and the Q width the widest, so an
offset the toolchain pools is never split instead.
Assisted-by: GLM 5.3 Flash
The pair encoder derived the imm7 divisor from the width suffix alone, so
the 128-bit FP pairs divided their offsets by eight and encoded twice the
distance. The Q spellings scale by sixteen like every other 128-bit
access; the differential kernel carries them now.
Assisted-by: GLM 5.3 Flash
FMOVQ routes through the MOV load/store machinery in the plain, post-index,
pre-index and static-symbol forms. The Q width carries its size in the opc
field, so the store spelling is opc=10 and the access scales by sixteen;
both come from helpers now instead of the size exponent. The static-symbol
form takes the toolchain's twelve-byte ADRP + ADD + access fallback with the
R_ADDRARM64 pair. The register-to-register and immediate forms stay
rejected, matching the toolchain's own table.
Assisted-by: GLM 5.3 Flash
The hint NOP opcode 0F 1C /r with a memory operand is CLDEMOTE, a
memory-only instruction the toolchain's own table carries; the decoder
rejects the encoding instead of naming it. The rejected-encoding side
of the supplementary table names it from the bytes, and the corpus row
0f1c03 pins the text in the unlisted fixture, round trip byte exact.
Assisted-by: GLM 5.3 Flash
x/arch reports the ADCX, ADOX, RDSEED, RDPID, TPAUSE, UMONITOR, UMWAIT
and ENDBR families with no error but the degenerate zero instruction,
which GoSyntax renders as Op(0) under its prefix decoration and with a
length of one. A supplementary naming table keyed by the opcode
pattern restores the toolchain's own spellings and lengths; the parity
fixtures pin all 41 corpus rows (ENDBR32 alone, which the toolchain
cannot spell, pins as bytes and text in the focused naming test).
Assisted-by: GLM 5.3 Flash
The two bank-crossing quadword moves take the mandatory prefix by
direction: the toolchain renders F3 0F D6 as MOVQ2DQ with the MMX
source and F2 0F D6 as MOVDQ2Q with the XMM source, and the encoder
emitted F2 for both, so MOVQ2DQ encoded MOVDQ2Q.
Assisted-by: GLM 5.3 Flash
The toolchain writes the source immediate straight into imm[31:12]
(riscv64.s: AUIPC 24287, X10 encodes 7ffff517), and rejects values
beyond the signed 20-bit span; the encoder divided by 4096 instead and
truncated silently, so the high bits of every large AUIPC and LUI were
lost.
Assisted-by: GLM 5.3 Flash
The toolchain's disassembler prints the segment-prefixed disp32
absolute as FS:0, but the bare-name branch read the segment register
alone and dropped the offset, so MOVQ FS:0, DX silently encoded a
register move. The colon-offset form now lowers to the same
segment-absolute operand the 0(FS) spelling takes.
Assisted-by: GLM 5.3 Flash
The renderer spells a conditional move CMOV plus the condition alone
(CMOVLE, CMOVG), the width carried by the operand registers, so
CMOVLE parsed as the size L and the condition E and encoded CMOVE.
A suffix that is itself a condition name now reads as that condition
with the width from the destination register, and the Plan 9
size-prefixed spellings keep their parse.
Assisted-by: GLM 5.3 Flash
The vector memory section stopped at the three hand-written shapes the
GOROOT kernels use: every other spelling the toolchain accepts, the
width variants, the constant-stride and indexed accesses, the segment
families, the fault-only-first loads, the whole-register moves and the
bit-mask pair were names without an encoder. The mnemonic now parses
into its own fields (direction, segment count, addressing mode, width,
fault-only-first and whole-register markers) and one encoder lays the
word down, with the optional V0 mask operand and the toolchain's
operand shapes. VSETVL joins the configuration settings. The
toolchain's whole vector memory section, six hundred and twenty-eight
statements of masked and unmasked forms, is a differential test against
the oracle, word for word.
Assisted-by: GLM 5.3 Flash
The C extension's own spellings were names the table carried and the
encoder refused: CLWSP stopped the corpus audit's riscv64 file first.
Thirty-eight mnemonics now encode directly to their halfword, with the
toolchain's operand spellings and validation: the stack loads and stores
pin their base to SP, the register-based loads, stores and arithmetic
carry prime registers, CLUI refuses zero and SP, CADDI4SPN scales by
four, CADDI16SP by sixteen, and CJ, CBEQZ and CBNEZ resolve their N(PC)
targets against the final layout, taking a two-byte placeholder in the
early passes so the offsets stay honest. CAND with an immediate is the
toolchain's C.ANDI spelling. The toolchain's whole C extension testdata
block is a differential test, halfword for halfword, beside a range test
at the toolchain's own boundaries.
Assisted-by: GLM 5.3 Flash
The Zba address generation, Zbb unary bit operations, Zbc carry-less
multiplication and Zbs single-bit families were names the table carried
and the encoder refused: thirty spellings plus RORI and XNOR fell over.
The register and immediate forms now encode as the toolchain does, the
unary operations carry their fixed rs2 constant, RORI lowers to ROR's
expansion (its reverse shift compressing like ROR's), XNOR XORs and
inverts in place, and ROL/ROLW rotate left through the same temporary
the toolchain uses, taking a register amount only as its own expansion
requires. The toolchain's whole testdata block for these families is
now a differential test: every word must agree byte for byte.
Assisted-by: GLM 5.3 Flash
SLLI $64 assembled with the amount silently masked into the six-bit
field where the toolchain rejects it, and the word forms took 0-63 where
they take 0-31. Both families now validate against their own width and
the check reads the immediate at full width, so a value the source
spelled beyond int32 cannot wrap into the range; the boundary is pinned
in a test.
Assisted-by: GLM 5.3 Flash
WORD and BYTE read their immediate through the truncating helper, so the
int32 wrap turned WORD $0xffffffff into -1 and rejected it, while WORD
$0x100000000 and BYTE $0x100000001 arrived pre-truncated and slipped
past the range check as small values. Both statements now read the
immediate the source wrote and bound it at the toolchain's own limits:
[0, 0xffffffff] for WORD, [0, 0xff] for BYTE, with the bounds pinned in
a test.
Assisted-by: GLM 5.3 Flash
The riscv64 assembler carried fourteen hand-picked CSR names where the
toolchain resolves three hundred and twenty-nine: a CSRR/CSRW family
instruction naming any privileged register beyond the few base ones came
out as unknown CSR. The table now carries the RISC-V privileged
specification's register set exactly as go tool asm spells it, and a
differential test assembles every name through both assemblers and
requires the words to agree byte for byte.
Assisted-by: GLM 5.3 Flash
Diagnostics no longer repeat the generated table's ignorance of a
registered mnemonic: unknown-instruction and unencodable-instruction
against a statement the registry encodes are filtered from the server's
own presentation, driven by asm.LookupExtension directly. The filter is
a no-op once lint learns the registry, so the two compose unchanged.
Assisted-by: GLM 5.3 Flash
Completion merges the extension registry's mnemonics beside the toolchain
entries, with operand shapes measured against the layer's own encoder, and
hover documents a registered mnemonic from its metadata: the extension
notice, the forms, the features, the fixed encodings and the manual
references.
Assisted-by: GLM 5.3 Flash
A line-for-line byte comparison of the whole amd64enc.s corpus against
go tool asm surfaced divergences the pass-only accounting never showed:
PEXTRW's GPR form swapped its fields, PUSHW took an imm32 where the
toolchain bounds the immediate to 16 bits, the double shift wrote the
unmasked register number into the reg field, VCOMISS carried a 0x66
prefix, RORX dropped the destination's R bit, and the variable bit
shifts used the manual's per-width opcodes where the toolchain
consolidates each row on one opcode with the W bit. The VEX forms the
toolchain prefers for plain vector registers (the SSE2/SSSE3/SSE4.1
AVX twins, the compare-with-predicate family, VMOVUPS, VSHUFPS, the
variable shifts) now encode under VEX, with EVEX left to the ZMM,
opmask and index-16+ spellings, and the mnemonics whose rows never
offer the 2-byte prefix force it. Every line is pinned through the new
corpus parity test (793 lines); the whole corpus file now assembles to
the toolchain's bytes at every commented line (10022 of 10022).
Assisted-by: GLM 5.3 Flash
The SSE3 horizontal and add-subtract pairs, the SSSE3 sign and
horizontal integers, the masked moves in both directions, the
reciprocity and test pairs, the AVX imm8 tail (blends, dot products,
inserts, rounds, MPSADBW, the string compares), the four-operand
variable blends with their /is4 mask byte, the scalar three-operand
moves, the MXCSR accessors, the VPERMIL register controls and the
variable word shifts, plus the BMI2 count forms over memory. Every
encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(852 lines); the /is4 byte carries the mask register number in its high
nibble, the layout the toolchain emits.
Assisted-by: GLM 5.3 Flash
The packed integer and float binaries, the imm8-controlled SSE4.1 forms,
the variable blends with their X0 mask, the high/low half moves, the
sign-mask extractions, the non-temporal stores, the MOVQ bank crossings
and their odd spellings, the MMX shifts and shuffle and the cache-line
mask stores, plus the scalar leaves LEAVE, INVPCID and the RTM controls.
Every encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(1465 lines). Two corpus-wide gaps fell out of the comparison: the
64-bit MOV immediate uses the zero-extending form across the unsigned
32-bit span, and the MMX-to-GPR MOVQ puts the bank register in reg.
Assisted-by: GLM 5.3 Flash
XSAVE, XSAVEOPT, XSAVEC and XSAVES with their restore twins, plain and
64, each pinned byte for byte against go tool asm through every corpus
line the toolchain's own amd64enc.s carries for the family (24 lines).
The toolchain emits XSAVEOPT without the manual's 0x66 prefix; the bytes
are the oracle, so the family carries none.
Assisted-by: GLM 5.3 Flash
The no-operand flag and system controls, the sign-extension pair, the
string primitives, the multi-byte no-ops, the cache controls, MOVBE, the
compare-exchange doubles, the random source and FS/GS base pairs, the
descriptor-table accesses, the 0F 00/01 register controls and the
LAR/LSL selector reads and far-segment loads, each pinned byte for byte
against go tool asm through every corpus line the toolchain's own
amd64enc.s carries for the families (279 lines).
Assisted-by: GLM 5.3 Flash
The x87 stack controls, the D8/DC arithmetic pair, the conditional moves,
the register compares, FADDDP, the memory loads and the FXSAVE pair, each
pinned byte for byte against go tool asm through every corpus line the
toolchain's own amd64enc.s carries for the family (78 lines).
Assisted-by: GLM 5.3 Flash