Commit Graph
199 Commits
Author SHA1 Message Date
petrbalvin 8080e0acef feat(arch): add the AVX512-FP16 FMA families to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 18:58:49 +02:00
petrbalvin acd30088af fix(asm): reject the operand-starved riscv64 spellings instead of panicking
Assisted-by: GLM 5.3
2026-10-07 13:53:37 +02:00
petrbalvin d03de62c07 feat(arch): add the amd64 fp16 packed imm8-control group
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin 8de1b371da feat(arch): add the amd64 fp16 packed conversion family
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin fb6d01a7d0 feat(arch): encode the amd64 embedded rounding and SAE decorations
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin 104bea036b feat(asm): encode the arm64 SVE gather loads and scatter stores
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin 6faf850793 feat(asm): encode the arm64 SVE2 crypto, counter and reduction families
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin dac0a5b51b docs(asm): state the seed layout of the arch fuzz targets exactly
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 965b33e5b0 test(asm): seed the riscv64 and loong64 assemble fuzz targets
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin b9dfb48d79 fix(asm): encode the loong64 64-bit-span 2RI14 offsets like the toolchain
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 41aa8edfd2 fix(asm): match the toolchain's loong64 logical immediate expansion
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 8fdc511d0d test(asm): add the loong64 error-parity catalogue
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e5035d92e9 fix(asm): reject the offset on the loong64 register-indexed memory form
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin b54547c08d fix(asm): reject the shifted-register compositions on loong64
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e9510e8a68 test(asm): pin the riscv64 tail against the toolchain byte for byte
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 9d50212a71 fix(asm): refuse the riscv64 width moves across register banks
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin f2892e4f59 feat(asm): emit the riscv64 local-exec TLS sequence
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin f8dbd4f017 fix(asm): lower the riscv64 immediate CSR pseudos onto their opcode forms
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e5aabf9801 fix(asm): encode the riscv64 FENCE predecessor and successor flags
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 36d1ac804a fix(asm): route the riscv64 register moves through the toolchain forms
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin da35883633 fix(asm): compress the riscv64 two-operand arithmetic and immediate tail
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 8bded3ea39 test(asm): pin the byte-form width reconciliation and its oracle
Table-driven rows for the renderer's spellings (the L suffix or none with
a byte register encodes the byte form, every register joining at its low
byte), the refusals (W, Q and the MOVD alias take no byte register) and a
differential kernel of the B-suffixed spellings assembled through both
gasm and go tool asm, byte for byte.

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 98a562d8b3 fix(asm): settle the byte-form width from the register operands
The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode.  operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX).  The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 33e7fdac98 fix(asm): enforce the arm64 TLBI, RPRFM, FCVT and integer-pair arities
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 7a69be8b59 fix(asm): reject the arm64 REGTMP spellings the toolchain refuses
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 2385bb7069 fix(asm): enforce the arm64 VLD/VST post-index contract
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 7bc80ccb54 fix(asm): emit nothing for the arm64 NOP pseudo-instruction
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 405e2427ed fix(asm): place the arm64 literal pool the way the toolchain flushes it
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 4fc96decc4 fix(asm): size the arm64 logical-immediate materialisation exactly
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 409c8b348d fix(asm): bound the arm64 VTBL table list before the destination read
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 2b2a72d54e fix(asm): encode the arm64 bitfield aliases with their own opc
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 8ab99c9b0c fix(asm): tighten the arm64 acceptance toward the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 82d741514a fix(asm): key the arm64 immediate class order on the ZR spelling
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin f20aa156d0 fix(asm): treat the arm64 $-8 frame as frameless and encode the RET forms
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin d853432dba fix(asm): route the arm64 logical immediates to ZR through REGTMP
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin fbdad8424f feat(asm): encode the arm64 FP immediate moves
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a9b54b6868 fix(asm): carry the arm64 immediate to ZR through MOVZ
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin d786b90fa1 feat(asm): lower the arm64 con(register) form to the ADD chain
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 9ef14bdb71 feat(asm): encode the arm64 SIMD arrangement bits
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 458cdd2066 feat(asm): encode the arm64 register-offset addressing forms
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a344399b81 fix(asm): set the CASALH opcode bit fifteen
The CASALH entry carried the CASB/CASH opcode pattern where the acquire
forms take the full fixed field, so the word differed from the
toolchain's in one opcode bit.  Pinned against the oracle word.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a0fa7e802c feat(asm): pool the arm64 offsets the split bands cannot carry
Offsets beyond the split bands ride a per-function literal pool the way
the toolchain lays one out: a PC-relative literal load into REGTMP, then
the register-offset access (the pair family adds the base addition), the
pooled words appended after the last instruction behind the UNDEF guard,
deduplicated by value with the sign- and width-aware load selection.

The same differential pass against the corpus exposed three wrong-code
bugs and fixes them: the logical-immediate period marker rode the wrong
position for every element below 64 bits, so the 32-bit forms encoded a
different constant than written; the plain register operand of an
ADD/SUB against SP took the shifted-register form where the toolchain
uses the extended one with the identity extend, silently truncating
through UXTB; and the AUTIA1716 and AUTIB1716 hint constants were the
PACIA and PACIB encodings.  An offset sweep across every band boundary
now pins all three against the live oracle.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 71e8dd550d feat(asm): split wide arm64 load and store offsets into REGTMP
Offsets the single-instruction forms cannot carry lower the way the
toolchain lowers them: ADD or SUB moves the whole distance into REGTMP
within the ±4095 band, and the 24-bit band above it splits into an ADD of
the high half and an access of the low half, with the pair family taking
the two-ADD sequence.  The split band follows loadStoreClass per width,
byte accesses taking the full 24 bits and the Q width the widest, so an
offset the toolchain pools is never split instead.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 2bd52eb7ad fix(asm): scale the FLDPQ and FSTPQ pair offsets by sixteen
The pair encoder derived the imm7 divisor from the width suffix alone, so
the 128-bit FP pairs divided their offsets by eight and encoded twice the
distance.  The Q spellings scale by sixteen like every other 128-bit
access; the differential kernel carries them now.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin daf7fad5b9 feat(asm): encode the arm64 Q-width FP load and store
FMOVQ routes through the MOV load/store machinery in the plain, post-index,
pre-index and static-symbol forms.  The Q width carries its size in the opc
field, so the store spelling is opc=10 and the access scales by sixteen;
both come from helpers now instead of the size exponent.  The static-symbol
form takes the toolchain's twelve-byte ADRP + ADD + access fallback with the
R_ADDRARM64 pair.  The register-to-register and immediate forms stay
rejected, matching the toolchain's own table.

Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin a6bd9c1ebe feat(asm): encode the LDEORAL acquire variants
Assisted-by: GLM 5.3 Flash
2026-10-07 02:36:24 +02:00
petrbalvin 6c4932c4ec feat(arch): the SVE2.1 Z-alias permutations and copies
Assisted-by: GLM 5.3 Flash
2026-10-07 02:27:35 +02:00
petrbalvin 8231302bca feat(arch): the SVE predicate family in the extended layer
Assisted-by: GLM 5.3 Flash
2026-10-07 02:26:19 +02:00
petrbalvin 9c951c232e feat(asm): assemble the extended instruction layer on arm64
Assisted-by: GLM 5.3 Flash
2026-10-07 02:23:37 +02:00
petrbalvin 470639cd68 test(asm): carry the fuzz pipeline to arm64
Assisted-by: GLM 5.3 Flash
2026-10-07 02:14:37 +02:00