Commit Graph
224 Commits
Author SHA1 Message Date
petrbalvin 0211d6672d feat(asm): expand riscv64 memory offsets beyond the 12-bit immediate
Assisted-by: GLM 5.3 Flash
2026-10-07 20:39:44 +02:00
petrbalvin 7aba29ac67 test(arch): pin the golden vectors of the last four families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:42 +02:00
petrbalvin aeb109a64c feat(arch): the vector-length arithmetic pseudo group
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin bd85f8838d feat(arch): the SVE2.1 last-active vector and compare families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 12a5cfff52 feat(arch): the rest of the SVE2 BFloat16 wall
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin cdc3a75c88 feat(arch): the SVE multiple-structure loads and stores
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 82dbf087a8 feat(arch): the SVE2 BFloat16 arithmetic core
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 0760cc91af feat(arch): the SVE2 three-source and bitwise combine families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin bd80499c74 feat(arch): the SVE2 shift-by-vector family
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin b93fccc075 feat(arch): the SVE2.1 pairwise and quadword-reduction families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 31ab584eee feat(arch): the SVE2.1 narrowing two-to-one family
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin ee2c6d51b3 fix(arch): keep Zdn out of the class bits of the predicated Z-alias source
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 6e49cbd099 feat(asm): assemble the extended instruction layer on amd64
Assisted-by: GLM 5.3 Flash
2026-10-07 20:24:51 +02:00
petrbalvin d4878524e8 test(asm): pin the clean GOROOT arm64 set against the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-07 19:54:22 +02:00
petrbalvin 07f4622ffe style(asm): range over the kernel generator's statement count
Assisted-by: GLM 5.3 Flash
2026-10-07 19:54:22 +02:00
petrbalvin b3ede5f952 test(asm): pin the mid-function pool flush against the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-07 19:54:22 +02:00
petrbalvin 5a5936d222 feat(asm): drain the arm64 literal pool mid-function at the distance bound
Assisted-by: GLM 5.3 Flash
2026-10-07 19:54:22 +02:00
petrbalvin 4632ac1bb9 feat(arch): add the AVX-VNNI-INT16 dot products to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 27ef71859b feat(arch): add VMINMAXSH to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 03f9ef0ac6 feat(arch): add the FP16 complex fused multiply-add to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 86cb785e58 fix(asm): read the MOV family registers against their banks
Assisted-by: GLM 5.3 Flash
2026-10-07 19:40:32 +02:00
petrbalvin d18195f581 test(asm): pin the riscv64 error parity against the toolchain catalogues
Assisted-by: GLM 5.3 Flash
2026-10-07 19:40:32 +02:00
petrbalvin f0942ff7f3 fix(asm): reject the operand shapes the toolchain rejects on riscv64
Assisted-by: GLM 5.3 Flash
2026-10-07 19:40:32 +02:00
petrbalvin 6c66f5bbd9 test(asm): encode the new FP16 families through the registry
Assisted-by: GLM 5.3 Flash
2026-10-07 19:00:44 +02:00
petrbalvin 509afbb6c9 feat(arch): add the FP16 complex multiply and minimum-maximum to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 18:59:09 +02:00
petrbalvin 8080e0acef feat(arch): add the AVX512-FP16 FMA families to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 18:58:49 +02:00
petrbalvin acd30088af fix(asm): reject the operand-starved riscv64 spellings instead of panicking
Assisted-by: GLM 5.3
2026-10-07 13:53:37 +02:00
petrbalvin d03de62c07 feat(arch): add the amd64 fp16 packed imm8-control group
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin 8de1b371da feat(arch): add the amd64 fp16 packed conversion family
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin fb6d01a7d0 feat(arch): encode the amd64 embedded rounding and SAE decorations
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin 104bea036b feat(asm): encode the arm64 SVE gather loads and scatter stores
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin 6faf850793 feat(asm): encode the arm64 SVE2 crypto, counter and reduction families
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin dac0a5b51b docs(asm): state the seed layout of the arch fuzz targets exactly
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 965b33e5b0 test(asm): seed the riscv64 and loong64 assemble fuzz targets
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin b9dfb48d79 fix(asm): encode the loong64 64-bit-span 2RI14 offsets like the toolchain
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 41aa8edfd2 fix(asm): match the toolchain's loong64 logical immediate expansion
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 8fdc511d0d test(asm): add the loong64 error-parity catalogue
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e5035d92e9 fix(asm): reject the offset on the loong64 register-indexed memory form
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin b54547c08d fix(asm): reject the shifted-register compositions on loong64
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e9510e8a68 test(asm): pin the riscv64 tail against the toolchain byte for byte
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 9d50212a71 fix(asm): refuse the riscv64 width moves across register banks
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin f2892e4f59 feat(asm): emit the riscv64 local-exec TLS sequence
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin f8dbd4f017 fix(asm): lower the riscv64 immediate CSR pseudos onto their opcode forms
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin e5aabf9801 fix(asm): encode the riscv64 FENCE predecessor and successor flags
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 36d1ac804a fix(asm): route the riscv64 register moves through the toolchain forms
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin da35883633 fix(asm): compress the riscv64 two-operand arithmetic and immediate tail
Assisted-by: GLM 5.3
2026-10-07 13:51:02 +02:00
petrbalvin 8bded3ea39 test(asm): pin the byte-form width reconciliation and its oracle
Table-driven rows for the renderer's spellings (the L suffix or none with
a byte register encodes the byte form, every register joining at its low
byte), the refusals (W, Q and the MOVD alias take no byte register) and a
differential kernel of the B-suffixed spellings assembled through both
gasm and go tool asm, byte for byte.

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 98a562d8b3 fix(asm): settle the byte-form width from the register operands
The suffixed scalar families derived the operand width from the mnemonic
alone, so a byte-spelled register under the L spelling or no suffix at all
encoded the widened form: XADDL DL, DL emitted 0F C1 where the byte form is
0F C0, CMPL AL, $7 emitted the 32-bit immediate form where the AL form is
3C 07, and CRC32 DL, R11 widened past the F0 byte opcode.  operandWidth now
reconciles the suffix with the operands: a byte register (AL, DL, R8B, ...)
forces the 8-bit form, which is the text the toolchain's own disassembly
prints for those encodings, while the W and Q spellings never ride a byte
register and are refused as go tool asm refuses them (MOVQ AL, AX).  The
shift count and the two- and three-operand IMUL forms stay out of the
reconciliation, and the byte accumulator short forms now belong to the AL
spelling alone, matching the toolchain's division (ADDB $3, AX is
80 c0 03, TESTB $7, AX is f6 c0 07).

Assisted-by: GLM 5.3
2026-10-07 13:49:58 +02:00
petrbalvin 33e7fdac98 fix(asm): enforce the arm64 TLBI, RPRFM, FCVT and integer-pair arities
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00
petrbalvin 7a69be8b59 fix(asm): reject the arm64 REGTMP spellings the toolchain refuses
Assisted-by: GLM 5.3
2026-10-07 13:49:49 +02:00