Commit Graph
234 Commits
Author SHA1 Message Date
petrbalvin e43fa39dc9 feat(asm): encode the explicit riscv64 compressed instructions
The C extension's own spellings were names the table carried and the
encoder refused: CLWSP stopped the corpus audit's riscv64 file first.
Thirty-eight mnemonics now encode directly to their halfword, with the
toolchain's operand spellings and validation: the stack loads and stores
pin their base to SP, the register-based loads, stores and arithmetic
carry prime registers, CLUI refuses zero and SP, CADDI4SPN scales by
four, CADDI16SP by sixteen, and CJ, CBEQZ and CBNEZ resolve their N(PC)
targets against the final layout, taking a two-byte placeholder in the
early passes so the offsets stay honest.  CAND with an immediate is the
toolchain's C.ANDI spelling.  The toolchain's whole C extension testdata
block is a differential test, halfword for halfword, beside a range test
at the toolchain's own boundaries.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin ddfa33ccc1 feat(asm): encode the riscv64 bit-manipulation families
The Zba address generation, Zbb unary bit operations, Zbc carry-less
multiplication and Zbs single-bit families were names the table carried
and the encoder refused: thirty spellings plus RORI and XNOR fell over.
The register and immediate forms now encode as the toolchain does, the
unary operations carry their fixed rs2 constant, RORI lowers to ROR's
expansion (its reverse shift compressing like ROR's), XNOR XORs and
inverts in place, and ROL/ROLW rotate left through the same temporary
the toolchain uses, taking a register amount only as its own expansion
requires.  The toolchain's whole testdata block for these families is
now a differential test: every word must agree byte for byte.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin d82ef33fa9 fix(asm): bound the riscv64 shift immediate at the instruction width
SLLI $64 assembled with the amount silently masked into the six-bit
field where the toolchain rejects it, and the word forms took 0-63 where
they take 0-31.  Both families now validate against their own width and
the check reads the immediate at full width, so a value the source
spelled beyond int32 cannot wrap into the range; the boundary is pinned
in a test.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin c3540f0549 fix(asm): read the riscv64 raw-data immediates at full width
WORD and BYTE read their immediate through the truncating helper, so the
int32 wrap turned WORD $0xffffffff into -1 and rejected it, while WORD
$0x100000000 and BYTE $0x100000001 arrived pre-truncated and slipped
past the range check as small values.  Both statements now read the
immediate the source wrote and bound it at the toolchain's own limits:
[0, 0xffffffff] for WORD, [0, 0xff] for BYTE, with the bounds pinned in
a test.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin bffe408afa feat(asm): resolve every CSR name the toolchain knows
The riscv64 assembler carried fourteen hand-picked CSR names where the
toolchain resolves three hundred and twenty-nine: a CSRR/CSRW family
instruction naming any privileged register beyond the few base ones came
out as unknown CSR.  The table now carries the RISC-V privileged
specification's register set exactly as go tool asm spells it, and a
differential test assembles every name through both assemblers and
requires the words to agree byte for byte.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:47:27 +02:00
petrbalvin 103864e8b2 feat(arch): add the VL packed FP16 forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:37:41 +02:00
petrbalvin aa9c7ca030 feat(arch): add the imm8 scalar FP16 controls to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:33:30 +02:00
petrbalvin 0354a1f4c1 feat(arch): scale and exponent-extract the scalar FP16 in the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 22055b9bf3 feat(arch): add the packed FP16 arithmetic to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin bc37ea5b79 feat(arch): add the AVX512-FP16 scalar family to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 28bea95128 feat(arch): add the amd64 extended-instruction layer with BF16 and VP2INTERSECT
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin c107b45933 fix(asm): reject duplicate symbol declarations like the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin a69f8cf4a8 fix(asm): match the toolchain's bytes across the corpus sweep
A line-for-line byte comparison of the whole amd64enc.s corpus against
go tool asm surfaced divergences the pass-only accounting never showed:
PEXTRW's GPR form swapped its fields, PUSHW took an imm32 where the
toolchain bounds the immediate to 16 bits, the double shift wrote the
unmasked register number into the reg field, VCOMISS carried a 0x66
prefix, RORX dropped the destination's R bit, and the variable bit
shifts used the manual's per-width opcodes where the toolchain
consolidates each row on one opcode with the W bit.  The VEX forms the
toolchain prefers for plain vector registers (the SSE2/SSSE3/SSE4.1
AVX twins, the compare-with-predicate family, VMOVUPS, VSHUFPS, the
variable shifts) now encode under VEX, with EVEX left to the ZMM,
opmask and index-16+ spellings, and the mnemonics whose rows never
offer the 2-byte prefix force it.  Every line is pinned through the new
corpus parity test (793 lines); the whole corpus file now assembles to
the toolchain's bytes at every commented line (10022 of 10022).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin cfc3abb752 feat(asm): encode the remaining amd64 VEX families
The SSE3 horizontal and add-subtract pairs, the SSSE3 sign and
horizontal integers, the masked moves in both directions, the
reciprocity and test pairs, the AVX imm8 tail (blends, dot products,
inserts, rounds, MPSADBW, the string compares), the four-operand
variable blends with their /is4 mask byte, the scalar three-operand
moves, the MXCSR accessors, the VPERMIL register controls and the
variable word shifts, plus the BMI2 count forms over memory.  Every
encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(852 lines); the /is4 byte carries the mask register number in its high
nibble, the layout the toolchain emits.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 6af3fd60d5 feat(asm): encode the legacy amd64 SSE and MMX families
The packed integer and float binaries, the imm8-controlled SSE4.1 forms,
the variable blends with their X0 mask, the high/low half moves, the
sign-mask extractions, the non-temporal stores, the MOVQ bank crossings
and their odd spellings, the MMX shifts and shuffle and the cache-line
mask stores, plus the scalar leaves LEAVE, INVPCID and the RTM controls.
Every encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(1465 lines).  Two corpus-wide gaps fell out of the comparison: the
64-bit MOV immediate uses the zero-extending form across the unsigned
32-bit span, and the MMX-to-GPR MOVQ puts the bank register in reg.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 257feace6e feat(asm): encode the amd64 XSAVE family
XSAVE, XSAVEOPT, XSAVEC and XSAVES with their restore twins, plain and
64, each pinned byte for byte against go tool asm through every corpus
line the toolchain's own amd64enc.s carries for the family (24 lines).
The toolchain emits XSAVEOPT without the manual's 0x66 prefix; the bytes
are the oracle, so the family carries none.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin f83bc8ddef feat(asm): encode the amd64 system, string and segment families
The no-operand flag and system controls, the sign-extension pair, the
string primitives, the multi-byte no-ops, the cache controls, MOVBE, the
compare-exchange doubles, the random source and FS/GS base pairs, the
descriptor-table accesses, the 0F 00/01 register controls and the
LAR/LSL selector reads and far-segment loads, each pinned byte for byte
against go tool asm through every corpus line the toolchain's own
amd64enc.s carries for the families (279 lines).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin f0b08ccea0 feat(asm): encode the amd64 x87 family
The x87 stack controls, the D8/DC arithmetic pair, the conditional moves,
the register compares, FADDDP, the memory loads and the FXSAVE pair, each
pinned byte for byte against go tool asm through every corpus line the
toolchain's own amd64enc.s carries for the family (78 lines).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 1040739fbb fix(asm): validate the loong64 ll/sc offset span like the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:27 +02:00
petrbalvin 5a8e9acbf3 feat(arch): add the extended-instruction layer with SVE arithmetic
Test / test (push) Successful in 3m38s
2026-10-02 20:39:33 +02:00
petrbalvin 2747fce7d3 feat(asm): encode the arm64 system registers and structure loads 2026-10-02 20:39:26 +02:00
petrbalvin e02918c17b fix(ci): keep the push suite inside the runner's memory and time budget
Test / test (push) Successful in 2m55s
Assisted-by: GLM 5.3 Flash
2026-10-02 17:20:31 +02:00
petrbalvin bc4ac93fd9 style(asm): reindent the evex comment gofmt asks for
Test / test (push) Failing after 21m29s
Assisted-by: GLM 5.3
2026-10-02 00:41:46 +02:00
petrbalvin 17cc49fee4 fix(asm): emit NOPTR data as its own symbol kind
Assisted-by: GLM 5.3
2026-10-02 00:40:43 +02:00
petrbalvin bafb2fd130 feat(asm): encode the amd64 and loong64 tails of the corpus testdata
Assisted-by: GLM 5.3
2026-10-02 00:40:43 +02:00
petrbalvin 4d01bb3ecf build: rename the module to sourcedock.dev/petrbalvin/gasm-sdk
Test / test (push) Successful in 4m18s
2026-09-26 11:08:43 +02:00
petrbalvin 74d6b90d69 fix(asm): read the arm64 move-wide immediate as an unsigned pattern
Assisted-by: GLM 5.3 Flash
2026-09-23 21:03:03 +02:00
petrbalvin 7b11c62f53 fix(asm): resolve negative numeric PC-relative jumps
Assisted-by: GLM 5.3 Flash
2026-09-23 21:02:50 +02:00
petrbalvin f720381d43 feat(asm): the segment-absolute and crash-store forms GOROOT writes
Test / test (push) Failing after 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-21 22:19:53 +02:00
petrbalvin 2c9042d62c feat(asm): PCALIGN alignment on amd64
Assisted-by: GLM 5.3 Flash
2026-09-21 22:00:30 +02:00
petrbalvin 82ef289d3a feat(asm): the immediate multiply and arm64 indirect branches GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:50:11 +02:00
petrbalvin 7246b0e002 feat(asm): the TLS access pair in the toolchain's one-instruction form
Assisted-by: GLM 5.3 Flash
2026-09-21 21:35:15 +02:00
petrbalvin 8cfd40aac8 feat(asm): the operand forms and defines GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:17:34 +02:00
petrbalvin 29ac03468e feat(amd64): floating-point immediates through a synthesised pool
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin bfb7701db1 feat(amd64): emit the quad-register EVEX families
Assisted-by: GLM 5.3 Flash
2026-09-21 02:02:19 +02:00
petrbalvin e8b6ff5d7c fix(parser): fold a signed parenthesised displacement expression
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-21 00:45:33 +02:00
petrbalvin 1456907000 feat(riscv64,loong64): operand tail, float DATA and honest port classification
Assisted-by: GLM 5.3 Flash
2026-09-21 00:44:47 +02:00
petrbalvin 522e6f2ae8 feat(parser): bracket register ranges, index-only VSIB and bare trailing immediates
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 81d4bd81e4 test(verify): register the wave kernels
Test / test (push) Failing after 2m20s
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:31 +02:00
petrbalvin 687678a2ea feat(elf): emit data relocations on arm64, riscv64 and loong64
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin b0f9071bf5 feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 81e2673923 feat(amd64): encode the AVX-512 and BMI corpus families
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 241e7256f6 fix(arm64): reject bare BTI with a diagnostic and accept the full family
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 6556b85abf feat(asm): symbol-valued DATA, division slash in symbols and plain semicolons
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 289cabe993 feat(loong64): encode the full LSX and LASX table
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin dce5d31462 feat(amd64): LOCK and REP prefixes, literal data pseudo-ops and ADJSP
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9dc3987e02 feat(riscv64,loong64): PCALIGN, branch relaxation and operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9b238a525a feat(arm64): wide immediates, SIMD compare and system operand forms
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 0629f5e2df feat(arm64): assemble PCALIGN padding and BYTE literal bytes
Test / test (push) Successful in 2m16s
Assisted-by: GLM 5.3 Flash
2026-09-20 11:49:05 +02:00
petrbalvin 6c672567f3 feat(amd64): assemble the double-shift and static-SB operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00