Commit Graph
432 Commits
Author SHA1 Message Date
petrbalvin aa9c7ca030 feat(arch): add the imm8 scalar FP16 controls to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:33:30 +02:00
petrbalvin 0354a1f4c1 feat(arch): scale and exponent-extract the scalar FP16 in the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 127e52de69 test(verify): accumulate the BF16 dot product on the metal
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 22055b9bf3 feat(arch): add the packed FP16 arithmetic to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin a2301b52de test(verify): run the amd64 extension encodings on the metal
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin bc37ea5b79 feat(arch): add the AVX512-FP16 scalar family to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 28bea95128 feat(arch): add the amd64 extended-instruction layer with BF16 and VP2INTERSECT
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 2d803e38d8 feat(lsp): leave the extension verdicts to the registry
Diagnostics no longer repeat the generated table's ignorance of a
registered mnemonic: unknown-instruction and unencodable-instruction
against a statement the registry encodes are filtered from the server's
own presentation, driven by asm.LookupExtension directly.  The filter is
a no-op once lint learns the registry, so the two compose unchanged.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:06:05 +02:00
petrbalvin dadeda144a feat(lsp): offer and document the registered extended mnemonics
Completion merges the extension registry's mnemonics beside the toolchain
entries, with operand shapes measured against the layer's own encoder, and
hover documents a registered mnemonic from its metadata: the extension
notice, the forms, the features, the fixed encodings and the manual
references.

Assisted-by: GLM 5.3 Flash
2026-10-07 00:06:05 +02:00
petrbalvin a90ec84bee fix(arch): narrow file names by go/build's suffix rule
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin c107b45933 fix(asm): reject duplicate symbol declarations like the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin a69f8cf4a8 fix(asm): match the toolchain's bytes across the corpus sweep
A line-for-line byte comparison of the whole amd64enc.s corpus against
go tool asm surfaced divergences the pass-only accounting never showed:
PEXTRW's GPR form swapped its fields, PUSHW took an imm32 where the
toolchain bounds the immediate to 16 bits, the double shift wrote the
unmasked register number into the reg field, VCOMISS carried a 0x66
prefix, RORX dropped the destination's R bit, and the variable bit
shifts used the manual's per-width opcodes where the toolchain
consolidates each row on one opcode with the W bit.  The VEX forms the
toolchain prefers for plain vector registers (the SSE2/SSSE3/SSE4.1
AVX twins, the compare-with-predicate family, VMOVUPS, VSHUFPS, the
variable shifts) now encode under VEX, with EVEX left to the ZMM,
opmask and index-16+ spellings, and the mnemonics whose rows never
offer the 2-byte prefix force it.  Every line is pinned through the new
corpus parity test (793 lines); the whole corpus file now assembles to
the toolchain's bytes at every commented line (10022 of 10022).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin cfc3abb752 feat(asm): encode the remaining amd64 VEX families
The SSE3 horizontal and add-subtract pairs, the SSSE3 sign and
horizontal integers, the masked moves in both directions, the
reciprocity and test pairs, the AVX imm8 tail (blends, dot products,
inserts, rounds, MPSADBW, the string compares), the four-operand
variable blends with their /is4 mask byte, the scalar three-operand
moves, the MXCSR accessors, the VPERMIL register controls and the
variable word shifts, plus the BMI2 count forms over memory.  Every
encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(852 lines); the /is4 byte carries the mask register number in its high
nibble, the layout the toolchain emits.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 6af3fd60d5 feat(asm): encode the legacy amd64 SSE and MMX families
The packed integer and float binaries, the imm8-controlled SSE4.1 forms,
the variable blends with their X0 mask, the high/low half moves, the
sign-mask extractions, the non-temporal stores, the MOVQ bank crossings
and their odd spellings, the MMX shifts and shuffle and the cache-line
mask stores, plus the scalar leaves LEAVE, INVPCID and the RTM controls.
Every encoding is pinned byte for byte against go tool asm through every
corpus line the toolchain's own amd64enc.s carries for the families
(1465 lines).  Two corpus-wide gaps fell out of the comparison: the
64-bit MOV immediate uses the zero-extending form across the unsigned
32-bit span, and the MMX-to-GPR MOVQ puts the bank register in reg.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 257feace6e feat(asm): encode the amd64 XSAVE family
XSAVE, XSAVEOPT, XSAVEC and XSAVES with their restore twins, plain and
64, each pinned byte for byte against go tool asm through every corpus
line the toolchain's own amd64enc.s carries for the family (24 lines).
The toolchain emits XSAVEOPT without the manual's 0x66 prefix; the bytes
are the oracle, so the family carries none.

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin f83bc8ddef feat(asm): encode the amd64 system, string and segment families
The no-operand flag and system controls, the sign-extension pair, the
string primitives, the multi-byte no-ops, the cache controls, MOVBE, the
compare-exchange doubles, the random source and FS/GS base pairs, the
descriptor-table accesses, the 0F 00/01 register controls and the
LAR/LSL selector reads and far-segment loads, each pinned byte for byte
against go tool asm through every corpus line the toolchain's own
amd64enc.s carries for the families (279 lines).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin f0b08ccea0 feat(asm): encode the amd64 x87 family
The x87 stack controls, the D8/DC arithmetic pair, the conditional moves,
the register compares, FADDDP, the memory loads and the FXSAVE pair, each
pinned byte for byte against go tool asm through every corpus line the
toolchain's own amd64enc.s carries for the family (78 lines).

Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 1040739fbb fix(asm): validate the loong64 ll/sc offset span like the toolchain
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:27 +02:00
petrbalvin 6cc6165c2b build(justfile): cap the fuzz recipe's workers
Assisted-by: GLM 5.3 Flash
2026-10-06 21:13:31 +02:00
petrbalvin 0768b4dc84 fix(cmd): locate GOROOT when the build is trimmed
Assisted-by: GLM 5.3 Flash
2026-10-06 20:24:39 +02:00
petrbalvin ae1baa0e61 ci: run the affordable gate set on push and publish at the tag
Test / test (push) Successful in 1m45s
Assisted-by: DeepSeek V4.1 Flash
2026-10-04 21:17:40 +02:00
petrbalvin 5a8e9acbf3 feat(arch): add the extended-instruction layer with SVE arithmetic
Test / test (push) Successful in 3m38s
2026-10-02 20:39:33 +02:00
petrbalvin 2747fce7d3 feat(asm): encode the arm64 system registers and structure loads 2026-10-02 20:39:26 +02:00
petrbalvin e02918c17b fix(ci): keep the push suite inside the runner's memory and time budget
Test / test (push) Successful in 2m55s
Assisted-by: GLM 5.3 Flash
2026-10-02 17:20:31 +02:00
petrbalvin 02a6359c1f ci: shrink the push pipeline to the affordable gate set
Test / test (push) Failing after 5m26s
Assisted-by: GLM 5.3 Flash
2026-10-02 16:47:52 +02:00
petrbalvin b4c1e133c0 ci: keep the GOOBJ link parity gate off the push pipeline
Test / test (push) Failing after 12m18s
Assisted-by: GLM 5.3 Flash
2026-10-02 16:15:02 +02:00
petrbalvin 1107928870 build(justfile): run the test recipes under the memory fence
Assisted-by: GLM 5.3 Flash
2026-10-02 16:15:02 +02:00
petrbalvin bc4ac93fd9 style(asm): reindent the evex comment gofmt asks for
Test / test (push) Failing after 21m29s
Assisted-by: GLM 5.3
2026-10-02 00:41:46 +02:00
petrbalvin c1bca7ce7e docs: record the development deltas in the changelog
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin fefb76beb9 docs(asm): correct the reference against the assemblers' behaviour
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin 42bc1669d7 feat(lsp): document directives on hover and widen completion
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin 69dcbec8ef feat(lint): eleven new rules over directives, data and addressing
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin f405cea5bc fix(cmd): stop the coverage run on a stray in-place trap
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin f57377abb9 fix(debug): handle mapping edges, stray traps and dying debuggees
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin dd1782c538 test(verify): gate GOOBJ link parity with cmd/link
Assisted-by: GLM 5.3
2026-10-02 00:40:54 +02:00
petrbalvin 17cc49fee4 fix(asm): emit NOPTR data as its own symbol kind
Assisted-by: GLM 5.3
2026-10-02 00:40:43 +02:00
petrbalvin bafb2fd130 feat(asm): encode the amd64 and loong64 tails of the corpus testdata
Assisted-by: GLM 5.3
2026-10-02 00:40:43 +02:00
petrbalvin 2f679326c2 style(testdata): canonicalise forms_amd64.s
Assisted-by: GLM 5.3
2026-10-02 00:40:20 +02:00
petrbalvin 96f2dd65b4 test(lexer): fuzz the token stream invariants
Assisted-by: GLM 5.3
2026-10-02 00:40:20 +02:00
petrbalvin ca887d3927 fix(parser): bound folding depth and macro expansion work
Assisted-by: GLM 5.3
2026-10-02 00:40:20 +02:00
petrbalvin 6570709226 fix(parser): peel stacked labels the way the formatter renders them
Assisted-by: GLM 5.3
2026-10-02 00:40:20 +02:00
petrbalvin 9a5217d9c1 fix(format): keep every token of a line in the canonical output
Assisted-by: GLM 5.3
2026-10-02 00:40:20 +02:00
petrbalvin 4d01bb3ecf build: rename the module to sourcedock.dev/petrbalvin/gasm-sdk
Test / test (push) Successful in 4m18s
2026-09-26 11:08:43 +02:00
petrbalvin 332c63e440 ci: compile-gate FreeBSD in the test pipeline
Test / test (push) Successful in 4m10s
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:49 +02:00
petrbalvin b306c210c6 feat(debug): port the debugger to FreeBSD
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:40 +02:00
petrbalvin b9015e1c2e fix(verify): make the executable mapping build on FreeBSD
Assisted-by: GLM 5.3 Flash
2026-09-25 21:46:31 +02:00
petrbalvin 26c5008136 fix(cmd): honour //go:build in the corpus audit
Test / test (push) Successful in 2m32s
Assisted-by: GLM 5.3 Flash
2026-09-23 21:03:16 +02:00
petrbalvin 74d6b90d69 fix(asm): read the arm64 move-wide immediate as an unsigned pattern
Assisted-by: GLM 5.3 Flash
2026-09-23 21:03:03 +02:00
petrbalvin 7b11c62f53 fix(asm): resolve negative numeric PC-relative jumps
Assisted-by: GLM 5.3 Flash
2026-09-23 21:02:50 +02:00
petrbalvin 8eed54b3da feat(lsp): quick fixes for the textflag include and the argument area
Test / test (push) Successful in 2m56s
Assisted-by: GLM 5.3 Flash
2026-09-23 20:23:35 +02:00