Commit Graph
53 Commits
Author SHA1 Message Date
petrbalvin 7aba29ac67 test(arch): pin the golden vectors of the last four families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:42 +02:00
petrbalvin aeb109a64c feat(arch): the vector-length arithmetic pseudo group
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin bd85f8838d feat(arch): the SVE2.1 last-active vector and compare families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 12a5cfff52 feat(arch): the rest of the SVE2 BFloat16 wall
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin cdc3a75c88 feat(arch): the SVE multiple-structure loads and stores
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 82dbf087a8 feat(arch): the SVE2 BFloat16 arithmetic core
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 0760cc91af feat(arch): the SVE2 three-source and bitwise combine families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin bd80499c74 feat(arch): the SVE2 shift-by-vector family
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin b93fccc075 feat(arch): the SVE2.1 pairwise and quadword-reduction families
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 31ab584eee feat(arch): the SVE2.1 narrowing two-to-one family
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin ee2c6d51b3 fix(arch): keep Zdn out of the class bits of the predicated Z-alias source
Assisted-by: GLM 5.3 Flash
2026-10-07 20:35:41 +02:00
petrbalvin 4632ac1bb9 feat(arch): add the AVX-VNNI-INT16 dot products to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 96000dd64d feat(arch): add the VEX encoder to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 47d561b229 refactor(arch): extract the displacement tail and SIB builders
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 27ef71859b feat(arch): add VMINMAXSH to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 03f9ef0ac6 feat(arch): add the FP16 complex fused multiply-add to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 19:49:15 +02:00
petrbalvin 509afbb6c9 feat(arch): add the FP16 complex multiply and minimum-maximum to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 18:59:09 +02:00
petrbalvin 8080e0acef feat(arch): add the AVX512-FP16 FMA families to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 18:58:49 +02:00
petrbalvin d03de62c07 feat(arch): add the amd64 fp16 packed imm8-control group
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin 8de1b371da feat(arch): add the amd64 fp16 packed conversion family
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin fb6d01a7d0 feat(arch): encode the amd64 embedded rounding and SAE decorations
Assisted-by: GLM 5.3
2026-10-07 13:51:48 +02:00
petrbalvin e56c04e9ee fix(asm): read three operands from the arm64 last-element form
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin 104bea036b feat(asm): encode the arm64 SVE gather loads and scatter stores
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin 6faf850793 feat(asm): encode the arm64 SVE2 crypto, counter and reduction families
Assisted-by: GLM 5.3
2026-10-07 13:51:10 +02:00
petrbalvin 6c4932c4ec feat(arch): the SVE2.1 Z-alias permutations and copies
Assisted-by: GLM 5.3 Flash
2026-10-07 02:27:35 +02:00
petrbalvin 8231302bca feat(arch): the SVE predicate family in the extended layer
Assisted-by: GLM 5.3 Flash
2026-10-07 02:26:19 +02:00
petrbalvin c2adde948f feat(arch): add the write mask to the packed amd64 destinations
Assisted-by: GLM 5.3 Flash
2026-10-07 02:21:57 +02:00
petrbalvin 11cac26508 feat(arch): add the scaled index to the amd64 memory operands
Assisted-by: GLM 5.3 Flash
2026-10-07 02:06:11 +02:00
petrbalvin ccb155437e feat(arch): add the {1toN} broadcast to the packed amd64 memory sources
Assisted-by: GLM 5.3 Flash
2026-10-07 02:06:11 +02:00
petrbalvin 8d611bfdaf feat(arch): add the remaining scalar FP16 memory forms to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 01:35:48 +02:00
petrbalvin 454a21f5b7 feat(arch): add the packed FP16 and BF16 memory forms to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 01:35:48 +02:00
petrbalvin d275dee3ae feat(arch): add the scalar FP16 memory forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 01:35:48 +02:00
petrbalvin 7d69dda874 feat(arch): add the amd64 memory-operand mechanism to the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 01:35:48 +02:00
petrbalvin 103864e8b2 feat(arch): add the VL packed FP16 forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:37:41 +02:00
petrbalvin aa9c7ca030 feat(arch): add the imm8 scalar FP16 controls to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:33:30 +02:00
petrbalvin 0354a1f4c1 feat(arch): scale and exponent-extract the scalar FP16 in the extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 22055b9bf3 feat(arch): add the packed FP16 arithmetic to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin bc37ea5b79 feat(arch): add the AVX512-FP16 scalar family to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin 28bea95128 feat(arch): add the amd64 extended-instruction layer with BF16 and VP2INTERSECT
Assisted-by: GLM 5.3 Flash
2026-10-07 00:07:58 +02:00
petrbalvin a90ec84bee fix(arch): narrow file names by go/build's suffix rule
Assisted-by: GLM 5.3 Flash
2026-10-06 23:59:47 +02:00
petrbalvin 5a8e9acbf3 feat(arch): add the extended-instruction layer with SVE arithmetic
Test / test (push) Successful in 3m38s
2026-10-02 20:39:33 +02:00
petrbalvin 4d01bb3ecf build: rename the module to sourcedock.dev/petrbalvin/gasm-sdk
Test / test (push) Successful in 4m18s
2026-09-26 11:08:43 +02:00
petrbalvin e9789ce3f4 chore(arch): regenerate the instruction tables 2026-09-21 20:15:55 +02:00
petrbalvin b0f9071bf5 feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin ca3fdce0e0 feat(arm64): encode pairs, atomics, crypto, system and NEON slices
Assisted-by: GLM 5.3 Flash
2026-09-20 06:44:51 +02:00
petrbalvin fc2d92eabd feat(amd64): encode the GOROOT instruction families
Assisted-by: GLM 5.3 Flash
2026-09-20 06:44:51 +02:00
petrbalvin ebdf14939f fix(loong64): FP immediates through R30 and unsigned branch forms
Assisted-by: GLM 5.3
2026-09-19 23:49:13 +02:00
petrbalvin 401386956c fix(arm64): encode shifts, divides and multiplies and align sizes with emission
Assisted-by: GLM 5.3
2026-09-19 23:49:07 +02:00
petrbalvin 1691c81095 style: replace em and en dashes across sources 2026-09-14 18:22:18 +02:00
petrbalvin be2ceaafb9 feat(amd64): encode legacy SSE packed binaries and imm8 shuffles
Test / vet (push) Successful in 47s
Test / test (push) Successful in 2m34s
Test / build (push) Successful in 41s
2026-08-27 17:15:10 +02:00