petrbalvin
|
718181e7a7
|
feat(arch): the FP8-to-halfword conversion pairs
Assisted-by: GLM 5.3 Flash
|
2026-10-07 21:39:45 +02:00 |
|
petrbalvin
|
107ca512b2
|
feat(arch): the ZCNOT unary pair and the vector counter step
Assisted-by: GLM 5.3 Flash
|
2026-10-07 21:39:45 +02:00 |
|
petrbalvin
|
7aba29ac67
|
test(arch): pin the golden vectors of the last four families
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:42 +02:00 |
|
petrbalvin
|
aeb109a64c
|
feat(arch): the vector-length arithmetic pseudo group
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
bd85f8838d
|
feat(arch): the SVE2.1 last-active vector and compare families
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
12a5cfff52
|
feat(arch): the rest of the SVE2 BFloat16 wall
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
cdc3a75c88
|
feat(arch): the SVE multiple-structure loads and stores
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
82dbf087a8
|
feat(arch): the SVE2 BFloat16 arithmetic core
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
0760cc91af
|
feat(arch): the SVE2 three-source and bitwise combine families
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
bd80499c74
|
feat(arch): the SVE2 shift-by-vector family
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
b93fccc075
|
feat(arch): the SVE2.1 pairwise and quadword-reduction families
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
31ab584eee
|
feat(arch): the SVE2.1 narrowing two-to-one family
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
ee2c6d51b3
|
fix(arch): keep Zdn out of the class bits of the predicated Z-alias source
Assisted-by: GLM 5.3 Flash
|
2026-10-07 20:35:41 +02:00 |
|
petrbalvin
|
4632ac1bb9
|
feat(arch): add the AVX-VNNI-INT16 dot products to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 19:49:15 +02:00 |
|
petrbalvin
|
96000dd64d
|
feat(arch): add the VEX encoder to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 19:49:15 +02:00 |
|
petrbalvin
|
47d561b229
|
refactor(arch): extract the displacement tail and SIB builders
Assisted-by: GLM 5.3 Flash
|
2026-10-07 19:49:15 +02:00 |
|
petrbalvin
|
27ef71859b
|
feat(arch): add VMINMAXSH to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 19:49:15 +02:00 |
|
petrbalvin
|
03f9ef0ac6
|
feat(arch): add the FP16 complex fused multiply-add to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 19:49:15 +02:00 |
|
petrbalvin
|
509afbb6c9
|
feat(arch): add the FP16 complex multiply and minimum-maximum to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 18:59:09 +02:00 |
|
petrbalvin
|
8080e0acef
|
feat(arch): add the AVX512-FP16 FMA families to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 18:58:49 +02:00 |
|
petrbalvin
|
d03de62c07
|
feat(arch): add the amd64 fp16 packed imm8-control group
Assisted-by: GLM 5.3
|
2026-10-07 13:51:48 +02:00 |
|
petrbalvin
|
8de1b371da
|
feat(arch): add the amd64 fp16 packed conversion family
Assisted-by: GLM 5.3
|
2026-10-07 13:51:48 +02:00 |
|
petrbalvin
|
fb6d01a7d0
|
feat(arch): encode the amd64 embedded rounding and SAE decorations
Assisted-by: GLM 5.3
|
2026-10-07 13:51:48 +02:00 |
|
petrbalvin
|
e56c04e9ee
|
fix(asm): read three operands from the arm64 last-element form
Assisted-by: GLM 5.3
|
2026-10-07 13:51:10 +02:00 |
|
petrbalvin
|
104bea036b
|
feat(asm): encode the arm64 SVE gather loads and scatter stores
Assisted-by: GLM 5.3
|
2026-10-07 13:51:10 +02:00 |
|
petrbalvin
|
6faf850793
|
feat(asm): encode the arm64 SVE2 crypto, counter and reduction families
Assisted-by: GLM 5.3
|
2026-10-07 13:51:10 +02:00 |
|
petrbalvin
|
6c4932c4ec
|
feat(arch): the SVE2.1 Z-alias permutations and copies
Assisted-by: GLM 5.3 Flash
|
2026-10-07 02:27:35 +02:00 |
|
petrbalvin
|
8231302bca
|
feat(arch): the SVE predicate family in the extended layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 02:26:19 +02:00 |
|
petrbalvin
|
c2adde948f
|
feat(arch): add the write mask to the packed amd64 destinations
Assisted-by: GLM 5.3 Flash
|
2026-10-07 02:21:57 +02:00 |
|
petrbalvin
|
11cac26508
|
feat(arch): add the scaled index to the amd64 memory operands
Assisted-by: GLM 5.3 Flash
|
2026-10-07 02:06:11 +02:00 |
|
petrbalvin
|
ccb155437e
|
feat(arch): add the {1toN} broadcast to the packed amd64 memory sources
Assisted-by: GLM 5.3 Flash
|
2026-10-07 02:06:11 +02:00 |
|
petrbalvin
|
8d611bfdaf
|
feat(arch): add the remaining scalar FP16 memory forms to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 01:35:48 +02:00 |
|
petrbalvin
|
454a21f5b7
|
feat(arch): add the packed FP16 and BF16 memory forms to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 01:35:48 +02:00 |
|
petrbalvin
|
d275dee3ae
|
feat(arch): add the scalar FP16 memory forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 01:35:48 +02:00 |
|
petrbalvin
|
7d69dda874
|
feat(arch): add the amd64 memory-operand mechanism to the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 01:35:48 +02:00 |
|
petrbalvin
|
103864e8b2
|
feat(arch): add the VL packed FP16 forms to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:37:41 +02:00 |
|
petrbalvin
|
aa9c7ca030
|
feat(arch): add the imm8 scalar FP16 controls to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:33:30 +02:00 |
|
petrbalvin
|
0354a1f4c1
|
feat(arch): scale and exponent-extract the scalar FP16 in the extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:07:58 +02:00 |
|
petrbalvin
|
22055b9bf3
|
feat(arch): add the packed FP16 arithmetic to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:07:58 +02:00 |
|
petrbalvin
|
bc37ea5b79
|
feat(arch): add the AVX512-FP16 scalar family to the amd64 extension layer
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:07:58 +02:00 |
|
petrbalvin
|
28bea95128
|
feat(arch): add the amd64 extended-instruction layer with BF16 and VP2INTERSECT
Assisted-by: GLM 5.3 Flash
|
2026-10-07 00:07:58 +02:00 |
|
petrbalvin
|
a90ec84bee
|
fix(arch): narrow file names by go/build's suffix rule
Assisted-by: GLM 5.3 Flash
|
2026-10-06 23:59:47 +02:00 |
|
petrbalvin
|
5a8e9acbf3
|
feat(arch): add the extended-instruction layer with SVE arithmetic
Test / test (push) Successful in 3m38s
|
2026-10-02 20:39:33 +02:00 |
|
petrbalvin
|
4d01bb3ecf
|
build: rename the module to sourcedock.dev/petrbalvin/gasm-sdk
Test / test (push) Successful in 4m18s
|
2026-09-26 11:08:43 +02:00 |
|
petrbalvin
|
e9789ce3f4
|
chore(arch): regenerate the instruction tables
|
2026-09-21 20:15:55 +02:00 |
|
petrbalvin
|
b0f9071bf5
|
feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
|
2026-09-20 21:17:20 +02:00 |
|
petrbalvin
|
ca3fdce0e0
|
feat(arm64): encode pairs, atomics, crypto, system and NEON slices
Assisted-by: GLM 5.3 Flash
|
2026-09-20 06:44:51 +02:00 |
|
petrbalvin
|
fc2d92eabd
|
feat(amd64): encode the GOROOT instruction families
Assisted-by: GLM 5.3 Flash
|
2026-09-20 06:44:51 +02:00 |
|
petrbalvin
|
ebdf14939f
|
fix(loong64): FP immediates through R30 and unsigned branch forms
Assisted-by: GLM 5.3
|
2026-09-19 23:49:13 +02:00 |
|
petrbalvin
|
401386956c
|
fix(arm64): encode shifts, divides and multiplies and align sizes with emission
Assisted-by: GLM 5.3
|
2026-09-19 23:49:07 +02:00 |
|