chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
petrbalvin committed 2026-10-07 22:37:42 +02:00
1 parent e5df0a9d84
commit a77dd12d2d
13 files changed
+217 -77

No files matched your search

+9
View File
@@ -100,6 +100,15 @@ encoder backlog that `gasm audit-instructions` measures. The families:
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **The extension layer.** The families the toolchain's own table carries
late or not at all assemble through the extension mechanism: BF16,
VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA
families, the complex multiply and complex FMA pairs, VMINMAXPH and
VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The
layer takes `k0` to `k7` write masks with merging and zeroing, the
`{1toN}` broadcast, the embedded rounding and `{sae}` decorations and the
imm8 controls, and its refusals surface as `extension-form` lint errors
with the reason the encoder would give.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and