-
v0.5.0 Stable
released this
2026-07-10 11:20:49 +00:00 | 275 commits to main since this releaseEVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage.Added
asm: EVEX (AVX-512) encoding — the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac
kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD,
VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32,
VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither.asm:AssembleFilenow accepts file-defined global (non-<>) symbols
too; a reference is external only when noGLOBLin the file defines it.
Fixed
asm: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
previously aVPBROADCASTD AX, Y30fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them.asm: the VEX encoder now rejects vector register indices 16–31 instead of
encoding a truncated (wrong) register.
Verified
- All 10 functions of the go-flac
avx512_amd64.skernel assemble
byte-identically to the Go toolchain's machine code (the disp32 of the one
VMOVDQU32 idx16(SB), Z13load is linker-filled in Go and resolved within
gasm's own image — checked to reach the right constant bytes). The AVX2
kernel's 17 functions remain byte-identical.
Downloads