feat(asm): add EVEX/AVX-512 encoding and assemble the AVX-512 kernel byte-identically

Assisted-by: Qwen 3.8 Max Preview
This commit is contained in:
2026-07-10 13:20:49 +02:00
parent 56ecc39539
commit 458cfb626e
9 changed files with 903 additions and 77 deletions
+11 -8
View File
@@ -206,13 +206,16 @@ memory destination r/m), the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`,
`VMULPD`, `VXORPD`, `VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`,
`VFMADD231PD`) and the no-operand `VZEROUPPER` — together with `VPERMD` and
the scalar families (`CMOVcc`, `SETcc`, `LZCNT`/`TZCNT`, the extending moves,
`CVTSx2SD`, `IMUL3`), covering every instruction the go-flac AVX2 kernels
use. Every encoding is validated two ways: by round-trip decoding through
`golang.org/x/arch`, and byte-for-byte against the machine code the real Go
assembler emits — a comparison that holds for the whole kernel: all 17
functions of the go-flac AVX2 file assemble to exactly the Go toolchain's
bytes, the lone exception being the displacements of the static-constant
loads, which the Go linker fills at link time.
`CVTSx2SD`, `IMUL3`) and the EVEX (AVX-512) prefix — the four-byte prefix with
5-bit register fields (Z0–Z31, X/Y 16–31), opmask registers as operands and
mask destinations, and the compressed disp8×N displacement, whose multiplier
follows the memory operand's size — covering every instruction the go-flac
AVX2 and AVX-512 kernels use. Every encoding is validated two ways: by
round-trip decoding through `golang.org/x/arch`, and byte-for-byte against
the machine code the real Go assembler emits — a comparison that holds for
whole functions: all 27 functions of both kernels assemble to exactly the Go
toolchain's bytes, the lone exception being the displacements of the
static-constant loads, which the Go linker fills at link time.
File-level assembly (`AssembleFile`) goes beyond single functions: it
materialises the file's static symbols (`GLOBL`/`DATA`) in a data section
@@ -220,7 +223,7 @@ behind the code and resolves references to them (`mask<>(SB)`) to
RIP-relative loads whose displacements point inside the resulting image, so
the bytes are self-consistent at any base address. External (non-file-local)
symbols are rejected: they need object-file emission, which — together with
EVEX / AVX-512 and the other architectures — is the rest of Phase 2.
EVEX masking/zeroing and the other architectures — is the rest of Phase 2.
## Extension points