-
v0.2.0 Stable
released this
2026-07-07 11:57:53 +00:00 | 278 commits to main since this releaseThe Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms.Added
asm: four new VEX (AVX/AVX2) operand forms, each validated by round-trip
decoding throughgolang.org/x/archand byte-for-byte against the
machine code the real Go assembler emits:- the immediate shuffle (
VPSHUFD,VPERMQ), - the three-operand-plus-immediate form (
VSHUFPD,VPERM2I128,
VINSERTI128), - the lane extract (
VEXTRACTI128,VEXTRACTF128— the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (
VMOVDQU,VMOVUPD,VMOVD,VMOVQ,
VMOVSD— each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler), - the no-operand
VZEROUPPER, andVPERMDin the NDS form, - the floating-point and FMA set (
VADDPD,VMULPD,VXORPD,
VUNPCKHPD, the scalarVADDSD/VMULSD,VCVTDQ2PD,VFMADD231PD).
With the scalar set and the earlier NDS / reg-rm / immediate-shift forms,
the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use.
- the immediate shuffle (
asm:CMPaccepts the immediate in the second operand position
(CMPL CX, $31) — the spelling the Go assembler accepts — encoding it
identically to the immediate-first form.
Fixed
asm: an unused VEX.vvvv field is now stored as1111(v̄vvv = 1111), as
the hardware requires — the previous value (0000) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite.
Downloads