• v0.10.0 Stable

    petrbalvin released this 2026-07-15 15:13:28 +00:00 | 270 commits to main since this release

    The EVEX floating-point and conversion set: the packed-double arithmetic,
    the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
    verified byte for byte against the Go assembler.

    Added

    • asm: the rest of the common EVEX/VEX floating-point set — packed double
      arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
      VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
      VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
      family in both VEX and EVEX — the EVEX scalar forms exist for masked and
      zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
    • asm: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
      EVEX; the destination sets the length for PS→PD), the EVEX form of
      VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
      (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
      spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
      source — a new operand form, since the destination is always XMM while
      VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
      memory source).
    • asm: masking and zeroing on every new form — the scalar SD/SS
      arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
      explicit K1–K7 operand and the .Z suffix the way Go writes them.

    Documented

    • VCVTPS2PD follows the Go assembler's encoding, which omits the F3
      mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the
      Go toolchain's machine code is the project's byte-for-byte oracle, and
      gasm reproduces it exactly (and round-trips through the x86 decoder, which
      shares the convention).

    Verified

    • 58 new ground-truth cases — every instruction extracted from the Go
      toolchain's own assembly (go build + an executable-segment dump), checked
      byte for byte and round-tripped through the decoder, covering disp8×N for
      the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32)
      memory operands, the 5-bit register fields and the masked/zeroing P2
      byte. All four go-flac/go-lz4 kernels still assemble byte-identically
      and lint clean.
    Downloads