• v0.16.0 Stable

    petrbalvin released this 2026-08-02 21:11:30 +00:00 | 263 commits to main since this release

    The scalar conversions between vector and general-purpose registers — the
    last of the amd64 EVEX instruction set.

    Added

    • asm: the GPR-interchanging conversions, byte for byte against the Go
      assembler (28 ground-truth cases including memory sources and extended
      GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
      in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
      (EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
      EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
      vector source sits in vvvv (three Plan 9 operands).
    Downloads
  • v0.15.0 Stable

    petrbalvin released this 2026-07-20 14:05:08 +00:00 | 265 commits to main since this release

    The last of the EVEX conversions and narrowing/extending moves — the EVEX
    instruction set is now complete save for the GPR-interchanging forms.

    Added

    • asm: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y
      spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
      VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
      VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
      VCVTQQ2PS X/Y.
    • asm: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and
      VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned
      narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB},
      VPMOVDB, VPMOVQW).
    • asm: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/
      D2M/Q2M), whose K register is a genuine operand rather than a mask and
      which therefore take no masking suffixes.
    Downloads
  • v0.14.0 Stable

    petrbalvin released this 2026-07-19 13:58:48 +00:00 | 266 commits to main since this release

    The floating-point helper and conversion tail of the AVX-512 set, plus
    gather and scatter with VSIB addressing — every encoding verified byte for
    byte against the Go assembler.

    Added

    • asm: the floating-point helpers — reciprocals and reciprocal square
      roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
      VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
      reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
      (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
      VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask
      destination).
    • asm: gather and scatter with VSIB addressing. The gathers take
      both Go spellings: the VEX form with a vector mask register (OP mask,
      vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst),
      where the EVEX L'L field follows the VSIB index register rather than the
      data register (a ZMM index with an YMM destination encodes L'L = 10, as
      the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
      only (OP src, K, vsib). All eight gather and eight scatter widths.
    • asm: the remaining conversions — VCVTQQ2PS (the 512-bit source sets
      the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
      VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
    Downloads
  • v0.13.0 Stable

    petrbalvin released this 2026-07-18 13:47:59 +00:00 | 267 commits to main since this release

    The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
    the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every
    encoding verified byte for byte against the Go assembler.

    Added

    • asm: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
      cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
      (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
      family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS —
      a new NDS-plus-immediate form with the K register in the reg field), the
      permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
      (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
      VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts),
      expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/
      Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/
      SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/
      KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select),
      the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the
      aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves
      (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
      remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
      VPMOVQB).
    • asm: the EVEX mnemonic suffixes the Go assembler accepts — the rounding
      modes .RN_SAE, .RD_SAE, .RU_SAE, .RZ_SAE (the EVEX b bit with the
      rounding control in L'L), suppress-all-exceptions .SAE, and memory
      broadcast .BCST (the b bit, the vector length preserved, disp8×N scaled
      by the element size) — each combinable with the .Z zeroing suffix,
      validated against the Go assembler's bytes, and rejected on instructions
      that do not support them.
    Downloads
  • v0.12.0 Stable

    petrbalvin released this 2026-07-17 16:57:04 +00:00 | 268 commits to main since this release

    GOOBJ emission: gasm-assembled functions drop into a go build without the
    Go assembler.

    Added

    • asm: GOOBJ object output. gasm asm --format goobj -p <pkgpath>
      writes the Go toolchain's own object format — the one cmd/link consumes
      directly: the functions as non-package symbols qualified with the package
      path (exactly as cmd/asm records assembly symbols), the GLOBL data,
      one serialized FuncInfo per function (argument/frame sizes, the asm
      func flag, the start line, the file table) and the four pc-value tables
      (pcsp, pcfile, pcline, pcinline). The pcsp table carries the
      real stack deltas: the assembler now tracks every stack-adjustment
      boundary through the prologue (PUSHQ BP, SUBQ $frame, SP) and each
      RET's epilogue, so frame-pointer functions unwind correctly. The
      object preamble — the version-and-experiment header the linker compares
      verbatim — is captured from the installed go tool asm, so the output is
      always consistent with the toolchain that links it.
    • asm: relocations against file-local GLOBL symbols become R_PCREL
      entries in the GOOBJ output, with the instruction's displacement field
      left zero for the linker to fill (as cmd/asm leaves it).

    Fixed

    • parser: 64-bit DATA literals above MaxInt64
      (DATA mask<>+8(SB)/8, $0x800f…) parse as unsigned and keep their bit
      pattern, instead of being rejected as non-integer.

    Verified

    • End-to-end: a gasm-emitted GOOBJ swapped into a go build in place of
      the toolchain's assembly object links and runs with output identical to
      the baseline binary (stack-argument calls and a GLOBL relocation
      resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as
      a GOOBJ that go tool nm reads back with every symbol intact.
    Downloads
  • v0.11.0 Stable

    petrbalvin released this 2026-07-16 18:52:20 +00:00 | 269 commits to main since this release

    Linkable object output: external symbols and relocatable ELF / Mach-O
    objects.

    Added

    • asm: object-file emission. gasm asm --format elf writes an
      ELF64 relocatable object and --format macho a Mach-O x86-64
      MH_OBJECT: a code section (.text / __TEXT,__text) and a data
      section (.data / __DATA,__data), a symbol table with one symbol per
      TEXT and GLOBL (file-local <> symbols local, the rest global), and
      one PC-relative relocation per static-symbol reference
      (R_X86_64_PC32 / X86_64_RELOC_SIGNED, the −4 addend the form needs).
      The ELF output is verified end-to-end: a gasm-emitted object links with
      a C driver and runs, resolving both a file-local constant and an
      external symbol; the Mach-O output is verified structurally with
      debug/macho.
    • asm: external symbol references. A reference to a symbol no
      GLOBL in the file defines no longer aborts assembly — it is recorded
      as an external relocation (Image.Externals, FuncLayout.Relocs) and
      becomes an undefined global symbol in the object output. The raw image
      format (--format raw, the default) still reports them: only an object
      file can represent a reference the linker must resolve.

    Changed

    • gasm asm takes a --format raw|elf|macho flag selecting what -o
      writes; without --format the behaviour is unchanged (the concatenated
      image).
    Downloads
  • v0.10.0 Stable

    petrbalvin released this 2026-07-15 15:13:28 +00:00 | 270 commits to main since this release

    The EVEX floating-point and conversion set: the packed-double arithmetic,
    the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
    verified byte for byte against the Go assembler.

    Added

    • asm: the rest of the common EVEX/VEX floating-point set — packed double
      arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
      VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
      VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
      family in both VEX and EVEX — the EVEX scalar forms exist for masked and
      zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
    • asm: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
      EVEX; the destination sets the length for PS→PD), the EVEX form of
      VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
      (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
      spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
      source — a new operand form, since the destination is always XMM while
      VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
      memory source).
    • asm: masking and zeroing on every new form — the scalar SD/SS
      arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
      explicit K1–K7 operand and the .Z suffix the way Go writes them.

    Documented

    • VCVTPS2PD follows the Go assembler's encoding, which omits the F3
      mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the
      Go toolchain's machine code is the project's byte-for-byte oracle, and
      gasm reproduces it exactly (and round-trips through the x86 decoder, which
      shares the convention).

    Verified

    • 58 new ground-truth cases — every instruction extracted from the Go
      toolchain's own assembly (go build + an executable-segment dump), checked
      byte for byte and round-tripped through the decoder, covering disp8×N for
      the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32)
      memory operands, the 5-bit register fields and the masked/zeroing P2
      byte. All four go-flac/go-lz4 kernels still assemble byte-identically
      and lint clean.
    Downloads
  • v0.9.0 0f3146ff2c

    v0.9.0 Stable

    petrbalvin released this 2026-07-14 19:03:26 +00:00 | 271 commits to main since this release

    AVX-512 masking and a wider EVEX integer set.

    Added

    • asm: EVEX masking the way Go writes it — an explicit K1–K7
      operand placed among the operands (merging mask), and a .Z mnemonic
      suffix for zeroing (VPADDD.Z Z1, Z2, K2, Z3). Supported across the NDS,
      reg/rm, immediate-shift, align, extract, convert and move forms, including
      masked comparisons with a K destination (VPCMPEQD Z0, Z3, K2, K1). K0 is
      rejected as an explicit mask, and .Z without a mask is an error, matching
      the Go assembler.
    • asm: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q,
      VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
      B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
      EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
      Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4]
      in X̄). All verified byte for byte against the Go assembler.
    • lint: masked EVEX forms (.Z suffix, K operands) are recognised by
      unknown-instruction and exempted from operand-count.

    Fixed

    • asm: EVEX register–register operands with indices 16–31 encoded rm[4]
      into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
      prefix bytes for X16+/Y16+ r/m operands.
    Downloads
  • v0.8.0 9370f9c3ee

    v0.8.0 Stable

    petrbalvin released this 2026-07-13 17:50:38 +00:00 | 272 commits to main since this release

    Standard CLI ergonomics.

    Added

    • gasm --help prints a proper top-level help (description, commands,
      flags, examples), and every subcommand now answers -h/--help with its
      own usage block (usage line, description, flag defaults), exiting 0. An
      unknown command points at gasm --help instead of dumping the whole usage.

    Changed

    • The version is primarily available as the standard gasm --version / -V
      flag; the gasm version spelling remains as an alias.
    Downloads
  • v0.7.0 e98680597d

    v0.7.0 Stable

    petrbalvin released this 2026-07-12 19:24:41 +00:00 | 273 commits to main since this release

    The formatter behaves like go fmt and canonicalises block separation.

    Added

    • gasm fmt now works like go fmt: with no arguments — or with a directory
      argument — it reformats every .s file below it in place and lists the
      changed files, skipping . and _ directories (.git, _refs, …).
      Explicit file arguments keep the -w / standard-output behaviour.

    Changed

    • format: canonical blank-line layout — a new block (a label, TEXT or
      GLOBL) is preceded by exactly one blank line, neither more nor less.
      Comments leading a block stay with it (the blank line goes before them),
      stacked labels share their block, the function's first label keeps hugging
      its TEXT, and runs of blank lines collapse to one. The output remains
      idempotent and round-trips through the parser. All four go-flac/go-lz4
      kernels were reformatted with this release and remain byte-identical when
      assembled.
    Downloads