-
v0.16.0 Stable
released this
2026-08-02 21:11:30 +00:00 | 263 commits to main since this releaseThe scalar conversions between vector and general-purpose registers — the
last of the amd64 EVEX instruction set.Added
asm: the GPR-interchanging conversions, byte for byte against the Go
assembler (28 ground-truth cases including memory sources and extended
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
vector source sits in vvvv (three Plan 9 operands).
Downloads
-
v0.15.0 Stable
released this
2026-07-20 14:05:08 +00:00 | 265 commits to main since this releaseThe last of the EVEX conversions and narrowing/extending moves — the EVEX
instruction set is now complete save for the GPR-interchanging forms.Added
asm: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
VCVTQQ2PS X/Y.asm: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and
VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned
narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB},
VPMOVDB, VPMOVQW).asm: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/
D2M/Q2M), whose K register is a genuine operand rather than a mask and
which therefore take no masking suffixes.
Downloads
-
v0.14.0 Stable
released this
2026-07-19 13:58:48 +00:00 | 266 commits to main since this releaseThe floating-point helper and conversion tail of the AVX-512 set, plus
gather and scatter with VSIB addressing — every encoding verified byte for
byte against the Go assembler.Added
asm: the floating-point helpers — reciprocals and reciprocal square
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask
destination).asm: gather and scatter with VSIB addressing. The gathers take
both Go spellings: the VEX form with a vector mask register (OP mask,
vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst),
where the EVEX L'L field follows the VSIB index register rather than the
data register (a ZMM index with an YMM destination encodes L'L = 10, as
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
only (OP src, K, vsib). All eight gather and eight scatter widths.asm: the remaining conversions — VCVTQQ2PS (the 512-bit source sets
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
Downloads
-
v0.13.0 Stable
released this
2026-07-18 13:47:59 +00:00 | 267 commits to main since this releaseThe wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every
encoding verified byte for byte against the Go assembler.Added
asm: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS —
a new NDS-plus-immediate form with the K register in the reg field), the
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts),
expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/
Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/
SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/
KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select),
the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the
aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
VPMOVQB).asm: the EVEX mnemonic suffixes the Go assembler accepts — the rounding
modes.RN_SAE,.RD_SAE,.RU_SAE,.RZ_SAE(the EVEX b bit with the
rounding control in L'L), suppress-all-exceptions.SAE, and memory
broadcast.BCST(the b bit, the vector length preserved, disp8×N scaled
by the element size) — each combinable with the.Zzeroing suffix,
validated against the Go assembler's bytes, and rejected on instructions
that do not support them.
Downloads
-
v0.12.0 Stable
released this
2026-07-17 16:57:04 +00:00 | 268 commits to main since this releaseGOOBJ emission: gasm-assembled functions drop into a
go buildwithout the
Go assembler.Added
asm: GOOBJ object output.gasm asm --format goobj -p <pkgpath>
writes the Go toolchain's own object format — the onecmd/linkconsumes
directly: the functions as non-package symbols qualified with the package
path (exactly ascmd/asmrecords assembly symbols), theGLOBLdata,
one serializedFuncInfoper function (argument/frame sizes, the asm
func flag, the start line, the file table) and the four pc-value tables
(pcsp,pcfile,pcline,pcinline). Thepcsptable carries the
real stack deltas: the assembler now tracks every stack-adjustment
boundary through the prologue (PUSHQ BP,SUBQ $frame, SP) and each
RET's epilogue, so frame-pointer functions unwind correctly. The
object preamble — the version-and-experiment header the linker compares
verbatim — is captured from the installedgo tool asm, so the output is
always consistent with the toolchain that links it.asm: relocations against file-localGLOBLsymbols becomeR_PCREL
entries in the GOOBJ output, with the instruction's displacement field
left zero for the linker to fill (ascmd/asmleaves it).
Fixed
parser: 64-bitDATAliterals aboveMaxInt64
(DATA mask<>+8(SB)/8, $0x800f…) parse as unsigned and keep their bit
pattern, instead of being rejected as non-integer.
Verified
- End-to-end: a gasm-emitted GOOBJ swapped into a
go buildin place of
the toolchain's assembly object links and runs with output identical to
the baseline binary (stack-argument calls and aGLOBLrelocation
resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as
a GOOBJ thatgo tool nmreads back with every symbol intact.
Downloads
-
v0.11.0 Stable
released this
2026-07-16 18:52:20 +00:00 | 269 commits to main since this releaseLinkable object output: external symbols and relocatable ELF / Mach-O
objects.Added
asm: object-file emission.gasm asm --format elfwrites an
ELF64 relocatable object and--format machoa Mach-O x86-64
MH_OBJECT: a code section (.text/__TEXT,__text) and a data
section (.data/__DATA,__data), a symbol table with one symbol per
TEXTandGLOBL(file-local<>symbols local, the rest global), and
one PC-relative relocation per static-symbol reference
(R_X86_64_PC32/X86_64_RELOC_SIGNED, the −4 addend the form needs).
The ELF output is verified end-to-end: a gasm-emitted object links with
a C driver and runs, resolving both a file-local constant and an
external symbol; the Mach-O output is verified structurally with
debug/macho.asm: external symbol references. A reference to a symbol no
GLOBLin the file defines no longer aborts assembly — it is recorded
as an external relocation (Image.Externals,FuncLayout.Relocs) and
becomes an undefined global symbol in the object output. The raw image
format (--format raw, the default) still reports them: only an object
file can represent a reference the linker must resolve.
Changed
gasm asmtakes a--format raw|elf|machoflag selecting what-o
writes; without--formatthe behaviour is unchanged (the concatenated
image).
Downloads
-
v0.10.0 Stable
released this
2026-07-15 15:13:28 +00:00 | 270 commits to main since this releaseThe EVEX floating-point and conversion set: the packed-double arithmetic,
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
verified byte for byte against the Go assembler.Added
asm: the rest of the common EVEX/VEX floating-point set — packed double
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
family in both VEX and EVEX — the EVEX scalar forms exist for masked and
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).asm: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
EVEX; the destination sets the length for PS→PD), the EVEX form of
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
source — a new operand form, since the destination is always XMM while
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
memory source).asm: masking and zeroing on every new form — the scalar SD/SS
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
explicit K1–K7 operand and the.Zsuffix the way Go writes them.
Documented
- VCVTPS2PD follows the Go assembler's encoding, which omits the F3
mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the
Go toolchain's machine code is the project's byte-for-byte oracle, and
gasm reproduces it exactly (and round-trips through the x86 decoder, which
shares the convention).
Verified
- 58 new ground-truth cases — every instruction extracted from the Go
toolchain's own assembly (go build + an executable-segment dump), checked
byte for byte and round-tripped through the decoder, covering disp8×N for
the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32)
memory operands, the 5-bit register fields and the masked/zeroing P2
byte. All four go-flac/go-lz4 kernels still assemble byte-identically
and lint clean.
Downloads
-
v0.9.0 Stable
released this
2026-07-14 19:03:26 +00:00 | 271 commits to main since this releaseAVX-512 masking and a wider EVEX integer set.
Added
asm: EVEX masking the way Go writes it — an explicitK1–K7
operand placed among the operands (merging mask), and a.Zmnemonic
suffix for zeroing (VPADDD.Z Z1, Z2, K2, Z3). Supported across the NDS,
reg/rm, immediate-shift, align, extract, convert and move forms, including
masked comparisons with a K destination (VPCMPEQD Z0, Z3, K2, K1). K0 is
rejected as an explicit mask, and.Zwithout a mask is an error, matching
the Go assembler.asm: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q,
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4]
in X̄). All verified byte for byte against the Go assembler.lint: masked EVEX forms (.Zsuffix, K operands) are recognised by
unknown-instructionand exempted fromoperand-count.
Fixed
asm: EVEX register–register operands with indices 16–31 encoded rm[4]
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
prefix bytes for X16+/Y16+ r/m operands.
Downloads
-
v0.8.0 Stable
released this
2026-07-13 17:50:38 +00:00 | 272 commits to main since this releaseStandard CLI ergonomics.
Added
gasm --helpprints a proper top-level help (description, commands,
flags, examples), and every subcommand now answers-h/--helpwith its
own usage block (usage line, description, flag defaults), exiting 0. An
unknown command points atgasm --helpinstead of dumping the whole usage.
Changed
- The version is primarily available as the standard
gasm --version/-V
flag; thegasm versionspelling remains as an alias.
Downloads
-
v0.7.0 Stable
released this
2026-07-12 19:24:41 +00:00 | 273 commits to main since this releaseThe formatter behaves like
go fmtand canonicalises block separation.Added
gasm fmtnow works likego fmt: with no arguments — or with a directory
argument — it reformats every.sfile below it in place and lists the
changed files, skipping.and_directories (.git,_refs, …).
Explicit file arguments keep the-w/ standard-output behaviour.
Changed
format: canonical blank-line layout — a new block (a label,TEXTor
GLOBL) is preceded by exactly one blank line, neither more nor less.
Comments leading a block stay with it (the blank line goes before them),
stacked labels share their block, the function's first label keeps hugging
itsTEXT, and runs of blank lines collapse to one. The output remains
idempotent and round-trips through the parser. All four go-flac/go-lz4
kernels were reformatted with this release and remain byte-identical when
assembled.
Downloads