-
v0.6.0 Stable
released this
2026-07-11 15:36:52 +00:00 | 274 commits to main since this releaseCalibrated to the Go ABI:
register-clobberstops reporting legal code, and
the encoder learns the legacy SSE moves.Changed
lint:register-clobberis now calibrated to the Go ABI
(cmd/compile/abi-internal.md), not the platform ABI. Go's stack-based
ABI0 has no System V style callee-saved registers — amd64BX,R12–R15
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
scratch, and hand-written kernels may clobber them freely. The rule now
audits only the registers Go fixes across calls: the frame pointer and the
goroutine pointer (amd64BP/R14, arm64R18/R28/R29, riscv64
X27, loong64R22), and the goroutine pointer is reported only when the
function can reach the runtime (is notNOSPLITor makes a call) — the
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
may use it, exactly as the runtime's own assembly does. Both go-flac
kernels now lint with zero diagnostics.
Fixed
lint: the liveness analysis took the destination operand to be the
first operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last
on every architecture Go supports. The def/use and save/restore
classification on those architectures was inverted.format: a comment that follows aRET(typically the next function's doc
comment) is no longer indented as if it were still inside the finished
function body.
Added
asm: the legacy (non-VEX) SSE moves —MOVOU/MOVO(the Plan 9 names
for MOVDQU/MOVDQA),MOVUPS/MOVAPS/MOVUPD/MOVAPDand the scalar
MOVSD/MOVSS— andVMOVDQU64in the EVEX set. All verified byte for
byte against the Go assembler.
Downloads
-
v0.5.0 Stable
released this
2026-07-10 11:20:49 +00:00 | 275 commits to main since this releaseEVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage.Added
asm: EVEX (AVX-512) encoding — the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac
kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD,
VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32,
VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither.asm:AssembleFilenow accepts file-defined global (non-<>) symbols
too; a reference is external only when noGLOBLin the file defines it.
Fixed
asm: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
previously aVPBROADCASTD AX, Y30fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them.asm: the VEX encoder now rejects vector register indices 16–31 instead of
encoding a truncated (wrong) register.
Verified
- All 10 functions of the go-flac
avx512_amd64.skernel assemble
byte-identically to the Go toolchain's machine code (the disp32 of the one
VMOVDQU32 idx16(SB), Z13load is linker-filled in Go and resolved within
gasm's own image — checked to reach the right constant bytes). The AVX2
kernel's 17 functions remain byte-identical.
Downloads
-
v0.4.0 Stable
released this
2026-07-09 13:56:03 +00:00 | 276 commits to main since this releaseThe standalone assembler reaches the whole go-flac AVX2 kernel: static
symbols assemble, and all 17 kernel functions now match the Go toolchain's
machine code byte for byte.Added
asm: file-level assembly —AssembleFileturns a parsed file into an
Image: the function bodies in source order followed by a data section
built from the file'sGLOBL/DATAdirectives (each symbol 16-aligned).asm: static-symbol (SB) operands —mask<>(SB)references encode as
RIP-relative loads with a patched disp32, resolved against the image layout
so the output is self-consistent and position-independent. External
(non-file-local) symbols are rejected with a clear error: they need
object-file emission.gasm asmprints the data section and symbol map alongside the functions
and writes the whole image (code + data) with-o.
Verified
- All 17 functions of the go-flac
avx2_amd64.skernel assemble
byte-identically to the Go toolchain's machine code; the only differing
bytes are the displacements of the twoVMOVDQU mask24<>(SB), X15loads,
which the Go linker fills at link time and gasm resolves within its own
image (checked to reach the right constant bytes).
Downloads
-
v0.3.0 Stable
released this
2026-07-08 10:51:35 +00:00 | 277 commits to main since this releaseThe assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that
avoid global symbols now assembles to exactly the Go assembler's bytes (the
two holdouts load a file-local constant throughSBand wait on relocation
support).Added
asm: the scalar instruction families the kernels use —CMOVccand
SETcc(conditions spelled exactly like the jumps),LZCNT/TZCNT
(legacyF3 0F BD/BC), the sign/zero-extending moves (MOVBLZX,MOVBQZX,
MOVWLZX,MOVWQZX,MOVWLSX,MOVLQSX),CVTSL2SD/CVTSQ2SD(the
legacy SSE encoding, as the Go assembler emits it), the traditional
three-operandIMUL3{W,L,Q}, and the variable-count vector shifts
(VPSRLQ X0, Y8, Y8— the count in an XMM register or memory takes the
ordinary NDS form).asm: jump relaxation — jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating the
layout to a fixed point (CALL is always rel32).asm: jump-to-jump folding — a conditional jump to a label whose only
instruction is an unconditional jump is redirected to the ultimate
target, replicating the Go toolchain's linker, which chases such chains
before it encodes branches.parser: leading negative displacements with a base and index
(LEAQ -4(DX)(R9*4), R9) parse into a fully populated address.
Fixed
asm:CMPwith a register or memory operand computed second − first
instead of first − second, silently inverting every condition that followed
(CMPQ SI, R10; JGEtested R10 ≥ SI). The encoding now always records
first − second —CMP r/m, rwith the first operand in r/m,CMP r, r/m
with the first operand in reg — and is byte-identical to the Go assembler.asm: register-to-registerMOVnow uses ther/m ← ropcode (reg =
source), the Go assembler's choice; the output is byte-identical.
Downloads
-
v0.2.0 Stable
released this
2026-07-07 11:57:53 +00:00 | 278 commits to main since this releaseThe Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms.Added
asm: four new VEX (AVX/AVX2) operand forms, each validated by round-trip
decoding throughgolang.org/x/archand byte-for-byte against the
machine code the real Go assembler emits:- the immediate shuffle (
VPSHUFD,VPERMQ), - the three-operand-plus-immediate form (
VSHUFPD,VPERM2I128,
VINSERTI128), - the lane extract (
VEXTRACTI128,VEXTRACTF128— the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (
VMOVDQU,VMOVUPD,VMOVD,VMOVQ,
VMOVSD— each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler), - the no-operand
VZEROUPPER, andVPERMDin the NDS form, - the floating-point and FMA set (
VADDPD,VMULPD,VXORPD,
VUNPCKHPD, the scalarVADDSD/VMULSD,VCVTDQ2PD,VFMADD231PD).
With the scalar set and the earlier NDS / reg-rm / immediate-shift forms,
the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use.
- the immediate shuffle (
asm:CMPaccepts the immediate in the second operand position
(CMPL CX, $31) — the spelling the Go assembler accepts — encoding it
identically to the immediate-first form.
Fixed
asm: an unused VEX.vvvv field is now stored as1111(v̄vvv = 1111), as
the hardware requires — the previous value (0000) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite.
Downloads
-
v0.1.0 Stable
released this
2026-07-06 07:49:50 +00:00 | 279 commits to main since this releaseInitial release — the Phase 1 foundation.
Added
token,lexer,ast,parser: a hand-written, error-tolerant front end
for Plan 9 assembly. The lexer splices C-preprocessor line continuations
(\before a newline) so multi-line#definemacros parse as one opaque
directive. Validated against the production AVX2/AVX-512 kernels in
go-libraries/go-flacand the Go runtime'ssrc/runtime/*.sfor all four
architectures, with zero parse errors.arch: register files and complete instruction tables for amd64,
arm64, riscv64 and loong64, with the middle-dot symbol separator and static
(<>) symbols. Instruction names are generated from the Go toolchain's own
assembler source (just gen) — theanamesopcode lists plus the common
opcodes and the per-architecture front-end aliases (arm64B/BL, the
.P/.Waddressing suffixes, loong64JAL, the x86 conditional-jump
spellings) — so every mnemonic the real assembler accepts is recognised.lint: conservative rules —unknown-instruction,operand-count,
undefined-label,duplicate-label,missing-ret,
missing-textflag-include,abi-argsizeandunreachable-code. Macro
invocations are recognised (in-file#definenames and underscore
identifiers) and the label/RET heuristics are suppressed in macro-using
files.abi-argsizeparses the// funcsignature with the Go parser and
checks the declared TEXT argument size against Go's ABI0 layout;
unreachable-codeflags dead code afterRET, suppressed where reachability
is undecidable (PC-relative jumps, register-indirect branches,#ifdef).
Register liveness is computed by dataflow over the control-flow graph (basic
blocks, def/use, iterative backward iteration) and drivesregister-clobber,
an audit that flags a callee-saved register written but never saved/restored.
funcdata-pcdatavalidates the structure ofFUNCDATA/PCDATAdirectives.
Zero error-severity diagnostics across the 90-file Go runtime corpus and the
production go-flac kernels (theregister-clobberaudit additionally reports
the go-flac kernels' unsaved callee-saved register use for review).format: an idempotent canonical formatter (operand spacing and per-function
mnemonic alignment) that preserves comments and round-trips through the
parser.lsp: a Language Server Protocol server over stdio providing completion,
hover documentation, document symbols, publish-diagnostics and semantic-token
highlighting.asm: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift —
covering the bulk of the integer SIMD set) validated by round-trip decoding
againstgolang.org/x/arch, and an assembler that drives the parser's AST
into the encoder with local-label resolution andFP/SPframe mapping
(plus Go prologue/epilogue generation), producing output byte-identical to the
Go assembler for the supported operand forms.cmd/gasm: thegasmbinary withtokens,parse,fmt,lint,asm
andlspsubcommands._gen: the generator that rebuilds the architecture instruction tables from
the Go toolchain source (just gen).
Downloads