• v0.6.0 1a01870695

    v0.6.0 Stable

    petrbalvin released this 2026-07-11 15:36:52 +00:00 | 274 commits to main since this release

    Calibrated to the Go ABI: register-clobber stops reporting legal code, and
    the encoder learns the legacy SSE moves.

    Changed

    • lint: register-clobber is now calibrated to the Go ABI
      (cmd/compile/abi-internal.md), not the platform ABI. Go's stack-based
      ABI0 has no System V style callee-saved registers — amd64 BX, R12–R15
      and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
      scratch, and hand-written kernels may clobber them freely. The rule now
      audits only the registers Go fixes across calls: the frame pointer and the
      goroutine pointer (amd64 BP/R14, arm64 R18/R28/R29, riscv64
      X27, loong64 R22), and the goroutine pointer is reported only when the
      function can reach the runtime (is not NOSPLIT or makes a call) — the
      ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
      may use it, exactly as the runtime's own assembly does. Both go-flac
      kernels now lint with zero diagnostics.

    Fixed

    • lint: the liveness analysis took the destination operand to be the
      first operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last
      on every architecture Go supports. The def/use and save/restore
      classification on those architectures was inverted.
    • format: a comment that follows a RET (typically the next function's doc
      comment) is no longer indented as if it were still inside the finished
      function body.

    Added

    • asm: the legacy (non-VEX) SSE moves — MOVOU/MOVO (the Plan 9 names
      for MOVDQU/MOVDQA), MOVUPS/MOVAPS/MOVUPD/MOVAPD and the scalar
      MOVSD/MOVSS — and VMOVDQU64 in the EVEX set. All verified byte for
      byte against the Go assembler.
    Downloads
  • v0.5.0 458cfb626e

    v0.5.0 Stable

    petrbalvin released this 2026-07-10 11:20:49 +00:00 | 275 commits to main since this release

    EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
    the Go toolchain, completing the production-kernel coverage.

    Added

    • asm: EVEX (AVX-512) encoding — the four-byte EVEX prefix with the
      5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
      V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
      operands and as mask destinations, and the compressed disp8×N displacement
      (the multiplier follows the memory operand's size, as the Go assembler's
      opcode tables prescribe). Covers every AVX-512 instruction the go-flac
      kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD,
      VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32,
      VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the
      extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
      broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
      and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
      — the kernels use neither.
    • asm: AssembleFile now accepts file-defined global (non-<>) symbols
      too; a reference is external only when no GLOBL in the file defines it.

    Fixed

    • asm: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
      previously a VPBROADCASTD AX, Y30 fell into the VEX encoder, which cannot
      represent indices above 15 and silently truncated them.
    • asm: the VEX encoder now rejects vector register indices 16–31 instead of
      encoding a truncated (wrong) register.

    Verified

    • All 10 functions of the go-flac avx512_amd64.s kernel assemble
      byte-identically to the Go toolchain's machine code (the disp32 of the one
      VMOVDQU32 idx16(SB), Z13 load is linker-filled in Go and resolved within
      gasm's own image — checked to reach the right constant bytes). The AVX2
      kernel's 17 functions remain byte-identical.
    Downloads
  • v0.4.0 56ecc39539

    v0.4.0 Stable

    petrbalvin released this 2026-07-09 13:56:03 +00:00 | 276 commits to main since this release

    The standalone assembler reaches the whole go-flac AVX2 kernel: static
    symbols assemble, and all 17 kernel functions now match the Go toolchain's
    machine code byte for byte.

    Added

    • asm: file-level assembly — AssembleFile turns a parsed file into an
      Image: the function bodies in source order followed by a data section
      built from the file's GLOBL/DATA directives (each symbol 16-aligned).
    • asm: static-symbol (SB) operands — mask<>(SB) references encode as
      RIP-relative loads with a patched disp32, resolved against the image layout
      so the output is self-consistent and position-independent. External
      (non-file-local) symbols are rejected with a clear error: they need
      object-file emission.
    • gasm asm prints the data section and symbol map alongside the functions
      and writes the whole image (code + data) with -o.

    Verified

    • All 17 functions of the go-flac avx2_amd64.s kernel assemble
      byte-identically to the Go toolchain's machine code; the only differing
      bytes are the displacements of the two VMOVDQU mask24<>(SB), X15 loads,
      which the Go linker fills at link time and gasm resolves within its own
      image (checked to reach the right constant bytes).
    Downloads
  • v0.3.0 a82f575aee

    v0.3.0 Stable

    petrbalvin released this 2026-07-08 10:51:35 +00:00 | 277 commits to main since this release

    The assembler reaches byte-identical parity with the Go toolchain on the
    production go-flac AVX2 kernels: every one of the 15 kernel functions that
    avoid global symbols now assembles to exactly the Go assembler's bytes (the
    two holdouts load a file-local constant through SB and wait on relocation
    support).

    Added

    • asm: the scalar instruction families the kernels use — CMOVcc and
      SETcc (conditions spelled exactly like the jumps), LZCNT/TZCNT
      (legacy F3 0F BD/BC), the sign/zero-extending moves (MOVBLZX, MOVBQZX,
      MOVWLZX, MOVWQZX, MOVWLSX, MOVLQSX), CVTSL2SD/CVTSQ2SD (the
      legacy SSE encoding, as the Go assembler emits it), the traditional
      three-operand IMUL3{W,L,Q}, and the variable-count vector shifts
      (VPSRLQ X0, Y8, Y8 — the count in an XMM register or memory takes the
      ordinary NDS form).
    • asm: jump relaxation — jumps start in the short (rel8) form and
      expand to rel32 when the settled displacement does not fit, iterating the
      layout to a fixed point (CALL is always rel32).
    • asm: jump-to-jump folding — a conditional jump to a label whose only
      instruction is an unconditional jump is redirected to the ultimate
      target, replicating the Go toolchain's linker, which chases such chains
      before it encodes branches.
    • parser: leading negative displacements with a base and index
      (LEAQ -4(DX)(R9*4), R9) parse into a fully populated address.

    Fixed

    • asm: CMP with a register or memory operand computed second − first
      instead of first − second, silently inverting every condition that followed
      (CMPQ SI, R10; JGE tested R10 ≥ SI). The encoding now always records
      first − second — CMP r/m, r with the first operand in r/m, CMP r, r/m
      with the first operand in reg — and is byte-identical to the Go assembler.
    • asm: register-to-register MOV now uses the r/m ← r opcode (reg =
      source), the Go assembler's choice; the output is byte-identical.
    Downloads
  • v0.2.0 39870f91f6

    v0.2.0 Stable

    petrbalvin released this 2026-07-07 11:57:53 +00:00 | 278 commits to main since this release

    The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
    and the moves, on top of the Phase 1 VEX forms.

    Added

    • asm: four new VEX (AVX/AVX2) operand forms, each validated by round-trip
      decoding through golang.org/x/arch and byte-for-byte against the
      machine code the real Go assembler emits:
      • the immediate shuffle (VPSHUFD, VPERMQ),
      • the three-operand-plus-immediate form (VSHUFPD, VPERM2I128,
        VINSERTI128),
      • the lane extract (VEXTRACTI128, VEXTRACTF128 — the YMM source occupies
        the ModRM.reg field, the XMM/memory destination the r/m field),
      • the direction-sensitive moves (VMOVDQU, VMOVUPD, VMOVD, VMOVQ,
        VMOVSD — each direction picks its own opcode and VEX.W; a vector→vector
        move uses the store-form layout, matching the Go assembler),
      • the no-operand VZEROUPPER, and VPERMD in the NDS form,
      • the floating-point and FMA set (VADDPD, VMULPD, VXORPD,
        VUNPCKHPD, the scalar VADDSD/VMULSD, VCVTDQ2PD, VFMADD231PD).
        With the scalar set and the earlier NDS / reg-rm / immediate-shift forms,
        the encoder now covers every integer, shuffle and FP instruction the
        go-flac AVX2 kernels use.
    • asm: CMP accepts the immediate in the second operand position
      (CMPL CX, $31) — the spelling the Go assembler accepts — encoding it
      identically to the immediate-first form.

    Fixed

    • asm: an unused VEX.vvvv field is now stored as 1111 (v̄vvv = 1111), as
      the hardware requires — the previous value (0000) made the two-operand
      reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
      and differ from the Go assembler's bytes. The round-trip decoder ignores
      the field on these instructions, which is why the byte-for-byte Go
      comparison (added this release) is now part of the test suite.
    Downloads
  • v0.1.0 d5a4a6de45

    v0.1.0 Stable

    petrbalvin released this 2026-07-06 07:49:50 +00:00 | 279 commits to main since this release

    Initial release — the Phase 1 foundation.

    Added

    • token, lexer, ast, parser: a hand-written, error-tolerant front end
      for Plan 9 assembly. The lexer splices C-preprocessor line continuations
      (\ before a newline) so multi-line #define macros parse as one opaque
      directive. Validated against the production AVX2/AVX-512 kernels in
      go-libraries/go-flac and the Go runtime's src/runtime/*.s for all four
      architectures, with zero parse errors.
    • arch: register files and complete instruction tables for amd64,
      arm64, riscv64 and loong64, with the middle-dot symbol separator and static
      (<>) symbols. Instruction names are generated from the Go toolchain's own
      assembler source (just gen) — the anames opcode lists plus the common
      opcodes and the per-architecture front-end aliases (arm64 B/BL, the
      .P/.W addressing suffixes, loong64 JAL, the x86 conditional-jump
      spellings) — so every mnemonic the real assembler accepts is recognised.
    • lint: conservative rules — unknown-instruction, operand-count,
      undefined-label, duplicate-label, missing-ret,
      missing-textflag-include, abi-argsize and unreachable-code. Macro
      invocations are recognised (in-file #define names and underscore
      identifiers) and the label/RET heuristics are suppressed in macro-using
      files. abi-argsize parses the // func signature with the Go parser and
      checks the declared TEXT argument size against Go's ABI0 layout;
      unreachable-code flags dead code after RET, suppressed where reachability
      is undecidable (PC-relative jumps, register-indirect branches, #ifdef).
      Register liveness is computed by dataflow over the control-flow graph (basic
      blocks, def/use, iterative backward iteration) and drives register-clobber,
      an audit that flags a callee-saved register written but never saved/restored.
      funcdata-pcdata validates the structure of FUNCDATA/PCDATA directives.
      Zero error-severity diagnostics across the 90-file Go runtime corpus and the
      production go-flac kernels (the register-clobber audit additionally reports
      the go-flac kernels' unsaved callee-saved register use for review).
    • format: an idempotent canonical formatter (operand spacing and per-function
      mnemonic alignment) that preserves comments and round-trips through the
      parser.
    • lsp: a Language Server Protocol server over stdio providing completion,
      hover documentation, document symbols, publish-diagnostics and semantic-token
      highlighting.
    • asm: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/
      ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
      AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift —
      covering the bulk of the integer SIMD set) validated by round-trip decoding
      against golang.org/x/arch, and an assembler that drives the parser's AST
      into the encoder with local-label resolution and FP/SP frame mapping
      (plus Go prologue/epilogue generation), producing output byte-identical to the
      Go assembler for the supported operand forms.
    • cmd/gasm: the gasm binary with tokens, parse, fmt, lint, asm
      and lsp subcommands.
    • _gen: the generator that rebuilds the architecture instruction tables from
      the Go toolchain source (just gen).
    Downloads