• v0.36.0
    Test / test (push) Successful in 32s
    Release / build (amd64, freebsd) (push) Successful in 52s
    Release / build (amd64, linux) (push) Successful in 13s
    Release / build (arm64, freebsd) (push) Successful in 51s
    Release / build (arm64, linux) (push) Successful in 29s
    Release / build (loong64, linux) (push) Successful in 29s
    Release / build (riscv64, linux) (push) Successful in 31s
    Release / release (push) Successful in 10s
    Stable

    petrbalvin released this 2026-10-07 20:44:47 +00:00 | 0 commits to main since this release

    Added

    • The SVE2 and SVE2.1 instruction sets on arm64. Over 350 SVE
      mnemonics across 537 encoded forms join the extension layer: the
      narrowing two-to-one arithmetic family, the SVE2 cryptographic set
      including ZADCLB, BFloat16 arithmetic, predicate counters and
      reductions, the pairwise and quadword forms, multiple-structure loads
      and stores (LD2/ST2 through LD4/ST4), shift by vector and the
      shift-immediate scheme with its element-size encoding, CLASTA and
      CLASTB, the last-active and compare predicate families, and the
      vector-length arithmetic ADDVL, ADDPL and RDVL. Every form derives
      from the toolchain's own encoder tables and is pinned against the
      GOROOT corpus words.
    • The complete AVX512-FP16 set and AVX-VNNI-INT16 on amd64. The
      extension layer reaches 191 entries: the packed and scalar FMA
      families (VFMADD, VFMSUB, VFMADDSUB and VFMSUBADD in the PH widths
      with their SH mirrors), the complex multiply and complex FMA pairs
      (VFMULC, VFCMULC, VFMADDC, VFCMADDC, packed and scalar, each
      conjugating the source its prefix names), VMINMAXPH and VMINMAXSH
      under their imm8 control, and the AVX-VNNI-INT16 dot products
      (VPDPWSUD, VPDPWSUDS, VPDPWUSD, VPDPWUSDS) through a new VEX
      encoding path beside the EVEX one. Every entry carries golden
      vectors byte-checked against GNU as.
    • Extension-layer instructions assemble straight from .s source on
      amd64.
      The registered mnemonics dispatch from the assembly front
      end with their operand grammar: k0 through k7 write masks with
      merging and zeroing, {1toN} broadcast, embedded rounding and
      {sae}, memory and scaled-index operands, and the imm8-control
      forms, the VEX forms beside the EVEX ones. gasm lint surfaces the
      layer's refusals as extension-form errors instead of letting a bad
      shape reach the encoder, and every distinct mnemonic is pinned from
      source text against the registry's bytes.
    • arm64 output byte-identical with the toolchain across the GOROOT
      tree.
      31 of the 36 files the differential harness walks (90
      functions, 27408 bytes) now assemble without a differing byte: the
      missing shapes are in (GETCALLERPC, the REM, REMW, UREM and UREMW
      family, DWORD, the FCVTHS, FCVTSH, FCVTDH and FCVTHD conversions, TLS
      local-exec loads as a single MOVZ with the TLS relocation), the
      literal pool drains mid-function when a function's literals would
      overflow the ±512 KiB load-literal range, and frame-relative
      addresses and flag-setting logicals materialise the way
      cmd/internal/obj/arm64 writes them.
    • riscv64 END and GETCALLERPC. END is accepted anywhere and emits
      nothing, and GETCALLERPC reads the return address the toolchain's
      rewrite describes: a leaf function reads LR, a framed body reads the
      prologue's save slot, with the compressed spellings where they fit.
    • The loong64 register-pair spelling. MULV R4:R5, R6 parses the
      way the toolchain's pair sugar does: the colon splits the operand and
      the halves swap, any instruction shape takes the spelling, and the
      malformed forms fail with the toolchain's own wording.
    • Fuzz targets across the tool. The assembler carries
      per-architecture FuzzAssemble targets and the front end carries
      targets over the lexer, parser, formatter and the extension
      registry; the seed corpora run as ordinary tests, and just fuzz
      drives the campaigns locally. Hardening the targets found caps GLOBL
      and DATA at 64 MiB per symbol and 128 MiB per section, so a crafted
      input cannot materialise unbounded memory.
    • The FreeBSD port of the debugger. gasm debug runs on FreeBSD on
      amd64 and arm64 with the same interactive surface as on Linux:
      breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
      register file), single-stepping, register and memory access, all behind
      the kernel's own ptrace requests, with tracee memory through PT_IO and
      stop reports through PT_LWPINFO. The JIT substrate maps executable
      memory through golang.org/x/sys/unix, so verify builds on FreeBSD
      too. The dispatched suite compile-gates both architectures, the
      release pipeline ships the two binaries, and live validation awaits a
      FreeBSD machine.
    • Workspace-wide navigation in the language server. gasm lsp indexes
      the .s files under the workspace root beyond the documents the editor
      has open, so go-to-definition, find references and workspace symbol search
      reach files that were never opened. An open buffer always shadows its
      disk copy, and watched-file events together with a per-query freshness
      check keep the index current.
    • Quick fixes for the textflag include and the argument area. The
      missing-textflag-include warning offers to add the include after the
      last one in the file, and the abi-argsize warning offers to set the
      TEXT argument area to the size the // func signature implies, computed
      by the new lint.ExpectedArgSize.
    • Eleven new lint rules over the directives, the data section and the
      sharpest addressing edges.
      missing-argsize flags a TEXT that
      declares no argument area its // func signature implies;
      noframe-frame-size a NOFRAME with a positive frame;
      unnamed-fp-reference a nameless 0(FP), which both assemblers
      reject; hardware-sp-addressing a negative offset off the hardware SP
      rather than the virtual frame; vex-sse-mixing a kernel that mixes VEX
      and legacy SSE spellings and pays the transition penalty;
      unnamed-result a ret+N(FP) the signature names; data-width,
      data-value-overflow, data-string-width, data-without-globl and
      data-exceeds-globl police the DATA width against its value type and
      the GLOBL size behind it. missing-ret now also flags a function
      whose tail can fall off its end even though a RET sits somewhere in the
      body, and invalid-textflag also reports a flag misplaced between TEXT
      and GLOBL.
    • Hover documentation and wider completions in the language server.
      Hovering a directive or pseudo-operation (TEXT, DATA, GLOBL, PCALIGN,
      FUNCDATA, PCDATA, the BYTE family) shows its grammar and rules in
      preference to the empty instruction-table entry, and completion offers
      those names beside the instruction set. The missing-argsize warning
      carries a quick fix that declares the argument area the signature
      implies ($0 becomes $0-16).
    • The GOOBJ link parity gate. A regression test assembles a kernel
      per architecture through gasm asm --format goobj, substitutes the
      object into a real go build's package archive, proves the archive
      carries it byte for byte and re-links with cmd/link on amd64, arm64,
      riscv64 and loong64. The binaries run, natively on amd64 and under
      qemu-user on arm64 and riscv64, and their output must match the
      toolchain-built baseline; loong64 is link-only. The pipeline installs
      qemu-user and runs the gate on every push.
    • The amd64 and loong64 encoders close four more corpus files. The
      whole-tree measure moves to 272 of 322 (84.5 %): amd64 gains the
      one-operand IMUL, the SSE compare family (CMPPD, CMPPS, CMPSS), RETFL,
      the LOOP family, the MMX register bank with its bank-crossing moves,
      MOVNTDQ, the CR and DR register moves, PUSH and POP of FS and GS, the
      (TLS) pseudo-base, the wait and cache controls (CLWB, CLDEMOTE,
      TPAUSE, UMONITOR, UMWAIT, RDPID, ENDBR64), indirect branches with the
      star spelling (JMP *(R12)(R13*4)), RET sym(SB) as the tail jump,
      the colon shift spelling (SHLL CX, R11:AX) and the EVEX and VEX forms
      of the rounds, AES key assist, string compares, extracts, blends and
      permutes; loong64 gains the acquire and release pair (LLACQ, SCREL,
      with the vector widths), the VMOVQ and XVMOVQ lane forms and the BYTE
      literal-data escape hatch the other architectures already take.

    Changed

    • TEXT flag operands behave like the toolchain's. A numeric or
      parenthesised flag list (TEXT ·f(SB), 4, $4096-0,
      (NOSPLIT|NOFRAME)) now suppresses the prologue and stack guard the
      names suppress, where the numbers were silently ignored and the
      prologue was emitted anyway; unknown flag names are rejected with the
      toolchain's wording instead of passing unnoticed; and the toolchain's
      own TEXT diagnostics fire, ABIInternal requires NOSPLIT and
      NOFRAME functions must have a frame size of 0.
    • The disassembler names what it used to print raw. The 195 amd64
      encodings that decoded to anonymous renders now decode to their
      instructions and round-trip, the BMI1 and BMI2 VEX set, RDSEED, RDPID,
      the WAITPKG family, ENDBR, CLDEMOTE, UD1 and RORX among them, with the
      arm64 HVC, SMC and SB forms and 74 loong64 rendering corrections
      beside.
    • Variadic macro parameters are rejected like the toolchain rejects
      them.
      A macro declaration whose parameter list the toolchain
      refuses no longer parses into a broken definition.
    • The corpus audit assembles like the build. A file's //go:build
      constraint decides which target architectures attempt it: cpu_x86.s is
      an x86 build alone, and the msan and goexperiment.runtimesecret trees
      are compiled by no supported build, so they leave the measured set
      instead of failing it. The headline now reads "assemble for every
      applicable target": every real-code GOROOT assembly file, the tree
      without testdata, assembles for all four architectures (250 of 250,
      100 %); over the whole tree including testdata the measure is 272 of
      322 (84.5 %).
    • The module moves to sourcedock.dev/petrbalvin/gasm-sdk. The
      repository and the module rename together with the product, now the
      GAsm Software Development Kit. Fresh installs become
      go install sourcedock.dev/petrbalvin/gasm-sdk/cmd/gasm@latest, and
      installs pinned to the old gasm-sdk path stop resolving once the
      repository takes the new name: reinstall from the new path. The
      binary stays gasm.

    Fixed

    • riscv64 memory offsets beyond the 12-bit immediate no longer
      truncate silently.
      LD 4096(X6), X5 encoded as LD X5, 0(X6) and
      read the wrong address; the expansion the toolchain performs (the
      upper bits into its temporary register, then the access) is emitted
      now, byte-identically, and offsets the toolchain refuses are refused
      with its wording.
    • riscv64 operand acceptance matches the toolchain's. 398 operand
      shapes the toolchain rejects assembled without complaint (a float
      register where an integer one is required, out-of-range CSRs, vector
      shapes against the wrong bank, MOV family widths), and 488 more
      produced the wrong diagnostic; all now fail with the toolchain's
      texts, and a 1030-case error-parity catalogue pins them against the
      toolchain's own negative files.
    • The debugger survives its failure paths. Killing a debuggee that
      had already died or sat ptrace-stopped could block the session
      forever; Kill now wakes the tracee, signals and reaps it without
      waiting on a parked child, so every launch-failure path returns
      instead of hanging the tool.
    • Encoder defects found by fuzzing and the differential audits. A
      GLOBL with a negative size panicked the assembler and a small crafted
      input materialised 3.92 GiB of DATA; certain SVE VTBL operands
      panicked; an over-strict guard refused BIC into RSP where the
      toolchain assembles it; byte-register operands now select the 8-bit
      forms the toolchain picks (XADDL DL, DL and its companions); and
      the loong64 shift compositions and indexed-offset forms that dropped
      fields re-encode faithfully.
    • Formatter edges around labels and macros. A stacked label that
      names a macro stays attached to it, macro content behind a label
      stays on its line, a selector-folded label resolves to its macro, an
      empty TEXT body formats, and a file ending in a comment no longer
      loses the comment.
    • Rename edits land in their own documents. A rename collected the
      ranges of every reference across the open documents but applied them all
      to the document that started it, so renaming a symbol used in a second
      file moved that file's text into the first. Each edit now applies to the
      document it was collected in.
    • Negative numeric PC-relative jumps. JMP -3(PC), the shape the
      runtime's exit loops write (sys_linux_amd64.s, sys_netbsd_amd64.s),
      resolved to nothing: only the forward forms counted. A negative count
      now walks the same instruction statements backwards, labels excluded,
      byte-identical with the toolchain.
    • The arm64 move-wide family reads its immediate as an unsigned
      pattern.
      MOVK $(40000<<48) folds to a negative int64 and was
      rejected; the toolchain picks the 16-bit lane from the 64-bit bit
      pattern, so the encoder now does the same, and a zero immediate is
      rejected where the toolchain rejects it.
    • NOPTR data emits its own symbol kind. A GLOBL with NOPTR was
      emitted as plain SDATA, the kind the linker holds to its Go type
      information requirement, so every gasm object carrying runtime-shaped
      data (GLOBL ·x(SB), NOPTR, ...) died in cmd/link with "missing Go
      type information". NOPTR data is now SNOPTRDATA, the toolchain's kind
      for pointer-free globals, and RODATA still wins where both flags
      appear, exactly as the toolchain chooses. The three end-to-end link
      tests that should have caught this substituted the gasm object after
      the package archive was already packed, so they passed vacuously; they
      now substitute inside the archive, prove the substitution byte for
      byte and re-link.
    • Latent encoder divergences against the toolchain, found by the
      whole-file differential harness.
      On loong64, MOVx $off(reg) lost
      its base register, SCQ swapped its operand fields, the vector lane
      inserts did not scale offsets by the element width and two immediate
      opcodes were mistyped; on amd64, NOP with operands encoded 0x90 where
      the toolchain emits nothing at all, the VEX gather length bit ignored
      the VSIB index width, four permute and extract families were pushed
      into EVEX where the toolchain stays VEX, and a reference to a static
      symbol no GLOBL defines failed where the toolchain defers it to the
      linker as an external relocation.
    • Seven front-end defects found by fuzzing. The formatter swallowed
      the statement after a leading block comment (/* head */ MOVQ AX, BX
      formatted to the comment alone), dropped trailing tokens after an
      #include header, and trimmed the trailing whitespace inside a string
      literal; the parser peeled only one label of a stacked pair (a: b:
      parsed differently than it formatted); deeply nested parentheses in an
      immediate overflowed the stack through the constant folder, which now
      stops at a bounded depth and falls back to the ordinary operand paths;
      macro expansion is linear in the invocation count instead of
      quadratic; and amplifying macros (the billion-laughs shape) stop at a
      work budget sized by the line and the macro table, reported as a
      diagnostic instead of running for hours.
    • The debugger handles the edges its tests now reach. Memory reads
      and writes cover exactly the requested bytes, so a request ending in
      the last page of a mapping no longer fails on the unmapped page behind
      it; disassembly shrinks its instruction window at a mapping's end
      instead of failing; a debuggee that dies before signalling readiness
      writes a failure notice the launcher reads, so launch fails fast with
      the reason instead of after the whole poll (and the poll no longer
      waits on the child, which could consume the SIGSTOP park and hang the
      launch); amd64 watchpoints acknowledge the sticky DR6 hit bits and
      clear the address register on release; the breakpoint listing is
      ordered by address so its numbering is stable; the REPL rejects bad
      counts, sizes and watchpoint types instead of silently guessing,
      reports stops on signals during stepping, and a breakpoint-class trap
      that matches no breakpoint and leaves the PC in place surfaces instead
      of spinning the continue loop and the coverage run forever.
    Downloads