• v0.34.0
    Test / test (push) Successful in 2m11s
    Release / gates (push) Successful in 2m11s
    Release / build (amd64, linux) (push) Successful in 1m13s
    Release / build (arm64, linux) (push) Successful in 1m10s
    Release / build (loong64, linux) (push) Successful in 1m12s
    Release / build (riscv64, linux) (push) Successful in 1m33s
    Release / release (push) Successful in 58s
    Stable

    petrbalvin released this 2026-09-19 23:44:37 +00:00 | 54 commits to main since this release

    Added

    • Indirect JMP and CALL on all four architectures. JMP AX,
      CALL AX, JMP (BX) and the memory forms encode at byte parity with
      the toolchain (FF /2 and FF /4 on amd64); arm64 lowers JMP (R0) to
      BR and accepts the raw BR/BLR spellings; riscv64 lowers JMP (X5) to
      JALR; loong64 accepts the raw JIRL rd, rj, off spelling the Go
      assembler cannot express. A frameless amd64 function containing a
      CALL now receives the toolchain's forced base-pointer frame. The
      riscv64 and loong64 verify trampolines join their ground-truth lists,
      and a lint check for control flow through registers and memory extends
      to the new forms.
    • gasm asm -GOARCH and gasm diff -GOARCH. The target
      architecture can be named explicitly instead of inferred from the
      file-name suffix, which is how the suffix-less majority of GOROOT's
      .s files (cpu_x86.s, stub.s, ...) become assemblable.
    • gasm audit-instructions --corpus [dir]. Assembles every .s
      file under a directory (default GOROOT/src) with the gasm encoder
      only: suffixed files for their architecture, suffix-less files for
      all four, as a GOARCH build would. Reports the headline number (127
      of 627 GOROOT files, 20.3 %, assemble for every target architecture),
      the per-architecture pass rates and the most common failure reasons
      with a representative file each,
      which drive the encodability backlog by frequency.
    • The parser and the formatter are fuzzed. Two targets carry the
      guarantee: no input makes the parser panic, and every input yields a
      file the rest of the toolkit can work on; formatting twice equals
      formatting once, and clean input stays clean. They seed from the
      repository's own kernels, and just fuzz drives the mutation engine
      on demand.
    • Man pages. docs/man carries gasm(1) and a page for every command
      except version, which gasm(1) documents itself, written in roff:
      synopsis, description, every flag with its default, exit status,
      worked examples and cross-references.
      just install-man compresses them into ~/.local/share/man (MANDIR
      overrides) and just uninstall-man removes them. A test builds the
      binary and compares each page's flags and synopsis with its own -h
      output, so the pages cannot drift from the CLI.

    Changed

    • Go 1.27.1 required. The module declares go 1.27.1, so building
      from source needs that patch release or newer.
    • Canonical just recipes. just gates is the definition of done
      (build, fmt-check, vet, test, race). install now builds and copies
      the binary into ~/.local/bin (BINDIR overrides) instead of
      downloading module dependencies, and install-bin is gone. The test
      gate sweeps the logic packages (arch through verify; the ptrace-bound
      debug and the thin cmd/gasm sit outside the coverage profile), so
      the coverage floor is computed over the product code and the number is
      identical locally and in CI; the two excluded packages' own tests run
      in the gate and in the pipelines, outside the floor. fuzz requires
      its target package.
    • The reported version comes from the build. gasm --version
      prints the version the toolchain recorded: the tag on a tagged
      checkout, a pseudo-version naming the commit below one, +dirty on a
      dirty tree and (devel) outside version control. Nothing is
      injected with -ldflags -X any more.
    • CI realigned with the gate set. The push pipeline runs the gates
      minus race in one job, in the gates order, with a cached Go setup and
      the module as the version source; a superseded run of the same branch
      is cancelled instead of queueing; every go test runs under a
      ten-minute bound that matches its job's; the race detector moved to a
      hand-dispatched workflow and runs in the local gate before a tag is
      cut, never on a push or a tag; the release builds without injection and
      its smoke test requires the recorded tag and rejects +dirty.
    • The documents follow the standard set. docs/ARCHITECTURE.md is
      organised as Overview, Packages, Data flow, State and lifetime and
      Dependencies, and carries a sequence diagram of the assembly path;
      docs/DEVELOPMENT.md lists every recipe in one table and documents the
      coverage floor, the CI and the release flow; docs/CLI.md gives the
      synopsis, the commands, every flag with its default, the exit codes and
      worked examples; CONTRIBUTING.md carries the Contributor terms and
      states the commit trailer form, the one-logical-change rule and the
      licence header rule; SECURITY.md states how a vulnerability is
      reported and what to expect. The repository's own assembly (the
      verify trampolines and the test kernels) is in gasm fmt canonical
      form.
    • The README states the project's purpose and status. It opens with
      a warning that the tool is an experiment under active development,
      version 0.x.x, free to change without warning, with 1.0.0 far off,
      and already in active use on real assembly work. It describes both
      goals (tooling for Plan 9 assembly, and Plan 9 assembly outside the
      Go toolchain), argues the case for the syntax in a new Why Plan 9
      assembly section, and carries a Direction section: extended
      instruction support, full GOOBJ and ELF compilation, Linux and
      FreeBSD, and the four architectures. A Validation status section
      states what has been executed where: amd64 on real hardware, the other
      three architectures under qemu-user emulation, the encoding parity on
      the host for all four, and the debugger's ptrace path on amd64 only.

    Fixed

    • riscv64 JALR silently jumped to the wrong register. The trampoline
      form JALR X0, 0(X5) read the memory operand's base as the destination,
      encoding a jump to X0 with no diagnostic; the destination is the first
      operand. The leaf detection shared the confusion, so affected functions
      also grew a bogus prologue. JALR X0, 0(X1) as written in the verify
      trampoline was mis-encoded since its introduction.
    • DATA lines demanded their GLOBL first. collectData processed the
      declarations in file order, but the Plan 9 convention puts every DATA
      line before its symbol's GLOBL; correctly ordered files (most of
      GOROOT's) failed with "no matching GLOBL". Two passes: symbols are
      registered before initialisers are applied.
    • The formatter lost idempotency on degenerate lines. Illegal tokens
      survived into the output, a label sharing its line with a
      non-instruction split into a line the parser rejects, stray-operand
      lines entered the alignment width computation, and rendered / *,
      > > sequences re-lexed as comments and shifts. The label, width and
      spacing rules now agree between passes.
    • gasm fmt deleted the | separators from TEXT and GLOBL flag
      lists.
      The lexer had no token for |, the formatter dropped the
      resulting illegal token, and an in-place format silently rewrote
      NOSPLIT|DUPOK as NOSPLIT DUPOK, which the Go assembler rejects.
      The bars now round-trip byte-identically, and ·foo<ABIInternal>(SB)
      parses its ABI marker instead of swallowing ABIInternal and SB
      into the flags, which produced false lint warnings on the standard
      runtime spelling.
    • A malformed TEXT declaration crashed gasm lint and the language
      server.
      A TEXT line without a symbol left a nil name that lint and
      the LSP dereferenced; both now carry on with a diagnostic. A branch
      to a label at the end of a function body panicked the liveness
      analysis the same way. A real NUL byte truncated the token stream
      (everything after it was dropped); it is an illegal token now, invalid
      UTF-8 no longer inflates byte offsets, and CRLF files format to
      uniform LF.
    • The class-2 stack guard branched four bytes past its target. When
      the underflow branch relaxed to its 32-bit form, its displacement was
      still computed as if the branch were two bytes long, so it landed
      inside the morestack CALL instead of the compare that decides it.
      The long form is reachable once a large frame carries a body of roughly
      a hundred bytes.
    • Immediate operands wrapped silently on amd64. Shift counts,
      immediates beyond the operand's width and displacements beyond int32
      truncated without a diagnostic (SHLQ $300 assembled as $44); they
      are range-checked now, matching go tool asm. EVEX scalar moves
      (VMOVSS Z1, Z2) accepted forms the toolchain rejects and emitted
      invalid encodings; PUSHW/POPW emit the 0x66-prefixed forms the
      toolchain emits; PUSHL is rejected as illegal in 64-bit mode; a bare
      zero-operand JE reports a diagnostic instead of panicking.
    • The arm64 shift and divide instructions encoded entirely different
      operations.
      LSL, LSR, ASR and ROR, immediate and register
      forms, all encoded as ORR; SDIV/UDIV sat in the wrong opcode
      space; MADD/MSUB never encoded the accumulate operand and silently
      read X0 for it. All now match the toolchain byte for byte (new
      differential kernels cover shifts, divides and multiplies), MADD
      takes its four operands in the toolchain's order, and shift amounts at
      or above the operand width are rejected.
    • arm64 multi-chunk immediates corrupted every branch that followed
      them.
      The size pass and the emitter disagreed on the expansion of
      constants with three or more non-zero 16-bit chunks and of MOVW $-1,
      so later label displacements were computed against the wrong offsets.
      The size now comes from the encoder itself. Large-frame stack guards
      (frames from roughly 64 KiB) branched to the wrong morestack entry,
      and the pcsp and DWARF CFA boundaries for materialised large frames
      are computed from the real prologue word counts.
    • arm64 immediates and addressing wrapped instead of erroring.
      Constants beyond the encodable range (ADD $0x100000000) wrapped to
      zero, memory offsets wrapped at 2^31, exclusive and atomic accesses
      silently ignored their offsets (LDXR 8(R1) read [R1]), and large
      register-based offsets were routed through SP instead of the operand's
      base. All four now either encode correctly or produce diagnostics.
    • riscv64 compressed stores with certain offsets wrote to the wrong
      address.
      The C.SD/C.SW/C.FSD immediate pattern dropped one bit, so
      any register-relative store with offset bit 4 or 5 set targeted a
      different address than the same-index load beside it. FENCE
      assembled as fence 0,0 instead of fence iorw, iorw. Branch and
      jump displacements beyond ±4 KiB / ±1 MiB wrapped silently; they are
      diagnostics now. The GOROOT width spellings (MOVW 4(SP), X9)
      compress to their C.LW/C.SW forms exactly as the toolchain lowers
      them, restoring byte parity for those shapes.
    • loong64 MOVW $c, Fd wrote a general register. The immediate was
      routed to the GPR of the F register's number (MOVW $2, F4 clobbered
      argument register R4), and the correct R30 + movgr2fr.w sequence was
      unreachable. Two-operand BLTU R4, label encoded as beqz
      (sometimes-taken where the toolchain's form is never-taken), and
      out-of-range FP constants and BSTRINS/BSTRPICK bit numbers wrapped
      silently; all are corrected or diagnosed.
    • ELF objects carried wrong relocation records. The amd64
      stack-guard TLS load relocated as R_X86_64_PC32 against the null
      symbol (every non-NOSPLIT object mislinked); arm64 SB references
      applied HI21 twice instead of the HI21/LO12 pair; riscv64 PCREL_LO12
      referenced the target instead of its AUIPC site, which the system
      linker rejects; riscv64 and loong64 e_flags declared the soft-float
      ABI, so standard linkers refused the merge.
    • ELF DWARF was unparseable. Eight abbrev-table constants were
      wrong, the version-5 line header carried DWARF2-shaped tables, the
      section count omitted .debug_frame (it sat past the section table,
      invisible to every tool), the CIE hardcoded one architecture's
      CFA and return-address registers for all four, and no DWARF address
      was ever relocated: the .rela.debug_info and .rela.debug_line
      records were computed and then discarded, so every address stayed
      zero after linking. The tables parse in readelf, the registers are
      per-architecture, .rela.debug_info, .rela.debug_line and
      .rela.debug_frame are emitted, and addresses resolve after the
      link; a data-only file emits a valid object instead of panicking, and
      the DWARF records the real source path.
    • GOOBJ cross-package references resolved against the wrong object.
      External package indices were zero-based against a table that
      reserves zero for the dummy invalid package, and symbol indices
      ignored the hashed definition blocks between the sections, so a
      reference into the first external package could bind to whatever
      object the loader saw first. An end-to-end cross-package link pins
      the chain. arm64 ADRP pairs now emit the toolchain's single 8-byte
      relocation (the previous twin 4-byte records were a hard link error),
      and symbols no longer claim the linkname flag the toolchain reserves
      for //go:linkname declarations.
    • The arm64 JIT trampolines saved a scratch register as the stack
      pointer.
      enterJIT and its checked twin stored R3, a plain
      caller-saved register on arm64, and restored RSP from it, so the
      first JIT call would have returned to a garbage stack. The loong64
      trampoline hands its leave address through the raw-symbol pattern the
      arm64 one uses, avoiding the ABI wrapper's prologue.
    • The checked-ABI report flagged legal frames. The red-zone canary
      sat 64 bytes below the entry stack, so any kernel with a larger
      declared frame reported "stack below SP written"; the guard now sizes
      itself from the kernel's frame. The amd64 JIT tests are gated to
      amd64 hosts (the suite previously SIGILL-crashed on the other three
      architectures), fuzz signatures wider than the TEXT frame report
      instead of panicking, --buf specifications are validated strictly
      (a typo no longer verifies against a zeroed buffer), and ABI0
      parameter sizes cover string and complex correctly.
    • amd64 hardware watchpoints never armed. The debug registers were
      poked at offsets inside user_regs_struct, corrupting five general
      registers while the REPL reported success; they now use the real
      u_debugreg window and stop on the watched address. loong64 watch
      goes through the kernel's HW_WATCH regset (riscv64 reports the
      kernel's interface as unsupported instead of failing obscurely).
    • gasm debug hung on the first faulting kernel. Genuine
      SIGSEGV/SIGBUS/SIGFPE/SIGILL stops were discarded as runtime noise
      and the faulting instruction restarted forever; faults now surface as
      reported stops. Conditional breakpoints with a false condition
      resumed mid-instruction, next and finish evaluated traps with
      stale registers and landed off instruction boundaries, and the
      breakpoint restore covered one byte of the four-byte traps (arm64
      silently skipped the instruction under it); the trap PCs follow the
      kernel's reporting on every architecture. regs reports YMM from
      the xstate (a struct-size overrun crashed FP register reads before),
      V register halves print correctly on arm64, unwatch accepts the
      architecture's slot range, x <addr> -8 no longer crashes, break
      conditions accept memory operands, and session scratch directories
      are cleaned up.
    • The language server died on one malformed frame and corrupted
      sources on rename.
      A bad Content-Length or an unparsable JSON
      body terminated the process instead of answering -32700 and
      continuing; rename and references covered the stripped name instead
      of the full ·name token, so renaming produced ·helper minus its
      last letter; documentHighlight never matched middle-dot symbols.
      Positions are UTF-16 code units in both directions (astral characters
      no longer shift columns) and responses always carry an explicit
      result member.
    • The linter now recognises the g spelling of the goroutine
      register.
      MOVD R0, g clobbered R28 on arm64 (and the equivalents
      on the other architectures) unflagged, and the numeric spellings the
      linter did track are rejected by the toolchain there, so g was the
      one spelling that escaped the audit. FUNCDATA and PCDATA literal
      indices are validated against the ranges the runtime defines.
    • Usage errors exit 2 uniformly. audit-instructions and
      scaffold argument errors and an unknown asm --format exited 1 (or,
      for --format without -o, exited 0 silently); the documented
      exit-2 contract now holds, asm -o no longer prints the hex dump it
      claimed to replace, and verify --ground-truth works for amd64
      kernels on non-amd64 hosts instead of refusing with JIT advice.
    • arm64 store-exclusive instructions read their operands in the
      toolchain's order.
      STXR treated the first register as the status
      register where go tool asm reads it as the data register, so the
      same source assembled to different code in the two assemblers; the
      pair forms (STXP, LDXP and their acquire/release variants) are
      accepted now, in the toolchain spelling.
    • Large arm64 frames matched the toolchain's sequences. A frame
      beyond the immediate range that is not a movcon constant (roughly
      64 KiB and up) made gasm verify report a false mismatch: the
      toolchain splits the prologue subtraction into two 12-bit immediates
      and materialises the non-leaf epilogue addition through the temporary
      register; gasm emits the same sequences and the spadj boundaries
      follow the real word counts.
    • The width spellings GOROOT uses assemble. MOVLQZX (four uses in
      runtime/asm_amd64.s), MOVBQSX, MOVWQSX, MOVBLSX, MOVBWSX,
      MOVBWZX and PMOVMSKB (the bytealg kernels) encode byte-identically
      with go tool asm, and the linter reports them encodable; a
      MOVLQZX is the plain 32-bit move, exactly as the toolchain lowers
      it.
    • verify --ground-truth no longer reports a mismatch for functions
      whose size is not a multiple of 16.
      The toolchain pads text symbols
      to 16-byte boundaries; the comparison now checks the padding is zero
      instead of comparing it, the same rule the test suite applies.
    • riscv64 accepts the g spelling of the goroutine register, like
      the other architectures, and the abi kernels use it; every verify
      kernel is now ground-truth checkable (the numeric X27 spelling the
      kernels used is one go tool asm rejects).
    • Two more spellings GOROOT uses now assemble. riscv64 FCLASSD
      (classify a float64 into an integer mask) is encodable, and a
      displacement written as a product (0*8(X5), the toolchain's own
      spelling in several kernels) parses instead of being rejected.
    • loong64 JIT execution enabled. The loong64 trampoline is now
      validated end to end under qemu-user emulation (plain and ABI-checked
      calls, goroutine-clobber detection), so gasm verify runs the JIT
      checks on loong64 hosts instead of forcing every loong64 kernel down
      the ground-truth path. The arm64 and riscv64 trampolines carry the
      same validation; the arm64 ABI test now seeds its kernel arguments
      (a zeroed block made the passthrough check meaningless), and the
      loong64 basic kernel's branch maze terminates on every path so the
      smoke sweep cannot spin on leftover register values.
    Downloads