• v0.35.0
    Test / test (push) Successful in 2m26s
    Release / gates (push) Successful in 2m25s
    Release / build (amd64, linux) (push) Successful in 1m15s
    Release / build (arm64, linux) (push) Successful in 1m16s
    Release / build (loong64, linux) (push) Successful in 1m16s
    Release / build (riscv64, linux) (push) Successful in 1m16s
    Release / release (push) Successful in 34s
    Stable

    petrbalvin released this 2026-09-21 23:33:11 +00:00 | 0 commits to main since this release

    Added

    • The go_asm.h generator. gasm asm generates the package's go_asm.h
      itself when an assembly file includes it: the Go files beside the source
      are type-checked for the target architecture and the constants and field
      offsets become assembler defines, so package-context files assemble with
      no compiler and no go build in the loop. -GOOS selects the
      type-checking GOOS for GOOS-specific files, and the corpus audit derives
      the GOOS from the file name.
    • ELF data relocations on arm64, riscv64 and loong64. gasm asm --format elf emits .rela.data for symbol-valued DATA initialisers on
      every architecture (amd64 carried them already), so standalone ELF
      objects link on all four targets.
    • Corpus failure listing. gasm audit-instructions --corpus --list
      prints every failing file with its failure reason, per architecture,
      instead of one representative file per reason.
    • DATA with symbol values and relaxed symbol spellings. DATA
      initialisers accept $symbol(SB) values, laid down as an absolute
      relocation at the data field (GOOBJ on all four architectures and ELF
      on all four as of this release), and U+2215 is accepted inside symbol
      package paths.
    • Macro expansion and include splicing. gasm asm, gasm diff and
      gasm audit-instructions now preprocess assembly the way the
      toolchain does: object and parameterised #define macros expand at
      the point of use, #undef and the #ifdef/#ifndef/#else/
      #endif family select branches, #include splices headers resolved
      through the source directory and the new repeatable -I flag, ;
      separates statements, and constant expressions left in operands
      ($(32-7), $~63, (index*4)(base)) fold at parse. Expansion
      happens only on the assembly path: gasm lint, gasm fmt and the
      language server keep reading the raw file.
    • Encoder coverage: the instruction families GOROOT's real code
      uses.
      The encoder now covers the
      instruction families GOROOT's real code uses that gasm lacked,
      byte-verified against go tool asm: on amd64 the carry ALU, the
      atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
      GFNI, ADX, BMI, the string primitives, the system set (CPUID, RDTSC,
      SYSCALL, fences, MXCSR) and the SSE/AVX/EVEX gaps; on arm64 the pair
      loads and stores (LDP/STP), acquire/release and LSE atomics, AES and
      SHA, the system operations, the bit ops and the NEON slice including
      structure loads and the literal-pool moves; on riscv64 the RV64A AMO
      family with aq/rl ordering, the Zbb pseudos with their RVC
      compressions, the FMA forms and the RVV slice with vsetvli/
      vsetivli; on loong64 the AM atomics with acquire/release forms, the
      LSX/LASX slice, the VMOVQ/XVMOVQ transfer family and FSEL.
      Also fixed on the way: arm64 CASD/CASW lacked an opcode bit, and
      riscv64 VSETVLI with an immediate length now canonicalises to
      vsetivli as the toolchain does.
    • Encoder coverage: quad-register AVX-512 and floating-point
      immediates.
      The encoder gains the
      quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
      VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
      floating-point immediates on the SSE scalar moves and arithmetic
      (the constant lands in a synthesised read-only pool, a positive zero
      collapses to XORPS exactly as the toolchain does), accept-and-ignore
      FUNCDATA and PCDATA, three-operand double shifts, static-symbol
      operands for the legacy SSE moves, and the pooled 64-bit immediate
      materialisation on riscv64. The parser carries bracketed register
      ranges, index-only VSIB memory operands and bare trailing immediates;
      macro substitution reaches parameters used with element suffixes
      (A.S4), and ; separates statements in plain files.
    • Per-architecture reference pages. docs/asm/
      gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
      roles the ABI fixes, addressing, operand order with every special form,
      constants and materialisation, alignment, fences and the relocations
      each target emits. An instruction inventory appendix per architecture
      is generated from the toolchain's own tables by just gen, and the
      regenerated tables recognise 147 more mnemonics than the previous
      release carried (arm64 107, riscv64 31, loong64 9).
    • The Plan 9 assembly language reference. docs/asm/
      opens the complete language reference with its common core: the lexicon,
      statement structure and constant expressions, the operand grammar with
      the pseudo-registers and symbol naming, the directives and the function
      flag vocabulary, preprocessing with #define and #include, and the
      Go-embedded layer (ABI0, prototypes, go_asm.h, funcdata.h and the
      runtime contract). Every claim is verified against go tool asm of
      Go 1.27.1 and gasm's differential tests; the per-architecture pages and
      generated instruction appendices follow.
    • GOOBJ format specification. docs/GOOBJ.md
      documents the Go object file format in full: both containers, the 96
      byte header and all 19 blocks, every structure with its byte
      offsets, symbol kinds and flag bits, all 106 relocation types with
      the weak variants, aux symbols, the FuncInfo payload, the pc-value
      table encoding, the content hashes and the builtin table, all
      verified byte for byte against objects produced by Go 1.27.1's own
      tools.

    Changed

    • The corpus audit measures like a build. Files named for a Go port
      gasm does not target (arm, 386, s390x, ...) are never attempted, because
      no supported build compiles them; the GOOS comes from the file name; and
      each target's go_asm.h is generated on the fly. The headline is reported
      over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
      for every target architecture and 295 of 303 on real code (97.4 %),
      against 108 of 627 over all files (17.2 %) that the previous release
      measured.

    Fixed

    • The operand forms GOROOT writes. Numeric PC-relative jumps
      (JEQ 2(PC), the park loop JMP 0(PC)) resolve with the toolchain's
      own instruction counting and fold jump-to-jump chains exactly as its
      branch optimiser does; symbol immediates (MOVQ $sym(SB), AX)
      assemble to the toolchain's RIP-relative LEA with an R_PCREL
      relocation; negated constant expressions in operands (ADJSP $-(REGS - 8), the shape the cgo ABI macros write) fold; the immediate
      multiply (IMULQ $1000000000, AX) encodes with the toolchain's
      0x69/0x6B selection; the TLS access pair assembles as the toolchain's
      one-instruction form (the bare MOVQ TLS, r load nops out and
      off(r)(TLS*1) folds to the segment-prefixed absolute whose disp32
      carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
      bare-register indirect branch (BL R9 beside BL (R9), both BLR) and
      the zero-immediate store (MOVD $0, mem through the zero register,
      rejecting non-zero immediates as the toolchain does); PCALIGN now
      aligns on amd64, padding with the toolchain's greedy
      single-instruction NOPs; the segment-absolute forms (MOVQ 0x30(GS), AX and the store direction) and the absolute crash-store
      (MOVL $0xf1, 0xf1) encode; and gasm asm predefines the
      GOARCH_<arch> and GOOS_<goos> macros the go command passes to
      go tool asm, so GOROOT headers' #ifdef GOARCH_amd64 platform
      blocks (go_tls.h's get_tls and friends) select as intended. The
      GOROOT corpus measure moves to 291 of 353 files assembling for every
      target architecture (82.4 %), 97.4 % of the real-code corpus, from
      70.8 % and 82.2 %.
    • Tool corrections across the pipeline. The formatter keeps square
      brackets in SIMD operands, statement separators and canonical macro
      bodies; the linter drops false positives on shift counts, SETcc
      spellings and ABIInternal references; the lexer treats a trailing
      carriage return as a line end so comment text stays idempotent; and
      arm64 rejects bare BTI with a diagnostic while accepting the full
      family.
    Downloads