33 KiB
Architecture
How gasm-devkit is put together and why.
Repository: sourcedock.dev/petrbalvin/gasm-devkit
Overview
Three design goals shape everything below.
- A real AST, not a grammar hack. The linter, analyser, assembler and language server all need to reason about assembly, not just colour it. So the centre of the toolkit is a hand-written lexer and a parser that produce a typed AST with source positions on every node.
- Architecture as data, not code. Per-architecture differences (amd64,
arm64, riscv64, loong64) live in register and instruction tables (
arch) and per-architecture encoders, rather than inif arch == …branches threaded through the analysis; the arch tests that remain are dispatch and policy points, such as which encoder a file name selects and which registers the liveness pass audits. The instruction tables are generated from the Go toolchain's own assembler source (just gen), so refreshing an architecture is a data operation, not a coding one. - Open integration surface. Everything the toolkit can do is reachable through two vendor-neutral interfaces: a CLI and an LSP server. No editor owns the toolkit; the toolkit is offered to editors on standard terms.
The components, and how data moves between them:
flowchart TD
SRC["source .s"] --> LEX["lexer<br/>token stream"]
LEX --> PAR["parser<br/>AST + diagnostics"]
LEX --> FMT["format<br/>re-space tokens"]
PAR --> LINT["lint<br/>static checks"]
PAR --> LSP["lsp server"]
LEX --> LSP
ARCH["arch tables<br/>amd64 / arm64 / riscv64 / loong64"] --> LINT
ARCH --> LSP
LINT --> LSP
PAR --> ASM["asm<br/>encoders, image, object emitters"]
ASM --> VER["verify<br/>JIT mapping, ABI checks, fuzzing"]
ASM --> DBG["debug<br/>ptrace session"]
VER --> DBG
DIS["disasm<br/>golang.org/x/arch"] --> DBG
FMT --> CLI["gasm CLI"]
LINT --> CLI
PAR --> CLI
LEX --> CLI
ASM --> CLI
VER --> CLI
DBG --> CLI
DIS --> CLI
LSP --> EDITOR["any LSP editor"]
The lexer is the shared foundation: the parser builds the AST from it, the formatter re-spaces its tokens directly, and the language server uses it for semantic highlighting. The packages follow a dependency chain: static analysis builds only on the AST, the standalone assembler emits object code, and both the dynamic analysis and the debugger consume the execution substrate the assembler provides.
Packages
| Package | Responsibility |
|---|---|
token |
token kinds and positions |
lexer |
hand-written scanner; permissive, and it never panics |
ast |
the typed syntax tree: declarations, lines, operands |
parser |
line-oriented parser producing the AST and its diagnostics |
arch |
register and instruction tables for the four architectures |
lint |
static checks over the AST |
format |
canonical formatter over the token stream |
lsp |
the language server |
asm |
standalone assembler: encoders, image layout, object emitters |
disasm |
disassembly backend over golang.org/x/arch |
verify |
JIT execution, ABI checks, differential fuzzing |
debug |
interactive ptrace debugger |
cmd/gasm |
the CLI |
_gen |
rebuilds the arch tables from the Go toolchain source |
The boundaries matter as much as the responsibilities: ast records syntax
only, and whether a name is a register or a label is left to arch, so the
parser stays architecture-agnostic. asm produces the machine code, verify
and debug are the two packages that map it executable (read-execute in
verify, read-write-execute in the debuggee), and cmd/gasm is the CLI, with
the verify sweep orchestration and the audit, scaffold and unified-diff
helpers beside its flags and output.
token and lexer
The scanner is hand-written and permissive: it never panics and maps anything
it cannot classify to an Illegal token, so every downstream tool still works
on malformed input. Newlines are significant tokens, because Plan 9 assembly
is line-oriented and the parser relies on line structure.
The middle dot (·, U+00B7) is treated as an identifier character so that
·funcName(SB) lexes as one symbol. Multi-character operators (<<, >>,
->) are recognised so arm64 shift operands scan correctly. A backslash
immediately before a newline is a C-preprocessor line continuation (used by
#define macros in the runtime .s files); the lexer splices the lines
together so a multi-line macro becomes one logical line the parser treats as an
opaque preprocessor directive.
ast and parser
The parser is line-oriented, matching how the Plan 9 assembler reads a file: it groups tokens into lines, classifies each line (directive, label, instruction, comment, preprocessor) and dispatches. A malformed line is reported and skipped; it never aborts the file.
Operands are parsed into a faithful, flat representation. The amd64
addressing modes (reg, $imm, (base), off(base), (base)(index*scale),
name+off(FP), name<>(SB)) are all captured structurally, and the original
token text is retained for fidelity.
A deliberate boundary: the AST records syntax only. Whether a bare
identifier is a register or a label is an architecture question, so it is
left to arch and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
Optional preprocessing
With Options{Expand: true} the parser runs a pre-parse pass
(preproc.go) that splices #include files (the source directory, then the
-I directories), expands object and parameterised #define macros,
applies #undef and the #ifdef/#ifndef/#else/#endif family, and
folds constant expressions left in operands. The go command's platform
macros (GOARCH_<arch>, GOOS_<goos>) arrive through Options.Predefines.
The assembly path (asm, diff, audit) expands; lint, fmt and the
language server read the raw file. The command layer adds the go_asm.h
generator (asmhdr.go): a file that includes go_asm.h gets the package's
defines type-checked out of its Go files for the target architecture and
GOOS, with no compiler in the loop.
arch
Register files are generated programmatically (the regular R8-R15,
X0-X15, Y0-Y15, Z0-Z31, K0-K7 ranges) plus the irregularly
named registers listed explicitly. Instruction names are generated from the
Go toolchain's own assembler source (cmd/internal/obj/<arch>/anames.go,
plus the common opcodes in cmd/internal/obj/util.go) by just gen, so the
tables always match what the real assembler accepts. The spellings the
toolchain's tables do not carry are hand-maintained instead: the front-end
alias lists in arch/arm64.go, arch/amd64.go and arch/loong64.go (the
arm64 B/BL branches among them), and the arm64 .P/.W load-store suffix
stripping in arch/arch.go. Each mnemonic maps to a summary and an optional
operand-count range; counts are recorded only where unambiguous (-1
disables the operand-count lint for that instruction) so the linter stays
silent rather than guess. For architectures with highly variable operand
forms (arm64, riscv64, loong64) relaxCounts clears those counts, leaving
RET and NOP with a range (RET alone on riscv64).
lint
Rules are conservative by design: silence beats a false positive. The rules
are unknown-instruction, operand-count, undefined-label,
duplicate-label, missing-ret, missing-textflag-include, abi-argsize,
unreachable-code, register-clobber, funcdata-pcdata, unused-label,
invalid-textflag, stack-imbalance, register-width-mismatch,
abi0-register-args, nonportable-register-name,
unencodable-instruction and reserved-register-write. Every diagnostic carries a stable
code so callers can disable rules individually, and arch-specific rules switch
off entirely when the target architecture cannot be inferred from the file name.
Two things keep the rules honest on real-world code:
- Pseudo-ops and macros are not instructions.
unknown-instructionknows the assembler pseudo-ops (BYTE,WORD,FUNCDATA,PCDATA, …) and recognises macro invocations: an in-file#definename, or any identifier containing an underscore (no Plan 9 mnemonic ever does). - Macro-heavy files get the label/RET heuristics turned off. Without a
preprocessor, labels a macro defines are invisible, so
undefined-labelandmissing-retare suppressed for files that use macros (an in-file#defineor a#includeof anything other thantextflag.h).missing-retalso treats a trailing unconditional jump andUNDEFas valid terminators.
The result is validated by TestGoRuntimeCorpus, which parses and lints every
src/runtime/*.s file the toolchain ships for all four architectures and
asserts zero parse errors and zero error-severity diagnostics.
Two deeper analyses sit on top of the AST:
abi-argsize. Hand-written kernels document their signature in a// func …comment above theTEXT. The linter parses that signature with the standard library's Go parser, lays out the parameters and results under Go's ABI0 stack rules (results begin on a word boundary after the parameters), and checks the total against the argument size declared in theTEXTdirective. It only runs for stack-argument functions (a non-zero declared arg area that is actually addressed throughFP), and aborts silently on a type whose size it cannot determine, so it never guesses.unreachable-code. Code after aRETand before the next label is dead. The check is suppressed for any function whose reachability cannot be decided statically: those using PC-relative jumps (JMP 2(PC)), register-indirect branches (JALR/JR/JIRL/BR/BLR, or aJMP/CALLthrough a register or memory operand), or living in a file with#ifdefconditionals.UNDEFis deliberately not a terminator: code after it is occasionally intentional metadata.register-clobber(register liveness). The linter builds the function's control-flow graph (basic blocks split at labels and after branches, with fall-through and jump-target edges), computes a conservative per-instruction register def/use, and runs the standard backward liveness iteration to a fixed point. On top of that it flags writes to the registers the Go ABI fixes across calls that are never saved and restored, calibrated fromcmd/compile/abi-internal.md, not the platform ABI: Go's stack-based ABI0 has no System V style callee-saved registers (amd64BX,R12-R15and the like are caller-saved or permanent scratch, and hand-written kernels may clobber them freely). The audited set is the frame pointer and the goroutine pointer per architecture (amd64BP/R14, arm64R18/R28/R29, riscv64X27, loong64R22); the goroutine pointer is reported only when the function can reach the runtime (it is notNOSPLITor makes a call), since the ABI0 transition machinery restores it on those paths, and NOSPLIT call-free leaves may use it (the runtime's own assembly does). It runs only on macro-free files, where no opaque macro can perform the save/restore.funcdata-pcdata.FUNCDATA $idx, sym(SB)andPCDATA $idx, $valare checked for well-formed operands (arity, immediate index and value, symbol reference) and a literal index is range-checked; a named index constant such as$PCDATA_StackMapIndexis accepted without a range check.
format
The formatter works on the token stream, not the AST, so it preserves
every line, comments and blanks included. It normalises indentation, operand
spacing, per-function mnemonic alignment and blank-line layout: a new block
(a label, TEXT or GLOBL) is preceded by exactly one blank line (comments
leading a block stay with it), runs of blanks collapse to one, and a RET
terminates the body so the next function's doc comment stays at column 0. It
is idempotent and its output always round-trips through the parser. With a
directory argument, or none, it reformats every .s file below it in
place and lists the files changed, the way go fmt does (. and _
directories are skipped).
lsp
The server speaks JSON-RPC 2.0 with Content-Length framing over any
io.Reader/io.Writer (normally stdin/stdout). It maintains an in-memory
document store, republishes diagnostics on every change, and provides:
- completion: instructions, registers, pseudo-registers, textflag macros and local labels;
- hover: instruction summaries and register descriptions from
arch; - document symbols:
TEXTfunctions with their labels, plusGLOBL/DATA; - semantic tokens: syntax highlighting delivered as LSP semantic tokens,
classified with the lexer plus
arch(instructions, registers by class, pseudo-registers, labels, immediates, comments, directives, textflag macros); - navigation: go-to-definition from a label reference to its definition, find references, document highlights of every use of the symbol under the cursor, rename, and workspace symbol search over the open documents;
- assists: document formatting through the
formatpackage, inlay hints (the frame size after the TEXT argument area), signature help (the callee's// funcsignature while the cursor is on aCALL), and code actions offering quick fixes for themissing-retandunused-labeldiagnostics. - document information: pull diagnostics (
textDocument/diagnostic), #include document links (resolved against the document directory, then$GOROOT/pkg/include) and folding ranges (one collapsible region per TEXT function body).
Semantic tokens are the key to editor-agnostic highlighting: the editor renders them from the standard LSP legend, so no editor-specific grammar is needed.
asm
The standalone assembler. Its core is an amd64 instruction encoder:
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
with the Plan 9 operand order (source first) mapped onto the x86 encoding.
Every encoding is validated by decoding it again with golang.org/x/arch, the
one module dependency, which also backs the gasm dis listings.
A RISC-V encoder (RV64IMAFDC + RVC compression) encodes the full
integer, atomic, float/double, FMA and CSR instruction sets with the MOV
pseudo-instruction and SB/global symbol references (AUIPC pairs with
R_RISCV_PCREL_HI20/LO12 relocations). The encoder compresses eligible
instructions to 16-bit RVC forms and is validated byte-for-byte against
GOARCH=riscv64 go tool asm.
A LoongArch encoder (LoongArch64) encodes the integer and
floating-point instruction sets with the dual-form arithmetic mnemonics (3R
vs 2RI12), the 16/21-bit branch families, the MOV pseudo-instruction and its
constant materialisation (the dcon classification driving lu12i.w/ori/lu32i.d/
lu52i.d expansions), the FP/SP frame mapping (autosize = align8(frame+8),
prologue storing the link register before and after the SP decrement) and
SB/global symbol references (pcalau12i pairs with R_LOONG64_ADDR_HI/LO
relocations). Like the RISC-V encoder it is validated byte-for-byte against
GOARCH=loong64 go tool asm, and its GOOBJ output is proven end-to-end by
substituting it into a cross-compiled go build and linking with cmd/link.
An AArch64 encoder (arm64) encodes the integer instruction set
with the data-processing (shifted register and immediate forms), load/store
(scaled unsigned immediate and unscaled9-bit immediate), conditional and
unconditional branches, the MOV pseudo-instruction and its constant
materialisation (MOVZ/MOVN/MOVK for wide immediates, ORR with logical bitmask
encoding for values like $1), the FP/SP frame mapping (autosize =
align16(frame+8), prologue using pre-index store for small frames and
STP+SUB for large frames) and SB/global symbol references (ADRP+ADD pairs with
R_ADDRARM64 relocations). Like the other encoders it is validated
byte-for-byte against GOARCH=arm64 go tool asm.
On top of the per-architecture encoders, every framed function carries the
stack-split guard: the prologue check against g.stackguard0 (small,
medium and large frame classes, the medium and large classes materialising
their offset through the architecture's temporary register and the large
class adding the SP-underflow branch) and the trailing morestack block
(save the link register, CALL runtime.morestack_noctxt, jump back to the
function entry). The auto-NOSPLIT rule, the frame classes, the large-frame
prologue and epilogue forms and the tail calls match the toolchain's
stacksplit and preprocess output byte for byte; a parity suite
assembles kernel files with gasm and the installed go tool asm and diffs
the bytes on all four architectures.
On top of the encoder, Assemble walks a parsed TEXT body, converts each
operand to an encoder operand, and lays the instructions out so local labels
resolve to relative jump offsets: jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating to a
fixed point, and jump-to-jump chains are folded (a conditional jump to a label
whose only instruction is an unconditional jump is redirected to the ultimate
target) exactly as the Go toolchain's linker does before it encodes branches.
The FP/SP pseudo-
registers are translated onto the hardware stack pointer: x+N(FP) becomes
(N+8)(SP) for a zero-frame function and (N+frame+16)(SP) once a frame
pointer is set up, with the matching Go prologue/epilogue generated, so the
output is byte-identical to the Go assembler for these cases. SIMD is handled
by a VEX (AVX/AVX2) encoder (the two- and three-byte VEX prefixes with XMM/YMM
registers) over nine operand forms plus a dedicated move encoder: the
three-operand NDS form, the two-operand reg/rm form, the immediate-shift form
(plus the variable-count shifts, which share the NDS shape with the count in
an XMM register or memory), the immediate shuffle form (VPSHUFD, VPERMQ),
the three-operand-plus-immediate form (VSHUFPD,
VPERM2I128, VINSERTI128), the lane-extract form (VEXTRACTI128,
VEXTRACTF128, where the YMM source occupies the reg field and the XMM or
memory destination r/m), the direction-sensitive moves (VMOVDQU, VMOVUPD,
VMOVD, VMOVQ, VMOVSD), the floating-point and FMA arithmetic: the
packed double operations (VADDPD/VSUBPD/VMULPD/VDIVPD/VMINPD/
VMAXPD), the unpacks (VUNPCKHPD/VUNPCKLPD), the scalar SD and SS
operations, VMOVDDUP, VXORPD, the width-changing conversions
(VCVTDQ2PS, VCVTPS2PD, VCVTDQ2PD, and the VCVTPD2DQX/Y and
VCVTTPD2DQX/Y spellings, whose length follows the wider source) and
VFMADD231PD, and the no-operand VZEROUPPER, together with VPERMD and
the scalar families (CMOVcc, SETcc, LZCNT/TZCNT, the extending moves,
CVTSx2SD, IMUL3) and the EVEX (AVX-512) prefix, the four-byte prefix with
5-bit register fields (Z0-Z31, X/Y 16-31, with the mod=11 quirk that carries
rm[4] in X̄), opmask registers (K0-K7 as operands, mask destinations and
explicit merging/zeroing masks, written the way Go writes them, as a K
operand among the operands plus a .Z mnemonic suffix), and the compressed
disp8×N displacement, whose multiplier follows the memory operand's size,
covering every instruction the go-flac and go-lz4 AVX2/AVX-512 kernels use,
plus the common AVX-512 F/BW integer set, the floating-point and conversion
set (the packed double and single arithmetic, the scalar SD/SS forms
(whose EVEX encodings serve masked and zeroing use), VMOVDDUP, the
replicating moves, and the width-changing conversions, including the
VCVTPD2DQ/VCVTTPD2DQ family whose length follows the wider source
operand), and the wider AVX-512 set: ternary logic, lane shuffles, inserts
and extracts, compares with an opmask destination, the permutes, the
expand/compress family, the broadcasts, the opmask-register instructions
(KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/KSHIFTL/KORTEST and KMOVQ), the aligned
moves and the remaining extending/narrowing moves, the floating-point
helper and conversion tail (VRCP14*, VRSQRT14*, VGETEXP*, VGETMANT*,
VSCALEF*, VRNDSCALE*, VREDUCE*, VFIXUPIMM*, VRANGE*, VFPCLASS* with an
opmask destination, and the VCVT* conversions, signed, unsigned and
truncating, including the length-suffixed X/Y spellings and the
mask/vector conversions VPMOVM2*/VPMOV2M, and the scalar conversions
between vector and general-purpose registers (VCVT{,T}S{D,S}2SI{,Q} and
the unsigned forms, VCVTSI2/VCVTUSI2*), and gather/scatter with VSIB addressing, both the
VEX spelling with a vector mask register and the EVEX spelling with an
explicit K mask, where the EVEX length follows the VSIB index register,
not the data register. The EVEX mnemonic
suffixes (rounding modes (.RN_SAE/.RD_SAE/.RU_SAE/.RZ_SAE),
suppress-all-exceptions (.SAE) and memory broadcast (.BCST)) set the EVEX
b bit and the L'L rounding-control field (broadcast keeps the vector length
and scales disp8 by the element size), and combine with the .Z zeroing
suffix. Every encoding is validated two ways: by
round-trip decoding through golang.org/x/arch, and byte-for-byte against
the machine code the real Go assembler emits; the parity suites carry that
comparison over whole kernel files on all four architectures, with the
relocation fields masked because the Go linker fills those displacements at
link time.
File-level assembly (AssembleFile) goes beyond single functions: it
materialises the file's static symbols (GLOBL/DATA) in a data section
behind the code and resolves references to them (mask<>(SB)) to
RIP-relative loads whose displacements point inside the resulting image, so
the bytes are self-consistent at any base address. References to symbols no
GLOBL defines are kept as relocations on the function layout, and the
object-file emitters turn the whole image into a linkable object: the ELF
writer (gasm asm --format elf) lays the code and data out as .text/.data
sections, exports a symbol per
TEXT and GLOBL (the <> ones local, the rest global), emits one
PC-relative relocation per static-symbol reference, undefined external
symbols included, and appends the DWARF5 debug sections
(.debug_abbrev, .debug_info, .debug_line, .debug_line_str, and a
.debug_frame CFI section on amd64) so addr2line and GDB/LLDB can debug
the output. The GOOBJ
emitter (gasm asm --format goobj) writes the format the Go linker consumes
directly: the functions as non-package symbols (the way cmd/asm records
assembly symbols), the GLOBL data, one FuncInfo per function and the
pc-value tables (pcsp built from the prologue and epilogue stack
boundaries, plus flat pcfile, pcline and pcinline tables), so a
gasm-assembled object drops into a go build in place of the toolchain's.
The object preamble (the version-and-experiment header the linker compares
verbatim) is captured from the installed go tool asm, so the output is
always consistent with the toolchain that links it. RISC-V and LoongArch
GOOBJ emission share this emitter: the loong64 marker with
R_LOONG64_ADDR_HI/LO relocation types, and the riscv64 marker with a single
R_RISCV_PCREL_ITYPE/STYPE relocation per AUIPC pair (plus R_RISCV_JAL for
CALL sym(SB)), the model cmd/asm
writes, not the ELF HI20/LO12 pair, and both link into a real go build for
their GOARCH. Per function, the emitter also writes the two DWARF
symbols the linker's DWARF pass reads verbatim: the subprogram DIE
(SDWARFFCN) and the .debug_line state-machine program (SDWARFLINES),
both built the way cmd/asm builds them (the DIE carries the
R_DWTXTADDR_U4 address reference; the line program one row per source-line
change, in the same special-opcode encoding), and the pc-value deltas are
in the architecture's MinLC units, as the runtime's pcvalue expects.
Cross-package external references resolve on all four architectures: when
img.Externals is non-empty, the GOOBJ emitter locates the referenced
package's .a archive via go list -json -export, reads the GOOBJ symbol
definitions the linker reads, and wires the resolved package and symbol
indices into the emission, so a gasm-assembled object links against the
compiled packages it references.
verify
The dynamic-analysis substrate. It JIT-loads assembled images into executable memory and invokes them directly, enabling differential testing, runtime ABI checks and coverage profiling.
The execution model is pure Go (stdlib only). Map copies machine code into
an anonymous syscall.Mmap mapping and enforces W^X (write the bytes, then
mprotect to read-execute). Call prepares a stack whose first word is the
address of an assembly trampoline (leaveJIT), lays the ABI0 argument
block after it, switches to that stack via enterJIT (which saves the Go
stack pointer in a package global and jumps to the target), and recovers
control when the function RETs into leaveJIT (which restores the Go stack
and returns). A 64-byte pad below the return address accommodates the
ABIInternal wrapper that the Go runtime interposes on assembly functions.
Every supported architecture carries its own hand-written trampoline pair
(trampoline_amd64.s, trampoline_arm64.s, trampoline_riscv64.s,
trampoline_loong64.s). The ABI-checked variant CallChecked exists for
every architecture too: enterJITChecked plants sentinels in the registers
the Go ABI fixes across calls (amd64 BP/R14, arm64 R29/R28, riscv64
X27, loong64 R22; the latter two keep no hardware frame pointer) and the
raw return trampoline leaveJITCheckedRaw verifies them, restoring the
saved registers before Go code resumes. All three non-amd64 trampolines
are validated end to end under qemu-user emulation, the loong64 one
through its raw-address leave handoff.
gasm verify runs the JIT checks when the host
matches the kernel's architecture and the toolchain comparisons
elsewhere.
Load / LoadSource / LoadAST parse, assemble and map a .s file in one
step, returning a Kernel whose CallFunc method marshals the argument block
by name. The image must be self-contained (no external relocations); the
assembler’s Image.Bytes() provides the code-and-data concatenation.
The gasm verify CLI subcommand exposes this: it loads a file, reports the
available functions and (with -smoke) calls each NOSPLIT function with zeroed
arguments to confirm the trampoline round-trips. The -smoke and -abi
sweeps run in parallel and each inside a child process, so a function that
faults is reported without ending the sweep; -abi is where the ABI check
lives, fuzzing each function with sentinel values in the registers the Go ABI
fixes across calls and a canary below SP, and reporting a violation on any
iteration. gasm verify --fuzz is the differential campaign instead: it
JIT-loads the kernel and the go tool asm build of the same kernel and
compares the output argument areas bit-for-bit, one child process per function
so a crash on a partial function is reported rather than fatal. When a fuzz
iteration crashes or mismatches, FuzzResult.CrashInput stores the exact input
for reproducibility. gasm verify --call <func> --buf name:size:pattern
invokes a single function with user-supplied buffers (patterns: zero, ones,
seq, or hex), printing the ABI0 argument block before and after the call,
useful for partial functions (e.g. decoders) that crash on random input but
should succeed on valid data; --args name=value,... supplies scalar
arguments (decimal or 0x hex) alongside the buffers. --save-corpus
records every input that crashes or mismatches as replayable JSON, and
--replay re-runs saved entries in isolated child processes, reporting
whether each reproduces. gasm verify --ground-truth compares the
assembled machine code byte-for-byte against go tool asm (relocation sites
masked), reporting any encoding drift.
debug
The interactive debugger (all four architectures). It launches the target
function in a child process that maps the JIT code, calls
PTRACE_TRACEME, and stops; the parent attaches via ptrace and controls
execution. Breakpoints are patched through /proc/pid/mem: the one-byte
INT3 on amd64, the four-byte break instruction on the other three (arm64
BRK #0, riscv64 ebreak, loong64 break 0).
The child pins its goroutine to the OS thread with runtime.LockOSThread
so the traced thread is the one executing JIT code. The REPL provides
single-step, register inspection (the GPRs on every architecture; on amd64 the
XMM set through PTRACE_GETFPREGS and the YMM set through PTRACE_GETREGSET
on NT_X86_XSTATE; on the other three the FP/SIMD regset through
PTRACE_GETREGSET on NT_PRFPREG),
label resolution, named buffer allocation with pattern filling
(--buf name:size:pattern: zero, ones, seq, or hex), and breakpoint
management. Breakpoints accept conditions
(break <label> if <reg> <op> <val>, including register-against-register
comparisons), and hardware watchpoints work on amd64 (the DR0-DR3 debug
registers), arm64 (NT_ARM_HW_WATCH) and loong64 (NT_LOONGARCH_HW_WATCH);
riscv64 reports that its kernel ptrace interface exposes no trigger regset.
The ptrace path is validated at run time on amd64, where the session tests are
built; arm64, riscv64 and loong64 compile and are covered by the
architecture-neutral units (label and line tables, the breakpoint manager).
For non-interactive use, --script runs REPL commands from a file (or
stdin) and exits, --timeout kills the debuggee when a run hangs (the
watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot
block it), and --cover runs to completion with a breakpoint on every
instruction and reports which instructions executed and how often.
Extending the toolkit
- New architecture: add an entry to the generator in
_gen, runjust gen, and add abuildXXX()register file plus a case inForArch. - New lint rule: add a function in
lintand a rule-code constant. - New LSP feature: add a method case in
dispatchand a handler.
Data flow
The main operation, assembling one file:
sequenceDiagram
participant User
participant CLI as gasm CLI
participant Parser as parser
participant Asm as asm
participant Go as go toolchain
User->>CLI: gasm asm --format goobj -p pkg -o k.o k_amd64.s
CLI->>Parser: Parse(path, src)
Parser-->>CLI: AST, diagnostics
CLI->>Asm: AssembleFile(AST)
Asm->>Asm: encode operands, settle label offsets, lay out data
Asm-->>CLI: Image, code and data and relocations
CLI->>Asm: GOObject(pkg, path)
Asm->>Go: go list -json -export, externals only
Go-->>Asm: package and symbol indices
Asm-->>CLI: Go object bytes
CLI-->>User: wrote N bytes to k.o
Errors are produced where the parse or the encoding fails and become values at
the CLI boundary: the parser returns a diagnostic list and never aborts a file,
AssembleFile returns an error, and cmd/gasm prints what it has to stderr
and returns a non-zero exit code. The formatter and the linter take different
inputs from the assembler: gasm fmt re-spaces the token stream
(format.Source lexes the source text itself) and gasm lint walks the parsed
AST, so neither depends on an encoding.
State and lifetime
- The analysis packages (
lexer,parser,format,lint,arch) hold only read-only lookup tables and no mutable state: every call allocates its own tokens and AST, and any number of goroutines may read thearchtables. - A
verify.Kernelowns one executable mapping, whichClosereleases. The JIT trampolines keep the Go stack pointer and the checked-call sentinels in package globals, so a call is a process-wide, one-at-a-time operation. Thegasm verifysweeps therefore run each function in a child process, which contains a crash and keeps the globals unshared. lsp.Serveris long-lived: it runs a single read and dispatch loop over the stream and touches its document store only from that loop, so one server serves one connection.- A
debug.Sessionowns a traced child process and pins its goroutine to the forking OS thread, because ptrace requests must stay on that thread.
Dependencies
golang.org/x/arch(v0.30.0) is the one module dependency: it is the disassembler backend (gasm disand the debugger's listings). The tests additionally decode through it to validate the encodings.- The Go toolchain, as an oracle and never as a library:
go tool asmsupplies the object preamble and the ground truth forgasm verify --ground-truth,go list -json -exportlocates the archives of the packages a GOOBJ object references, and_genparses$GOROOT/src/cmd/internal/obj/<arch>/anames.goto rebuild the tables. - Linux process interfaces for the dynamic work:
mmapandmprotectfor the JIT mapping, ptrace with/proc/pid/memfor the debugger. That is whyverifyruns a JIT check only when the host architecture matches the kernel's, and whydebugis Linux-only.