feat(asm): byte-identical go-flac AVX2 assembly with scalar families and jump relaxation
Assisted-by: Qwen 3.8 Max Preview
This commit is contained in:
+23
-17
@@ -182,33 +182,39 @@ Every encoding is validated by decoding it again with `golang.org/x/arch` — th
|
||||
one module dependency, used in tests only and never linked into the binary.
|
||||
|
||||
On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each
|
||||
operand to an encoder operand, and lays the instructions out in two passes so
|
||||
local labels resolve to fixed rel32 jump offsets. The `FP`/`SP` pseudo-
|
||||
operand to an encoder operand, and lays the instructions out so local labels
|
||||
resolve to relative jump offsets: jumps start in the short (rel8) form and
|
||||
expand to rel32 when the settled displacement does not fit, iterating to a
|
||||
fixed point, and jump-to-jump chains are folded (a conditional jump to a label
|
||||
whose only instruction is an unconditional jump is redirected to the ultimate
|
||||
target) exactly as the Go toolchain's linker does before it encodes branches.
|
||||
The `FP`/`SP` pseudo-
|
||||
registers are translated onto the hardware stack pointer — `x+N(FP)` becomes
|
||||
`(N+8)(SP)` for a zero-frame function and `(N+frame+16)(SP)` once a frame
|
||||
pointer is set up, with the matching Go prologue/epilogue generated — so the
|
||||
output is byte-identical to the Go assembler for these cases. SIMD is handled
|
||||
by a VEX (AVX/AVX2) encoder — the two- and three-byte VEX prefixes with XMM/YMM
|
||||
registers — across seven operand forms: the three-operand NDS form, the
|
||||
two-operand reg/rm form, the immediate-shift form, the immediate shuffle form
|
||||
(`VPSHUFD`, `VPERMQ`), the three-operand-plus-immediate form (`VSHUFPD`,
|
||||
registers — across eight operand forms: the three-operand NDS form, the
|
||||
two-operand reg/rm form, the immediate-shift form (plus the variable-count
|
||||
shifts, which share the NDS shape with the count in an XMM register or
|
||||
memory), the immediate shuffle form (`VPSHUFD`, `VPERMQ`), the
|
||||
three-operand-plus-immediate form (`VSHUFPD`,
|
||||
`VPERM2I128`, `VINSERTI128`), the lane-extract form (`VEXTRACTI128`,
|
||||
lane-extract form (`VEXTRACTI128`,
|
||||
`VEXTRACTF128`, where the YMM source occupies the reg field and the XMM or
|
||||
memory destination r/m), the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`,
|
||||
`VMOVD`, `VMOVQ`, `VMOVSD`), the floating-point and FMA arithmetic (`VADDPD`,
|
||||
`VMULPD`, `VXORPD`, `VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`,
|
||||
`VFMADD231PD`) and the no-operand `VZEROUPPER` — together with `VPERMD`,
|
||||
covering every integer, shuffle and FP instruction the go-flac AVX2 kernels
|
||||
use. Every encoding is validated two ways: by round-trip decoding
|
||||
through `golang.org/x/arch`, and byte-for-byte against the machine code the
|
||||
real Go assembler emits (which also locks the v̄vvv = 1111 rule for unused
|
||||
vvvv fields — a value the hardware rejects with #UD and the decoder silently
|
||||
ignores). This increment covers register / memory / immediate / FP-frame
|
||||
operands, local-label jumps and these VEX SIMD forms; EVEX / AVX-512, `SB`
|
||||
(global symbol) operands (relocations), a handful of scalar gaps the kernels
|
||||
hit (`CMOVcc`, `SETcc`, `LZCNT`, `MOVSX`/`MOVZX`) and object-file emission
|
||||
are the rest of Phase 2.
|
||||
`VFMADD231PD`) and the no-operand `VZEROUPPER` — together with `VPERMD` and
|
||||
the scalar families (`CMOVcc`, `SETcc`, `LZCNT`/`TZCNT`, the extending moves,
|
||||
`CVTSx2SD`, `IMUL3`), covering every instruction the go-flac AVX2 kernels use
|
||||
apart from global-symbol loads. Every encoding is validated two ways: by
|
||||
round-trip decoding through `golang.org/x/arch`, and byte-for-byte against the
|
||||
machine code the real Go assembler emits — a comparison that now holds for
|
||||
whole functions: every kernel function that avoids `SB` operands assembles to
|
||||
exactly the Go toolchain's bytes. This increment covers register / memory /
|
||||
immediate / FP-frame operands, local-label jumps and these VEX SIMD forms;
|
||||
EVEX / AVX-512, `SB` (global symbol) operands (relocations) and object-file
|
||||
emission are the rest of Phase 2.
|
||||
|
||||
## Extension points
|
||||
|
||||
|
||||
Reference in New Issue
Block a user