Assisted-by: GLM 5.3 Flash
1.8 KiB
Measurement: the request path without waste
- Date: 2026-09-22
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), otherwise idle
- Toolchain: go1.27.1, no build flags
- Command:
go test ./internal/nfs4server/ -run '^$' -bench=Wire -benchmem -count=5
Baseline
The commit c50cd0b, measured in the same session as the new code, both sides five runs back to back. The baseline is the wire baseline of 2026-09-22-wire-baseline.md plus the client file commands.
Result
Median of five runs per side, one session.
| Benchmark | Latency old | Latency new | Change | Allocs old | Allocs new | Change |
|---|---|---|---|---|---|---|
BenchmarkWireRead64K |
92.0 µs | 81.5 µs | -11.4 % | 78 | 63 | -19 % |
BenchmarkWireWrite64K |
61.4 µs | 72.2 µs | see note | 76 | 62 | -18 % |
BenchmarkWireGetattr |
14.2 µs | 13.5 µs | -4.5 % | 85 | 70 | -18 % |
BenchmarkWireLookup |
15.2 µs | 15.3 µs | 0 % | 81 | 66 | -19 % |
Bytes per operation fell 27 percent on READ and 13 percent on WRITE. The
changes behind the numbers: the COMPOUND answer accumulates in one buffer
with the header patched in place instead of copying every operation result
twice, READ fills the reply buffer through the backend's ReadInto instead
of an intermediate allocation and drops the second attribute call for the
end of file flag, WRITE hands the request record's own bytes to the backend
instead of copying them out, and the outgoing record marking buffers recycle
through a bounded pool.
The WRITE latency column carries an honest warning: the machine's session to session variance on this benchmark exceeds the effect being measured. Within a single session the order of the two sides flipped twice; the deterministic counters, allocations and bytes, are the trustworthy part of the WRITE row, and the READ row's improvement is well outside the noise.