feat: full NFSv4.2 server and client in pure Go
Test / test (push) Successful in 2m4s
Release / gates (push) Successful in 2m5s
Release / build (amd64, freebsd) (push) Successful in 1m27s
Release / build (amd64, linux) (push) Successful in 1m22s
Release / build (amd64, netbsd) (push) Successful in 1m19s
Release / build (amd64, openbsd) (push) Successful in 1m20s
Release / build (arm64, darwin) (push) Successful in 1m21s
Release / build (arm64, freebsd) (push) Successful in 1m26s
Release / build (arm64, linux) (push) Successful in 1m25s
Release / build (arm64, netbsd) (push) Successful in 1m31s
Release / build (arm64, openbsd) (push) Successful in 1m27s
Release / build (loong64, linux) (push) Successful in 1m37s
Release / build (riscv64, linux) (push) Successful in 1m21s
Release / release (push) Successful in 40s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-21 18:51:17 +02:00
commit a9b8039ef7
153 changed files with 34403 additions and 0 deletions
@@ -0,0 +1,31 @@
# Measurement: the descriptor cache in internal/nfsfs
- Date: 2026-09-22
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), idle
- Toolchain: go1.27.1, no build flags
- Command: `go test ./internal/nfsfs/ -run '^$' -bench=. -benchmem -count=5`
## Baseline
The commit b8758b7, the head of development before the descriptor cache: every
READ, WRITE and SYNC opened the registered path, verified it with a stat and
closed the descriptor again, per operation. The cache keeps idle descriptors of
regular files in a bounded LRU and revalidates the identity on every use, so the
measurements answer one question: what the open and close per operation cost.
## Result
Median of five runs, same session, same machine.
| Benchmark | Baseline | With cache | Change |
|---|---|---|---|
| `BenchmarkRead64K` | 10888 ns/op | 9047 ns/op | -16.9 % |
| `BenchmarkWrite64K` | 5881 ns/op | 5236 ns/op | -11.0 % |
| `BenchmarkGetattr` | 461.4 ns/op | 437.2 ns/op | -5.2 % |
| `BenchmarkLookup` | 1216 ns/op | 1134 ns/op | -6.7 % |
The two data operations are the ones the cache touches, and they gain 11 to 17
percent per operation. GETATTR and LOOKUP run the same code as before the cache
(they resolve paths through Lstat either way), so their shifts are code layout
noise of the same binary, not an effect to claim; both sit within the run to run
spread the five repetitions showed.
+27
View File
@@ -0,0 +1,27 @@
# Measurement: the paginated READDIR
- Date: 2026-09-22
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), otherwise idle
- Toolchain: go1.27.1, no build flags
- Command: `go test ./internal/nfsfs/ -run '^$' -bench=ReadDirPage -benchmem -count=1 -benchtime=50x`
## Baseline
The commit 340783e, measured in the same session as the new code. The
benchmark pages a 10 000 entry directory 64 entries at a time, which is the
shape of a client listing a large directory through READDIR pages: the
baseline re listed and re sorted the whole directory for every page.
## Result
Median free, 50 iterations per side, one session.
| Benchmark | Baseline | With the listing cache | Change |
|---|---|---|---|
| `BenchmarkReadDirPage64` | 1863611 ns/op, 1402 KiB/op, 20290 allocs/op | 62506 ns/op, 40 KiB/op, 267 allocs/op | -96.6 % |
A page of 64 costs 30 times less once the sorted order is cached and
revalidated against the directory's modification time, and the cost no
longer grows with the size of the directory: the numbers above are the
boundary case, where a page paid for listing and sorting ten thousand names
to serve sixty four of them.
+39
View File
@@ -0,0 +1,39 @@
# Measurement: the request path without waste
- Date: 2026-09-22
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), otherwise idle
- Toolchain: go1.27.1, no build flags
- Command: `go test ./internal/nfs4server/ -run '^$' -bench=Wire -benchmem -count=5`
## Baseline
The commit c50cd0b, measured in the same session as the new code, both sides
five runs back to back. The baseline is the wire baseline of
[2026-09-22-wire-baseline.md](2026-09-22-wire-baseline.md) plus the client
file commands.
## Result
Median of five runs per side, one session.
| Benchmark | Latency old | Latency new | Change | Allocs old | Allocs new | Change |
|---|---|---|---|---|---|---|
| `BenchmarkWireRead64K` | 92.0 µs | 81.5 µs | -11.4 % | 78 | 63 | -19 % |
| `BenchmarkWireWrite64K` | 61.4 µs | 72.2 µs | see note | 76 | 62 | -18 % |
| `BenchmarkWireGetattr` | 14.2 µs | 13.5 µs | -4.5 % | 85 | 70 | -18 % |
| `BenchmarkWireLookup` | 15.2 µs | 15.3 µs | 0 % | 81 | 66 | -19 % |
Bytes per operation fell 27 percent on READ and 13 percent on WRITE. The
changes behind the numbers: the COMPOUND answer accumulates in one buffer
with the header patched in place instead of copying every operation result
twice, READ fills the reply buffer through the backend's `ReadInto` instead
of an intermediate allocation and drops the second attribute call for the
end of file flag, WRITE hands the request record's own bytes to the backend
instead of copying them out, and the outgoing record marking buffers recycle
through a bounded pool.
The WRITE latency column carries an honest warning: the machine's session to
session variance on this benchmark exceeds the effect being measured. Within
a single session the order of the two sides flipped twice; the deterministic
counters, allocations and bytes, are the trustworthy part of the WRITE row,
and the READ row's improvement is well outside the noise.
+30
View File
@@ -0,0 +1,30 @@
# Measurement: the wire level baseline
- Date: 2026-09-22
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), idle
- Toolchain: go1.27.1, no build flags
- Command: `go test ./internal/nfs4server/ -run '^$' -bench=Wire -benchmem -count=5`
## Baseline
The commit 354760b, the head of development: the wire benchmarks are new, so this
report is the baseline every later optimisation of the request path measures
against. One benchmark iteration is one COMPOUND of the client library against
the server handler over a loopback connection, sessions included.
## Result
Median of five runs, same session, same machine.
| Benchmark | Throughput | Latency | Allocations |
|---|---|---|---|
| `BenchmarkWireRead64K` | 805 MB/s | 81.4 µs/op | 78 allocs, 628 KiB/op |
| `BenchmarkWireWrite64K` | 1062 MB/s | 61.7 µs/op | 76 allocs, 428 KiB/op |
| `BenchmarkWireGetattr` | - | 14.2 µs/op | 85 allocs, 4.1 KiB/op |
| `BenchmarkWireLookup` | - | 15.0 µs/op | 81 allocs, 4.4 KiB/op |
READ of a 64 KiB chunk is slower than WRITE of the same chunk, and the
allocation columns show where the request path spends its memory: around 80
allocations per COMPOUND regardless of the operation, on top of the data copies
the read and write paths make. Both facts are the starting point for the
optimisations of the request path; neither is a claim about any other setup.