feat: full NFSv4.2 server and client in pure Go
Test / test (push) Successful in 2m4s
Release / gates (push) Successful in 2m5s
Release / build (amd64, freebsd) (push) Successful in 1m27s
Release / build (amd64, linux) (push) Successful in 1m22s
Release / build (amd64, netbsd) (push) Successful in 1m19s
Release / build (amd64, openbsd) (push) Successful in 1m20s
Release / build (arm64, darwin) (push) Successful in 1m21s
Release / build (arm64, freebsd) (push) Successful in 1m26s
Release / build (arm64, linux) (push) Successful in 1m25s
Release / build (arm64, netbsd) (push) Successful in 1m31s
Release / build (arm64, openbsd) (push) Successful in 1m27s
Release / build (loong64, linux) (push) Successful in 1m37s
Release / build (riscv64, linux) (push) Successful in 1m21s
Release / release (push) Successful in 40s
Test / test (push) Successful in 2m4s
Release / gates (push) Successful in 2m5s
Release / build (amd64, freebsd) (push) Successful in 1m27s
Release / build (amd64, linux) (push) Successful in 1m22s
Release / build (amd64, netbsd) (push) Successful in 1m19s
Release / build (amd64, openbsd) (push) Successful in 1m20s
Release / build (arm64, darwin) (push) Successful in 1m21s
Release / build (arm64, freebsd) (push) Successful in 1m26s
Release / build (arm64, linux) (push) Successful in 1m25s
Release / build (arm64, netbsd) (push) Successful in 1m31s
Release / build (arm64, openbsd) (push) Successful in 1m27s
Release / build (loong64, linux) (push) Successful in 1m37s
Release / build (riscv64, linux) (push) Successful in 1m21s
Release / release (push) Successful in 40s
Assisted-by: GLM 5.3 Flash
This commit is contained in:
@@ -0,0 +1,119 @@
|
||||
# Architecture
|
||||
|
||||
How nfs is put together. Every node, package and arrow below exists in the
|
||||
source tree; nothing is aspirational.
|
||||
|
||||
## Overview
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
nfs[cmd/nfs] --> Client[internal/nfsclient]
|
||||
nfsd[cmd/nfsd] --> Server[internal/server]
|
||||
nfsd --> Backend[internal/nfsfs]
|
||||
Server --> Dispatch[internal/nfs4server]
|
||||
Dispatch --> Backend
|
||||
Dispatch --> Wire[internal/nfs4]
|
||||
Dispatch --> Record[internal/rpc]
|
||||
Client --> Wire
|
||||
Client --> Record
|
||||
Wire --> XDR[internal/xdr]
|
||||
Record --> XDR
|
||||
```
|
||||
|
||||
The project speaks NFSv4.2 only, on one TCP port, with the full state model
|
||||
of RFC 8881 under the extensions of RFC 7862 and the add-ons of RFC 8276,
|
||||
RFC 7861 and RFC 9289. The server carries the stateless operations, the
|
||||
session machinery, open and lock state, delegations with their back channel
|
||||
recalls, the optional operations, extended and named attributes, directory
|
||||
delegations with notifications, pNFS in the metadata server role, and the
|
||||
security layers: RPCSEC_GSS v1 and v3 with Kerberos in pure Go, and
|
||||
RPC-with-TLS with in place connection upgrade. The client mirrors the same
|
||||
surface and serves as the second oracle against the server. The two
|
||||
commands carry the roles: nfsd serves one local directory tree, nfs speaks
|
||||
to a server from the command line.
|
||||
|
||||
## Packages
|
||||
|
||||
| Package | Responsibility |
|
||||
|---|---|
|
||||
| `cmd/nfsd` | flags, the version report, the signal wiring; no logic |
|
||||
| `cmd/nfs` | the client command: the argument handling and the compound building for ls, cat, put and stat |
|
||||
| `internal/server` | the accept loop, connection lifetime and shutdown |
|
||||
| `internal/nfs4server` | the COMPOUND dispatcher: the file handle register, the session machinery, the operations, the mapping of backend errors to statuses |
|
||||
| `internal/nfs4` | the NFSv4.2 wire vocabulary: numbers, bitmap4, fattr4, the COMPOUND codec and the per operation arguments and results |
|
||||
| `internal/nfsfs` | the virtual filesystem interface and the local backend with dev and ino based handles |
|
||||
| `internal/rpc` | ONC RPC: record marking, the call and reply headers, AUTH_SYS credentials |
|
||||
| `internal/xdr` | the RFC 4506 primitives: integers, booleans, opaque values and strings with their padding |
|
||||
| `internal/nfsclient` | the client half; it shares the wire packages with the server and serves as the second oracle |
|
||||
| `internal/krb5` | the Kerberos crypto profiles and GSS tokens of RFC 3961, 3962, 4120 and 4121, verified against the test vectors of the RFCs and the MIT krb5 suite |
|
||||
| `internal/rdma` | the RPC-over-RDMA framing of RFC 8166: the fixed header, the chunk lists and the stream adapter; verbs live outside pure Go |
|
||||
|
||||
The boundaries follow the layering of the protocol stack: a layer speaks
|
||||
only to the one below it, and the wire packages know nothing about sockets.
|
||||
The dispatcher decides nothing about storage; the backend interface owns
|
||||
that.
|
||||
|
||||
## Data flow
|
||||
|
||||
The main operation today is one COMPOUND through the whole stack, from the
|
||||
client that is also the project's oracle:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant C as internal/nfsclient
|
||||
participant S as nfsd
|
||||
participant D as internal/nfs4server
|
||||
participant F as internal/nfsfs
|
||||
C->>S: record: COMPOUND, PUTROOTFH, LOOKUP, GETFH, GETATTR, READ
|
||||
S->>D: the record is reassembled, the call decoded
|
||||
D->>F: Root, Lookup, Getattr, Read
|
||||
F-->>D: handle, attributes, bytes
|
||||
D-->>C: one result per operation, the failing one ends the array
|
||||
```
|
||||
|
||||
READDIR carries the paging machinery of the standard: the client's cookie
|
||||
and verifier are checked against the backend's order, the entries are
|
||||
packed under the maxcount budget, and the response reports where the next
|
||||
page starts. The error paths are named: a clean end of stream before the
|
||||
first header is io.EOF, a stream that stops mid record is
|
||||
io.ErrUnexpectedEOF, a record beyond the limit is ErrRecordTooLarge, and
|
||||
every backend failure becomes the NFS4ERR status its sentinel names.
|
||||
|
||||
## State and lifetime
|
||||
|
||||
- The Server lives for the process and Serve blocks for that long; a
|
||||
cancellation of the context closes the listener and Serve returns nil.
|
||||
- Each accepted connection runs on its own goroutine and is owned by its
|
||||
Handle.
|
||||
- The file handle register is per COMPOUND: PUTROOTFH, PUTFH, SAVEFH and
|
||||
RESTOREFH move handles through it, and nothing of it survives the call.
|
||||
- Session state lives in the session store: the slot table with the reply
|
||||
cache, the lease of every client, the open and lock state with their
|
||||
stateids, delegations, directory delegations and the GSS contexts of
|
||||
RPCSEC_GSS. Referral stubs and named attribute handles are synthetic and
|
||||
live in their own stores.
|
||||
- The handles of the local backend encode the device and inode number and
|
||||
resolve through an in memory map persisted on demand: a handle from
|
||||
before a server restart resolves again once the mapping is loaded back,
|
||||
and one whose object is gone answers NFS4ERR_STALE.
|
||||
- WRITE answers FILE_SYNC with the boot verifier, so no unstable writes
|
||||
outlive a restart and the client keeps no replay debt. The stateid of a
|
||||
WRITE is validated: a real one must name a live OPEN of the current file,
|
||||
and the anonymous forms pass without state.
|
||||
- WRITE and CREATE reach the backend through the nfsfs Writer interface; a
|
||||
backend that does not implement it is answered NFS4ERR_ROFS. The local
|
||||
backend implements both halves: CREATE makes directories, symlinks,
|
||||
fifos, sockets and device nodes, and a regular file is the business of
|
||||
OPEN.
|
||||
|
||||
## Dependencies
|
||||
|
||||
One dependency outside the standard library: interpres
|
||||
(`sourcedock.dev/petrbalvin/interpres/v2`), the TOML reader of the
|
||||
configuration file, itself built on the standard library alone. The
|
||||
protocol stack is carried by hand written code, because the server
|
||||
targets Linux on amd64, arm64, loong64 and riscv64; FreeBSD,
|
||||
OpenBSD and NetBSD on amd64 and arm64; and darwin on arm64,
|
||||
and a dependency
|
||||
that breaks one of those platforms is a dependency the project cannot
|
||||
carry.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Benchmarking
|
||||
|
||||
How the nfs project is measured. Every number a document, a README or a changelog
|
||||
quotes comes from here and nowhere else.
|
||||
|
||||
## The method
|
||||
|
||||
- Two levels are measured, and they answer different questions. The backend level
|
||||
benchmarks the `internal/nfsfs` filesystem layer on its own: what one READ, WRITE,
|
||||
GETATTR or LOOKUP costs against a local directory. The wire level benchmarks whole
|
||||
NFS sessions: the server binary and the client library over a loopback connection,
|
||||
which adds the RPC, XDR and session layers on top of the backend.
|
||||
- What is deliberately left out: kernel NFS mounts, the network beyond loopback, and
|
||||
any comparison against other NFS servers. Those are interop questions, not
|
||||
benchmark questions.
|
||||
- The machine is idle, named, and stays the same across comparable reports.
|
||||
- The toolchain is named with its version and its build flags.
|
||||
- Comparisons run inside one process, with the order of the two sides alternated
|
||||
where a comparison is the point. Differences under two percent are noise, not
|
||||
results.
|
||||
- A change is measured against its baseline, not against a memory of how fast it
|
||||
used to be. The baseline run is part of the measurement, and both sides land in
|
||||
the same report.
|
||||
|
||||
## Running
|
||||
|
||||
```sh
|
||||
just bench
|
||||
```
|
||||
|
||||
The recipe sweeps `./internal/...` with `-benchmem -count=5`.
|
||||
|
||||
A first look at one target, before the full battery is worth the time:
|
||||
|
||||
```sh
|
||||
go test -run '^$' -bench 'BenchmarkRead64K' -benchtime=1x ./internal/nfsfs
|
||||
```
|
||||
|
||||
The full battery runs once, deliberately, on an idle machine. A benchmark command
|
||||
is capped at about two minutes per round; longer sweeps are split.
|
||||
|
||||
## Reports
|
||||
|
||||
Reports live in `docs/_results/`, one file per measurement round, named
|
||||
`YYYY-MM-DD-subject.md`, and follow [BENCHMARK_TEMPLATE.md](BENCHMARK_TEMPLATE.md).
|
||||
A report carries its numbers, its machine, its toolchain and the exact command. A
|
||||
number without its provenance is not a result, and a performance claim without a
|
||||
report behind it is left out of the documentation rather than softened into an
|
||||
adjective.
|
||||
+90
@@ -0,0 +1,90 @@
|
||||
# Command line
|
||||
|
||||
The reference below is taken from the programs' own help output. If the two
|
||||
disagree, the programs are right and this file is a defect. Both commands
|
||||
carry a manpage: [man/nfsd.1](../man/nfsd.1) and [man/nfs.1](../man/nfs.1).
|
||||
|
||||
## Synopsis
|
||||
|
||||
```sh
|
||||
nfsd [-addr addr] [-export dir] [-version]
|
||||
nfs [-addr host:port] [-concurrency 1-8] version | ls [path] | cat path | put local remote | get remote local | rm path | mkdir path | stat path | selftest
|
||||
```
|
||||
|
||||
## nfsd
|
||||
|
||||
The server exports one local directory tree over NFSv4.2 on a single TCP
|
||||
port and serves it read and write. It exits cleanly on SIGINT and SIGTERM.
|
||||
|
||||
| Flag | Default | Effect |
|
||||
|---|---|---|
|
||||
| `-addr` | `:2049` | the TCP address to listen on |
|
||||
| `-export` | | the directory to serve; required, and the path must be an existing directory |
|
||||
| `-ro` | `false` | serve the export read only: every mutation answers NFS4ERR_ROFS, the reads of every half work unchanged |
|
||||
| `-root-squash` | `false` | map a client claiming uid 0 onto nobody (65534): no superuser grant, and objects root creates carry nobody |
|
||||
| `-tls-cert` | | the certificate chain in PEM for RPC-with-TLS (RFC 9289); requires `-tls-key` |
|
||||
| `-tls-key` | | the private key in PEM for RPC-with-TLS; requires `-tls-cert` |
|
||||
| `-log-ops` | `false` | log every operation to stderr as `nfs: OP status N duration` |
|
||||
| `-max-connections` | `0` | cap on live connections; a connection above the cap closes at once; `0` means no cap |
|
||||
| `-state-dir` | | directory for persisted client state: file handles and opens are written there as they change, a restart loads them back and serves a grace window, so a client reclaims its open with CLAIM_PREVIOUS; without it nothing persists |
|
||||
| `-config` | | configuration file in TOML, see [CONFIGURATION.md](CONFIGURATION.md); never read unless named, the flags override it |
|
||||
| `-version` | | print the version and exit |
|
||||
|
||||
```sh
|
||||
$ ./bin/nfsd -export /srv/demo -addr 127.0.0.1:2049
|
||||
nfsd: serving /srv/demo on 127.0.0.1:2049
|
||||
```
|
||||
|
||||
## nfs
|
||||
|
||||
The client runs one operation against an NFSv4.2 server and works against
|
||||
any server that speaks the minor version, `nfsd` included. Paths address the
|
||||
server's namespace from its root.
|
||||
|
||||
| Command | Purpose |
|
||||
|---|---|
|
||||
| `version` | print the version |
|
||||
| `ls [path]` | list a directory, the root by default; one line per entry: name, size, octal mode |
|
||||
| `cat path` | stream a file to standard output |
|
||||
| `put local remote` | write a local file to the server; creates with mode 0644 and truncates first, so a shorter file leaves no tail |
|
||||
| `get remote local` | copy a remote file into a local file; the local file is truncated first, so a shorter remote leaves no tail |
|
||||
| `rm path` | remove one object from the server |
|
||||
| `mkdir path` | make one directory on the server |
|
||||
| `stat path` | print the type, size, octal mode and modification time of an object |
|
||||
| `selftest` | run the whole operation matrix against the server: mkdir, touch, write and compare 64 KiB, list, rename, symlink, a nested directory, ownership of files created as another uid, setattr, and the removals; one line per check plus a summary, exit 1 when any check fails. The work directory is removed on success and left in place on failure |
|
||||
|
||||
| Flag | Default | Effect |
|
||||
|---|---|---|
|
||||
| `-addr` | `127.0.0.1:2049` | the server address |
|
||||
| `-concurrency` | `1` | compounds in flight for `get` and `put`, 1 to 8; more slots move several chunks at once over the one connection |
|
||||
|
||||
```sh
|
||||
$ ./bin/nfs -addr 127.0.0.1:2049 ls
|
||||
hello.txt 5 644
|
||||
$ ./bin/nfs -addr 127.0.0.1:2049 cat /hello.txt
|
||||
ahoj
|
||||
```
|
||||
|
||||
## Exit codes
|
||||
|
||||
| Code | nfsd | nfs |
|
||||
|---|---|---|
|
||||
| `0` | clean shutdown on SIGINT or SIGTERM, or the version print | the operation completed |
|
||||
| `1` | no export, a bad export path, a listen failure or a listener failure | the dial, the session or the operation failed |
|
||||
| `2` | | wrong arguments: no, unknown or starved subcommand |
|
||||
|
||||
## Examples
|
||||
|
||||
Serve a tree and read it from another terminal:
|
||||
|
||||
```sh
|
||||
mkdir -p /srv/demo && echo "ahoj" > /srv/demo/hello.txt
|
||||
./bin/nfsd -export /srv/demo -addr 127.0.0.1:2049
|
||||
```
|
||||
|
||||
```sh
|
||||
./bin/nfs -addr 127.0.0.1:2049 ls
|
||||
./bin/nfs -addr 127.0.0.1:2049 cat /hello.txt
|
||||
./bin/nfs -addr 127.0.0.1:2049 put README.md /readme.md
|
||||
./bin/nfs -addr 127.0.0.1:2049 stat /readme.md
|
||||
```
|
||||
@@ -0,0 +1,58 @@
|
||||
# Configuration
|
||||
|
||||
nfsd reads its configuration from the file the `-config` flag names, in TOML.
|
||||
Without `-config` no file is read and every setting comes from the flags and
|
||||
the built-in defaults; the file is never looked for in a default location.
|
||||
|
||||
## File
|
||||
|
||||
A complete example with every key present:
|
||||
|
||||
```toml
|
||||
listen = ":2049"
|
||||
log-ops = false
|
||||
state-dir = ""
|
||||
max-connections = 0
|
||||
|
||||
[tls]
|
||||
cert = ""
|
||||
key = ""
|
||||
|
||||
[[export]]
|
||||
path = "/srv/demo"
|
||||
read-only = false
|
||||
root-squash = false
|
||||
```
|
||||
|
||||
## Keys
|
||||
|
||||
| Key | Type | Default | Effect |
|
||||
|---|---|---|---|
|
||||
| `listen` | string | `":2049"` | the TCP address to listen on, the `-addr` flag |
|
||||
| `log-ops` | boolean | `false` | log every operation to stderr, the `-log-ops` flag |
|
||||
| `state-dir` | string | `""` | the directory for persisted handles and opens, the `-state-dir` flag; empty means nothing persists |
|
||||
| `max-connections` | integer | `0` | the cap on live connections, the `-max-connections` flag; `0` means no cap |
|
||||
| `tls.cert` | string | `""` | the certificate chain in PEM for RPC-with-TLS (RFC 9289), the `-tls-cert` flag |
|
||||
| `tls.key` | string | `""` | the private key in PEM for RPC-with-TLS, the `-tls-key` flag |
|
||||
| `export.path` | string | | the directory to serve; required, the `-export` flag |
|
||||
| `export.read-only` | boolean | `false` | serve the export read only, the `-ro` flag |
|
||||
| `export.root-squash` | boolean | `false` | map a client claiming uid 0 onto nobody (65534), the `-root-squash` flag; the default keeps the trust AUTH_SYS gives to the claim, and operators serving untrusted clients are advised to turn it on |
|
||||
|
||||
The `[[export]]` array carries exactly one table: this server serves one
|
||||
export. A future release that serves several exports lifts the count without
|
||||
changing the schema.
|
||||
|
||||
## Precedence
|
||||
|
||||
The command line flags win, then the file, then the built-in defaults. A flag
|
||||
present on the command line overrides the file even when it carries the
|
||||
default value, so `-ro=false` keeps a `read-only = true` from the file at
|
||||
`false`. A key the file leaves out yields to the flag default.
|
||||
|
||||
## Validation
|
||||
|
||||
A file that fails is a failed start up. A syntax error is reported with the
|
||||
file and the line: `nfsd: /etc/nfsd/nfsd.toml:2: expected '=' after key`. A
|
||||
key the schema does not carry is rejected, so a typo never slips through as
|
||||
an ignored setting. A file without exactly one `[[export]]`, or one without
|
||||
`path`, ends the start up with a message naming the file and the count.
|
||||
@@ -0,0 +1,168 @@
|
||||
# Deployment
|
||||
|
||||
How nfsd runs in production.
|
||||
|
||||
## Topology
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
C1[NFS client] --> S[nfsd]
|
||||
C2[NFS client] --> S
|
||||
S --> Disk[(export tree)]
|
||||
S --> State[(state dir)]
|
||||
```
|
||||
|
||||
One nfsd process serves one exported directory tree to any number of NFSv4.2
|
||||
clients over the single TCP port 2049. There is no portmapper, no mountd and
|
||||
no separate locking protocol: a client mounts `nfs://host:2049/` directly and
|
||||
everything rides the one connection or its successors.
|
||||
|
||||
## Requirements
|
||||
|
||||
- a Linux host, on amd64, arm64, loong64 or riscv64; the binary is static, no
|
||||
runtime libraries
|
||||
- port 2049 free; it is not a privileged port, so the service does not need
|
||||
root
|
||||
- the exported directory must exist before start; the service identity needs
|
||||
read access to it, and write access where clients may write
|
||||
- a writable state directory, when persistence is on, writable by the service
|
||||
identity alone (mode 0700)
|
||||
- write access to the export tree requires one of: the service runs as root,
|
||||
or the service holds `CAP_CHOWN` and `CAP_MKNOD`, or the operator accepts
|
||||
the identity behaviour described under Privileges
|
||||
|
||||
## Build
|
||||
|
||||
```sh
|
||||
just build
|
||||
```
|
||||
|
||||
The binaries land in `bin/nfsd` and `bin/nfs`.
|
||||
|
||||
## Run
|
||||
|
||||
```sh
|
||||
bin/nfsd -config /etc/nfsd/nfsd.toml
|
||||
```
|
||||
|
||||
The configuration file is described in [CONFIGURATION.md](CONFIGURATION.md);
|
||||
it carries the listen address, the export, the state directory, the
|
||||
connection cap, the TLS key pair and the root squash switch. The flags
|
||||
override the file.
|
||||
|
||||
## Service unit
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=NFSv4.2 server for one export
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=notify
|
||||
User=nfsd
|
||||
Group=nfsd
|
||||
ExecStart=/usr/local/bin/nfsd -config /etc/nfsd/nfsd.toml
|
||||
StateDirectory=nfsd
|
||||
AmbientCapabilities=CAP_CHOWN CAP_MKNOD
|
||||
CapabilityBoundingSet=CAP_CHOWN CAP_MKNOD
|
||||
NoNewPrivileges=yes
|
||||
ProtectSystem=strict
|
||||
ReadWritePaths=/srv/export
|
||||
Restart=on-failure
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
`Type=notify` is real readiness: the server writes `READY=1` to
|
||||
`$NOTIFY_SOCKET` once the listener is up, and systemd considers the unit
|
||||
started at that moment, not at the fork. `StateDirectory=nfsd` creates
|
||||
`/var/lib/nfsd` owned by the service identity; point the configuration's
|
||||
`state-dir` at it and handles and opens survive a restart inside the grace
|
||||
window.
|
||||
|
||||
### Privileges
|
||||
|
||||
The port needs no privilege, so the unit runs under a dedicated identity and
|
||||
names exactly two capabilities:
|
||||
|
||||
- `CAP_CHOWN` lets the server hand a freshly created object to the identity
|
||||
the client presented. Without it the object keeps the service identity;
|
||||
the server still answers the client's claim as the owner attribute, but
|
||||
the on disk owner is the service one. This is a deliberate operator
|
||||
decision, documented here and not hidden: serving untrusted clients
|
||||
without `CAP_CHOWN` changes whose identity new files carry on disk.
|
||||
- `CAP_MKNOD` serves the special objects a CREATE may carry, character and
|
||||
block devices among them. Without it those creations fail.
|
||||
|
||||
Neither capability lets the service read a file it could not already reach.
|
||||
|
||||
### Root squash
|
||||
|
||||
Serve untrusted clients with `root-squash = true` in the export: a client
|
||||
claiming uid 0 acts as nobody (65534), loses the superuser grant, and its
|
||||
objects carry nobody. The default is `false`, which keeps the trust AUTH_SYS
|
||||
hands to the claim; an operator who controls every client may keep it.
|
||||
|
||||
## Production configuration
|
||||
|
||||
```toml
|
||||
listen = ":2049"
|
||||
state-dir = "/var/lib/nfsd"
|
||||
max-connections = 256
|
||||
log-ops = false
|
||||
|
||||
[tls]
|
||||
cert = "/etc/nfsd/cert.pem"
|
||||
key = "/etc/nfsd/key.pem"
|
||||
|
||||
[[export]]
|
||||
path = "/srv/export"
|
||||
root-squash = true
|
||||
```
|
||||
|
||||
The TLS key pair enables RPC-with-TLS of RFC 9289; a client that skips the
|
||||
upgrade is refused. The certificate comes from the operator's PKI; no
|
||||
credential belongs in this repository or its configuration examples.
|
||||
|
||||
## Firewall
|
||||
|
||||
One port in, no outbound requirement beyond what the clients reach the back
|
||||
channel on: the server calls the client back on the client's connection, so
|
||||
no inbound port per client is needed.
|
||||
|
||||
```sh
|
||||
firewall-cmd --permanent --add-port=2049/tcp && firewall-cmd --reload
|
||||
```
|
||||
|
||||
## Upgrade
|
||||
|
||||
```sh
|
||||
just build
|
||||
install -m 755 bin/nfsd /usr/local/bin/nfsd
|
||||
systemctl restart nfsd
|
||||
```
|
||||
|
||||
With `state-dir` set the restart is a recovery, not a loss: the new process
|
||||
loads the handle map and the opens, and clients inside the grace window
|
||||
reclaim with CLAIM_PREVIOUS. Without it, clients re establish their sessions
|
||||
and re open; their mounted trees keep working through the file handles the
|
||||
backend re registers.
|
||||
|
||||
## Rollback
|
||||
|
||||
Reinstall the previous binary and restart; the state directory format has
|
||||
one version so far, so a downgrade reloads the same state. Not rehearsed
|
||||
against a released predecessor yet: rehearse before relying on it.
|
||||
|
||||
## Monitoring
|
||||
|
||||
- the unit's ready state: `systemctl is-active nfsd`
|
||||
- the operation log, when `log-ops` is on: every operation, its status and
|
||||
its duration on standard error, journald's `journalctl -u nfsd` picks it up
|
||||
- the connection cap answering refusals shows up as clients reconnecting;
|
||||
a steady refusal rate means the cap or the client count is wrong
|
||||
- a healthy idle server logs nothing and holds no CPU: check
|
||||
`systemctl status nfsd` for a flat memory figure and `ss -tnp sport = 2049`
|
||||
for the connected clients
|
||||
@@ -0,0 +1,90 @@
|
||||
# Development
|
||||
|
||||
How to work on nfs.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Go 1.27.1, the newest stable release.
|
||||
- [just](https://github.com/casey/just) for the recipes.
|
||||
- gcc, for the race detector in `just gates`.
|
||||
|
||||
## Setup
|
||||
|
||||
```sh
|
||||
git clone https://sourcedock.dev/petrbalvin/nfs.git
|
||||
cd nfs
|
||||
just build
|
||||
```
|
||||
|
||||
## Recipes
|
||||
|
||||
Every recipe in the project's file, and what it does. Taken from the file itself, so
|
||||
the names and the list match it exactly.
|
||||
|
||||
| Recipe | What it does |
|
||||
|---|---|
|
||||
| `just gates` | the definition of done: build, format check, vet, modernisation, the test suite with the coverage floor, and the race detector |
|
||||
| `just build` | compiles `cmd/nfsd` and `cmd/nfs` into `bin/nfsd` and `bin/nfs`, zero errors and zero warnings |
|
||||
| `just test` | the full suite with no test cache and the 80 percent coverage floor |
|
||||
| `just race` | the same suite under the race detector |
|
||||
| `just unit ./internal/xdr 'TestName'` | a fast scoped run for iterating |
|
||||
| `just fuzz FuzzXdr ./internal/xdr 60s` | a time boxed fuzz of one target in one package |
|
||||
| `just bench` | the benchmarks, five counts, allocation stats on |
|
||||
| `just fmt` | gofmt over the tree, in place |
|
||||
| `just fmt-check` | zero diff, prints nothing when everything is formatted |
|
||||
| `just vet` | `go vet` and `go fix -diff` |
|
||||
| `just clean` | removes `bin/` and `coverage.out` |
|
||||
| `just install` | builds, then copies `bin/nfsd` and `bin/nfs` into the user's bin directory |
|
||||
| `just uninstall` | removes the installed binary |
|
||||
| `just run` | runs the program in place; nfsd exits at once until it is given an export, so a real run passes flags to the built binary: `./bin/nfsd -export DIR` |
|
||||
| `just dev` | the same as `run`, for now |
|
||||
|
||||
The test and bench recipes sweep `./internal/...`, and not
|
||||
the whole tree: the thin `cmd/nfsd` and `cmd/nfs` count as zero coverage and
|
||||
would drag the floor below 80 percent on their own. The protocol logic lives
|
||||
under `internal/`.
|
||||
|
||||
## Running a single test
|
||||
|
||||
```sh
|
||||
go test -run TestName ./package
|
||||
```
|
||||
|
||||
Add `-v` for the sub-test names, and `-race` when the change touches concurrency.
|
||||
`-count=1` defeats the test cache when a result looks stale.
|
||||
|
||||
## Coverage
|
||||
|
||||
```sh
|
||||
just test
|
||||
go tool cover -func=coverage.out
|
||||
```
|
||||
|
||||
The `total:` line is the number that matters, and it stays at 80 percent or more.
|
||||
|
||||
## Benchmarks
|
||||
|
||||
```sh
|
||||
just bench
|
||||
```
|
||||
|
||||
Benchmark on an idle machine, and compare only runs made in one process against each other.
|
||||
|
||||
## Debugging the build
|
||||
|
||||
```sh
|
||||
go build -gcflags='-m' ./... # inlining decisions
|
||||
go build -gcflags='-S' ./... # what the compiler generated
|
||||
```
|
||||
|
||||
## Continuous integration
|
||||
|
||||
Workflows live in `.gitea/workflows/` and run on the project's own runners. They are
|
||||
written by hand rather than through `just`, but they enforce the same set of gates, so a
|
||||
green `just gates` locally is the fastest way to a green pipeline.
|
||||
|
||||
## Releases
|
||||
|
||||
Releases are cut by merging `development` into `main` and tagging `vX.Y.Z`. The tag
|
||||
drives the release workflow, which builds the assets and publishes the notes it
|
||||
extracted from `CHANGELOG.md`.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Measurement: the descriptor cache in internal/nfsfs
|
||||
|
||||
- Date: 2026-09-22
|
||||
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), idle
|
||||
- Toolchain: go1.27.1, no build flags
|
||||
- Command: `go test ./internal/nfsfs/ -run '^$' -bench=. -benchmem -count=5`
|
||||
|
||||
## Baseline
|
||||
|
||||
The commit b8758b7, the head of development before the descriptor cache: every
|
||||
READ, WRITE and SYNC opened the registered path, verified it with a stat and
|
||||
closed the descriptor again, per operation. The cache keeps idle descriptors of
|
||||
regular files in a bounded LRU and revalidates the identity on every use, so the
|
||||
measurements answer one question: what the open and close per operation cost.
|
||||
|
||||
## Result
|
||||
|
||||
Median of five runs, same session, same machine.
|
||||
|
||||
| Benchmark | Baseline | With cache | Change |
|
||||
|---|---|---|---|
|
||||
| `BenchmarkRead64K` | 10888 ns/op | 9047 ns/op | -16.9 % |
|
||||
| `BenchmarkWrite64K` | 5881 ns/op | 5236 ns/op | -11.0 % |
|
||||
| `BenchmarkGetattr` | 461.4 ns/op | 437.2 ns/op | -5.2 % |
|
||||
| `BenchmarkLookup` | 1216 ns/op | 1134 ns/op | -6.7 % |
|
||||
|
||||
The two data operations are the ones the cache touches, and they gain 11 to 17
|
||||
percent per operation. GETATTR and LOOKUP run the same code as before the cache
|
||||
(they resolve paths through Lstat either way), so their shifts are code layout
|
||||
noise of the same binary, not an effect to claim; both sit within the run to run
|
||||
spread the five repetitions showed.
|
||||
@@ -0,0 +1,27 @@
|
||||
# Measurement: the paginated READDIR
|
||||
|
||||
- Date: 2026-09-22
|
||||
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), otherwise idle
|
||||
- Toolchain: go1.27.1, no build flags
|
||||
- Command: `go test ./internal/nfsfs/ -run '^$' -bench=ReadDirPage -benchmem -count=1 -benchtime=50x`
|
||||
|
||||
## Baseline
|
||||
|
||||
The commit 340783e, measured in the same session as the new code. The
|
||||
benchmark pages a 10 000 entry directory 64 entries at a time, which is the
|
||||
shape of a client listing a large directory through READDIR pages: the
|
||||
baseline re listed and re sorted the whole directory for every page.
|
||||
|
||||
## Result
|
||||
|
||||
Median free, 50 iterations per side, one session.
|
||||
|
||||
| Benchmark | Baseline | With the listing cache | Change |
|
||||
|---|---|---|---|
|
||||
| `BenchmarkReadDirPage64` | 1863611 ns/op, 1402 KiB/op, 20290 allocs/op | 62506 ns/op, 40 KiB/op, 267 allocs/op | -96.6 % |
|
||||
|
||||
A page of 64 costs 30 times less once the sorted order is cached and
|
||||
revalidated against the directory's modification time, and the cost no
|
||||
longer grows with the size of the directory: the numbers above are the
|
||||
boundary case, where a page paid for listing and sorting ten thousand names
|
||||
to serve sixty four of them.
|
||||
@@ -0,0 +1,39 @@
|
||||
# Measurement: the request path without waste
|
||||
|
||||
- Date: 2026-09-22
|
||||
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), otherwise idle
|
||||
- Toolchain: go1.27.1, no build flags
|
||||
- Command: `go test ./internal/nfs4server/ -run '^$' -bench=Wire -benchmem -count=5`
|
||||
|
||||
## Baseline
|
||||
|
||||
The commit c50cd0b, measured in the same session as the new code, both sides
|
||||
five runs back to back. The baseline is the wire baseline of
|
||||
[2026-09-22-wire-baseline.md](2026-09-22-wire-baseline.md) plus the client
|
||||
file commands.
|
||||
|
||||
## Result
|
||||
|
||||
Median of five runs per side, one session.
|
||||
|
||||
| Benchmark | Latency old | Latency new | Change | Allocs old | Allocs new | Change |
|
||||
|---|---|---|---|---|---|---|
|
||||
| `BenchmarkWireRead64K` | 92.0 µs | 81.5 µs | -11.4 % | 78 | 63 | -19 % |
|
||||
| `BenchmarkWireWrite64K` | 61.4 µs | 72.2 µs | see note | 76 | 62 | -18 % |
|
||||
| `BenchmarkWireGetattr` | 14.2 µs | 13.5 µs | -4.5 % | 85 | 70 | -18 % |
|
||||
| `BenchmarkWireLookup` | 15.2 µs | 15.3 µs | 0 % | 81 | 66 | -19 % |
|
||||
|
||||
Bytes per operation fell 27 percent on READ and 13 percent on WRITE. The
|
||||
changes behind the numbers: the COMPOUND answer accumulates in one buffer
|
||||
with the header patched in place instead of copying every operation result
|
||||
twice, READ fills the reply buffer through the backend's `ReadInto` instead
|
||||
of an intermediate allocation and drops the second attribute call for the
|
||||
end of file flag, WRITE hands the request record's own bytes to the backend
|
||||
instead of copying them out, and the outgoing record marking buffers recycle
|
||||
through a bounded pool.
|
||||
|
||||
The WRITE latency column carries an honest warning: the machine's session to
|
||||
session variance on this benchmark exceeds the effect being measured. Within
|
||||
a single session the order of the two sides flipped twice; the deterministic
|
||||
counters, allocations and bytes, are the trustworthy part of the WRITE row,
|
||||
and the READ row's improvement is well outside the noise.
|
||||
@@ -0,0 +1,30 @@
|
||||
# Measurement: the wire level baseline
|
||||
|
||||
- Date: 2026-09-22
|
||||
- Machine: AMD Ryzen AI Max+ Pro 395 (32 threads), idle
|
||||
- Toolchain: go1.27.1, no build flags
|
||||
- Command: `go test ./internal/nfs4server/ -run '^$' -bench=Wire -benchmem -count=5`
|
||||
|
||||
## Baseline
|
||||
|
||||
The commit 354760b, the head of development: the wire benchmarks are new, so this
|
||||
report is the baseline every later optimisation of the request path measures
|
||||
against. One benchmark iteration is one COMPOUND of the client library against
|
||||
the server handler over a loopback connection, sessions included.
|
||||
|
||||
## Result
|
||||
|
||||
Median of five runs, same session, same machine.
|
||||
|
||||
| Benchmark | Throughput | Latency | Allocations |
|
||||
|---|---|---|---|
|
||||
| `BenchmarkWireRead64K` | 805 MB/s | 81.4 µs/op | 78 allocs, 628 KiB/op |
|
||||
| `BenchmarkWireWrite64K` | 1062 MB/s | 61.7 µs/op | 76 allocs, 428 KiB/op |
|
||||
| `BenchmarkWireGetattr` | - | 14.2 µs/op | 85 allocs, 4.1 KiB/op |
|
||||
| `BenchmarkWireLookup` | - | 15.0 µs/op | 81 allocs, 4.4 KiB/op |
|
||||
|
||||
READ of a 64 KiB chunk is slower than WRITE of the same chunk, and the
|
||||
allocation columns show where the request path spends its memory: around 80
|
||||
allocations per COMPOUND regardless of the operation, on top of the data copies
|
||||
the read and write paths make. Both facts are the starting point for the
|
||||
optimisations of the request path; neither is a claim about any other setup.
|
||||
Reference in New Issue
Block a user