Initial commit
Test / test (push) Successful in 7m5s
Release / gates (push) Successful in 7m28s
Release / build (amd64, freebsd) (push) Successful in 2m52s
Release / build (amd64, linux) (push) Successful in 2m46s
Release / build (arm64, freebsd) (push) Successful in 2m22s
Release / build (arm64, linux) (push) Successful in 2m38s
Release / build (loong64, linux) (push) Successful in 2m7s
Release / build (riscv64, linux) (push) Successful in 2m17s
Release / release (push) Successful in 1m0s

Assisted-by: GLM 5.3
This commit is contained in:
2026-09-29 10:03:32 +02:00
commit f8ed33df83
206 changed files with 44165 additions and 0 deletions
+92
View File
@@ -0,0 +1,92 @@
# Benchmarking
How **Volumen** is measured. Every number a document, a README or a changelog
quotes comes from here and nowhere else.
## The method
The benchmarks live next to the code they measure. The tree carries exactly
two, and nothing else is measured:
| Benchmark | What it measures |
|---|---|
| `BenchmarkRender` | the Markdown pipeline, from source to sanitised HTML and a table of contents, for a medium body and a large one |
| `BenchmarkServer` | the wired server over real HTTP: a post list, a single post, the tag cloud, a tag feed and the sitemap, against a corpus of five hundred posts |
`BenchmarkRender` in `internal/markdown/markdown_test.go` has two
sub-benchmarks and calls `SetBytes`, so a result reads as input bytes per
second: the medium input is 2400 bytes of Markdown, the large one eighteen
copies of it, 43200 bytes.
`BenchmarkServer` in `internal/app/bench_test.go` builds a server over a
temporary content directory and drives it through an `httptest` server, so it
covers the whole chain rather than one function: routing, the middleware, the
session layer, the store, the payload builders and the renderer. It is the
workload a release is judged on, and the one a profile is recorded from.
- The machine is the development workstation, idle: AMD Ryzen AI MAX+ PRO 395
with Radeon 8060S, 16 cores and 32 threads, 117 GiB of memory, Fedora Linux
44 (`linux/amd64`). A loaded box times whatever else is running, and the
fastest sample can land on the wrong function.
- The toolchain is `go1.27.1 linux/amd64`. Benchmarks run through `go test`,
which builds the package's test binary; the command build's `-trimpath` and
`-buildvcs=true` are not part of a benchmark's build.
- Comparisons run inside one process. A loaded machine and separate processes
of identical binaries differ by more than the effects being measured, so
A/B runs alternate the two sides rather than run one after the other, and
the counts are compared through their medians, with the allocations and the
bytes per operation alongside the times. Differences within a few percent
of the spread between runs are noise; only a difference beyond that is a
result.
- When timing is hopeless, the allocation and byte counts are the result.
- A profile says where the time goes, and it is only read from an idle
machine. Record one with `-cpuprofile` on `BenchmarkServer`, then `go tool
pprof -top` it.
- Profile-guided optimisation: a profile is committed at
`cmd/volumen/default.pgo`, and the toolchain consumes it automatically when
it builds the command (measured: a `-pgo=off` build and a default build of
the same tree produce different binaries). A benchmark's test binary never
sees it, because the profile belongs to the main package, so the early A/B
run of `BenchmarkServer` compared two plain builds and could not have shown
a difference; its numbers stand as the plain build's, and they understate
the shipped binary. The comparison measured on 2026-09-25 against the two
real builds (five hundred posts, the six JSON and feed endpoints
interleaved, rate limiting off, both arms served the same twelve thousand
six hundred requests) puts the profile's effect beyond the noise: server
CPU over the warm mix is about six percent lower, the per-request median
about thirty-two percent lower on `site`, twenty-three percent on the post
list and fifteen percent on a single post, while the tag feed, the sitemap
and the cold store scan are unchanged within the noise. The profile
predates the search-relevance and custom-fields changes to the hot path.
## Running
```sh
just bench
```
The recipe runs the whole module with five counts:
```sh
go test -run '^$' -bench=. -benchmem -count=5 ./...
```
`-benchmem` is not optional: allocations per operation are part of the
result. A first look at one target, before the full battery is worth the
time:
```sh
go test -run '^$' -bench 'BenchmarkServer' -benchmem -benchtime=1x -count=1 ./internal/app
```
The full battery runs once, deliberately, on an idle machine. A benchmark
command is capped at about two minutes per round; longer sweeps are split.
A server benchmark also spends time in the kernel, so read the allocation
numbers alongside the time: a change that halves allocations and leaves the
time flat has moved the cost to the filesystem.
## Reports
The repository stores no benchmark reports. A performance claim in
`CHANGELOG.md` is measured with the method above on the change that makes
it, on the named machine, and the number travels with the claim.