93 lines
4.5 KiB
Markdown
93 lines
4.5 KiB
Markdown
# Benchmarking
|
|||
|
|
|
||
|
|
How **Volumen** is measured. Every number a document, a README or a changelog
|
||
|
|
quotes comes from here and nowhere else.
|
||
|
|
|
||
|
|
## The method
|
||
|
|
|
||
|
|
The benchmarks live next to the code they measure. The tree carries exactly
|
||
|
|
two, and nothing else is measured:
|
||
|
|
|
||
|
|
| Benchmark | What it measures |
|
||
|
|
|---|---|
|
||
|
|
| `BenchmarkRender` | the Markdown pipeline, from source to sanitised HTML and a table of contents, for a medium body and a large one |
|
||
|
|
| `BenchmarkServer` | the wired server over real HTTP: a post list, a single post, the tag cloud, a tag feed and the sitemap, against a corpus of five hundred posts |
|
||
|
|
|
||
|
|
`BenchmarkRender` in `internal/markdown/markdown_test.go` has two
|
||
|
|
sub-benchmarks and calls `SetBytes`, so a result reads as input bytes per
|
||
|
|
second: the medium input is 2400 bytes of Markdown, the large one eighteen
|
||
|
|
copies of it, 43200 bytes.
|
||
|
|
|
||
|
|
`BenchmarkServer` in `internal/app/bench_test.go` builds a server over a
|
||
|
|
temporary content directory and drives it through an `httptest` server, so it
|
||
|
|
covers the whole chain rather than one function: routing, the middleware, the
|
||
|
|
session layer, the store, the payload builders and the renderer. It is the
|
||
|
|
workload a release is judged on, and the one a profile is recorded from.
|
||
|
|
|
||
|
|
- The machine is the development workstation, idle: AMD Ryzen AI MAX+ PRO 395
|
||
|
|
with Radeon 8060S, 16 cores and 32 threads, 117 GiB of memory, Fedora Linux
|
||
|
|
44 (`linux/amd64`). A loaded box times whatever else is running, and the
|
||
|
|
fastest sample can land on the wrong function.
|
||
|
|
- The toolchain is `go1.27.1 linux/amd64`. Benchmarks run through `go test`,
|
||
|
|
which builds the package's test binary; the command build's `-trimpath` and
|
||
|
|
`-buildvcs=true` are not part of a benchmark's build.
|
||
|
|
- Comparisons run inside one process. A loaded machine and separate processes
|
||
|
|
of identical binaries differ by more than the effects being measured, so
|
||
|
|
A/B runs alternate the two sides rather than run one after the other, and
|
||
|
|
the counts are compared through their medians, with the allocations and the
|
||
|
|
bytes per operation alongside the times. Differences within a few percent
|
||
|
|
of the spread between runs are noise; only a difference beyond that is a
|
||
|
|
result.
|
||
|
|
- When timing is hopeless, the allocation and byte counts are the result.
|
||
|
|
- A profile says where the time goes, and it is only read from an idle
|
||
|
|
machine. Record one with `-cpuprofile` on `BenchmarkServer`, then `go tool
|
||
|
|
pprof -top` it.
|
||
|
|
- Profile-guided optimisation: a profile is committed at
|
||
|
|
`cmd/volumen/default.pgo`, and the toolchain consumes it automatically when
|
||
|
|
it builds the command (measured: a `-pgo=off` build and a default build of
|
||
|
|
the same tree produce different binaries). A benchmark's test binary never
|
||
|
|
sees it, because the profile belongs to the main package, so the early A/B
|
||
|
|
run of `BenchmarkServer` compared two plain builds and could not have shown
|
||
|
|
a difference; its numbers stand as the plain build's, and they understate
|
||
|
|
the shipped binary. The comparison measured on 2026-09-25 against the two
|
||
|
|
real builds (five hundred posts, the six JSON and feed endpoints
|
||
|
|
interleaved, rate limiting off, both arms served the same twelve thousand
|
||
|
|
six hundred requests) puts the profile's effect beyond the noise: server
|
||
|
|
CPU over the warm mix is about six percent lower, the per-request median
|
||
|
|
about thirty-two percent lower on `site`, twenty-three percent on the post
|
||
|
|
list and fifteen percent on a single post, while the tag feed, the sitemap
|
||
|
|
and the cold store scan are unchanged within the noise. The profile
|
||
|
|
predates the search-relevance and custom-fields changes to the hot path.
|
||
|
|
|
||
|
|
## Running
|
||
|
|
|
||
|
|
```sh
|
||
|
|
just bench
|
||
|
|
```
|
||
|
|
|
||
|
|
The recipe runs the whole module with five counts:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go test -run '^$' -bench=. -benchmem -count=5 ./...
|
||
|
|
```
|
||
|
|
|
||
|
|
`-benchmem` is not optional: allocations per operation are part of the
|
||
|
|
result. A first look at one target, before the full battery is worth the
|
||
|
|
time:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go test -run '^$' -bench 'BenchmarkServer' -benchmem -benchtime=1x -count=1 ./internal/app
|
||
|
|
```
|
||
|
|
|
||
|
|
The full battery runs once, deliberately, on an idle machine. A benchmark
|
||
|
|
command is capped at about two minutes per round; longer sweeps are split.
|
||
|
|
A server benchmark also spends time in the kernel, so read the allocation
|
||
|
|
numbers alongside the time: a change that halves allocations and leaves the
|
||
|
|
time flat has moved the cost to the filesystem.
|
||
|
|
|
||
|
|
## Reports
|
||
|
|
|
||
|
|
The repository stores no benchmark reports. A performance claim in
|
||
|
|
`CHANGELOG.md` is measured with the method above on the change that makes
|
||
|
|
it, on the named machine, and the number travels with the claim.
|