Files
volumen/docs/BENCHMARKING.md
petrbalvin f8ed33df83
Test / test (push) Successful in 7m5s
Release / gates (push) Successful in 7m28s
Release / build (amd64, freebsd) (push) Successful in 2m52s
Release / build (amd64, linux) (push) Successful in 2m46s
Release / build (arm64, freebsd) (push) Successful in 2m22s
Release / build (arm64, linux) (push) Successful in 2m38s
Release / build (loong64, linux) (push) Successful in 2m7s
Release / build (riscv64, linux) (push) Successful in 2m17s
Release / release (push) Successful in 1m0s
Initial commit
Assisted-by: GLM 5.3
2026-09-29 10:03:32 +02:00

4.5 KiB

Benchmarking

How Volumen is measured. Every number a document, a README or a changelog quotes comes from here and nowhere else.

The method

The benchmarks live next to the code they measure. The tree carries exactly two, and nothing else is measured:

Benchmark What it measures
BenchmarkRender the Markdown pipeline, from source to sanitised HTML and a table of contents, for a medium body and a large one
BenchmarkServer the wired server over real HTTP: a post list, a single post, the tag cloud, a tag feed and the sitemap, against a corpus of five hundred posts

BenchmarkRender in internal/markdown/markdown_test.go has two sub-benchmarks and calls SetBytes, so a result reads as input bytes per second: the medium input is 2400 bytes of Markdown, the large one eighteen copies of it, 43200 bytes.

BenchmarkServer in internal/app/bench_test.go builds a server over a temporary content directory and drives it through an httptest server, so it covers the whole chain rather than one function: routing, the middleware, the session layer, the store, the payload builders and the renderer. It is the workload a release is judged on, and the one a profile is recorded from.

  • The machine is the development workstation, idle: AMD Ryzen AI MAX+ PRO 395 with Radeon 8060S, 16 cores and 32 threads, 117 GiB of memory, Fedora Linux 44 (linux/amd64). A loaded box times whatever else is running, and the fastest sample can land on the wrong function.
  • The toolchain is go1.27.1 linux/amd64. Benchmarks run through go test, which builds the package's test binary; the command build's -trimpath and -buildvcs=true are not part of a benchmark's build.
  • Comparisons run inside one process. A loaded machine and separate processes of identical binaries differ by more than the effects being measured, so A/B runs alternate the two sides rather than run one after the other, and the counts are compared through their medians, with the allocations and the bytes per operation alongside the times. Differences within a few percent of the spread between runs are noise; only a difference beyond that is a result.
  • When timing is hopeless, the allocation and byte counts are the result.
  • A profile says where the time goes, and it is only read from an idle machine. Record one with -cpuprofile on BenchmarkServer, then go tool pprof -top it.
  • Profile-guided optimisation: a profile is committed at cmd/volumen/default.pgo, and the toolchain consumes it automatically when it builds the command (measured: a -pgo=off build and a default build of the same tree produce different binaries). A benchmark's test binary never sees it, because the profile belongs to the main package, so the early A/B run of BenchmarkServer compared two plain builds and could not have shown a difference; its numbers stand as the plain build's, and they understate the shipped binary. The comparison measured on 2026-09-25 against the two real builds (five hundred posts, the six JSON and feed endpoints interleaved, rate limiting off, both arms served the same twelve thousand six hundred requests) puts the profile's effect beyond the noise: server CPU over the warm mix is about six percent lower, the per-request median about thirty-two percent lower on site, twenty-three percent on the post list and fifteen percent on a single post, while the tag feed, the sitemap and the cold store scan are unchanged within the noise. The profile predates the search-relevance and custom-fields changes to the hot path.

Running

just bench

The recipe runs the whole module with five counts:

go test -run '^$' -bench=. -benchmem -count=5 ./...

-benchmem is not optional: allocations per operation are part of the result. A first look at one target, before the full battery is worth the time:

go test -run '^$' -bench 'BenchmarkServer' -benchmem -benchtime=1x -count=1 ./internal/app

The full battery runs once, deliberately, on an idle machine. A benchmark command is capped at about two minutes per round; longer sweeps are split. A server benchmark also spends time in the kernel, so read the allocation numbers alongside the time: a change that halves allocations and leaves the time flat has moved the cost to the filesystem.

Reports

The repository stores no benchmark reports. A performance claim in CHANGELOG.md is measured with the method above on the change that makes it, on the named machine, and the number travels with the claim.