3.3 KiB
Benchmarking
How Tensor's performance is measured. The numbers a reader quotes must be reproducible by following this document; anything else is an impression, not a result.
The tool
The benchmarks live beside the code they measure as Go benchmark
functions (func BenchmarkXxx(b *testing.B)), four hundred and
seventy-one of them across the packages: the array kernels and
parallel scheduler in internal/core and internal/engine, the dense
and sparse solvers in linalg, the transforms and filters in
signal, the differentiable core in grad, the integrators in
integrate, the optimisers in optim, the statistics in stats, the
collectives in spmd and the reader and writer round trips in io.
Run the whole suite with:
just bench
which runs go test -run '^$' -bench=. -benchmem -count=5 over every
logic package. One package at a time:
go test ./linalg/ -bench 'BenchmarkSolve' -benchmem -count=5 -run xxx
-benchmem is not optional: allocations per operation are part of the
result. A kernel whose allocations grow has regressed even when its
time did not.
The discipline
- One process, A or B. Two runs of two different binaries differ by more than the effect being measured. When comparing a change inside one revision, run both variants inside one process, or interleave the sub-benchmarks behind a package-level switch.
- Across revisions, interleave the rounds. A release against the
head tree is necessarily two binaries;
just bench-reportruns the two sides in alternating order over four rounds, so a host that penalises the first run of a pair penalises both sides equally. - Idle machine. A loaded machine profiles and times whatever ran last. Close everything; treat any run sharing the box with other work as void.
- Five counts, median.
just benchtakes five counts; report the median and the spread. One to two percent is noise. - Deterministic inputs. Every benchmark builds its inputs from the seeded generator or fixed literals, so a number is tied to a revision, not to a dice roll.
- Complexity, not folklore. A claim that a kernel is O(n log n) belongs next to the measurements that show the scaling (two or three sizes), not as an adjective.
What is exact, what is fast
Performance numbers say nothing about correctness. The correctness
dossier lives in the test suite: TestOracle pins the raw-bit digest
of one fixed workload per domain (arch-specific, see
oracle_test.go), and every solver test carries a residual or an
exact-reference check. A benchmark result without the gates green is
not a result.
Reports
One live report exists: docs/benchmarks/release-vs-head.md, the
newest release tag against the working tree. Regenerate it with:
just bench-report
The recipe checks out the latest v* tag in a scratch worktree, runs
the representative set bench_set names over four interleaved rounds
on both revisions, and rewrites the file. The set holds one benchmark
per kernel family; a name either revision lacks is left out rather
than counted. Run it on an idle machine and commit the file it writes
alongside the release it describes. Per-change exploration numbers
belong in the commit's own review, not in a growing pile of report
files; the repository carries the one comparison that matters, the
release a reader has against the tree as it stands.