83 lines
3.3 KiB
Markdown
83 lines
3.3 KiB
Markdown
# Benchmarking
|
|
|
|
How Tensor's performance is measured. The numbers a reader quotes must
|
|
be reproducible by following this document; anything else is an
|
|
impression, not a result.
|
|
|
|
## The tool
|
|
|
|
The benchmarks live beside the code they measure as Go benchmark
|
|
functions (`func BenchmarkXxx(b *testing.B)`), four hundred and
|
|
seventy-one of them across the packages: the array kernels and
|
|
parallel scheduler in `internal/core` and `internal/engine`, the dense
|
|
and sparse solvers in `linalg`, the transforms and filters in
|
|
`signal`, the differentiable core in `grad`, the integrators in
|
|
`integrate`, the optimisers in `optim`, the statistics in `stats`, the
|
|
collectives in `spmd` and the reader and writer round trips in `io`.
|
|
Run the whole suite with:
|
|
|
|
```sh
|
|
just bench
|
|
```
|
|
|
|
which runs `go test -run '^$' -bench=. -benchmem -count=5` over every
|
|
logic package. One package at a time:
|
|
|
|
```sh
|
|
go test ./linalg/ -bench 'BenchmarkSolve' -benchmem -count=5 -run xxx
|
|
```
|
|
|
|
`-benchmem` is not optional: allocations per operation are part of the
|
|
result. A kernel whose allocations grow has regressed even when its
|
|
time did not.
|
|
|
|
## The discipline
|
|
|
|
- **One process, A or B.** Two runs of two different binaries differ
|
|
by more than the effect being measured. When comparing a change
|
|
inside one revision, run both variants inside one process, or
|
|
interleave the sub-benchmarks behind a package-level switch.
|
|
- **Across revisions, interleave the rounds.** A release against the
|
|
head tree is necessarily two binaries; `just bench-report` runs the
|
|
two sides in alternating order over four rounds, so a host that
|
|
penalises the first run of a pair penalises both sides equally.
|
|
- **Idle machine.** A loaded machine profiles and times whatever ran
|
|
last. Close everything; treat any run sharing the box with other
|
|
work as void.
|
|
- **Five counts, median.** `just bench` takes five counts; report the
|
|
median and the spread. One to two percent is noise.
|
|
- **Deterministic inputs.** Every benchmark builds its inputs from the
|
|
seeded generator or fixed literals, so a number is tied to a
|
|
revision, not to a dice roll.
|
|
- **Complexity, not folklore.** A claim that a kernel is O(n log n)
|
|
belongs next to the measurements that show the scaling (two or
|
|
three sizes), not as an adjective.
|
|
|
|
## What is exact, what is fast
|
|
|
|
Performance numbers say nothing about correctness. The correctness
|
|
dossier lives in the test suite: `TestOracle` pins the raw-bit digest
|
|
of one fixed workload per domain (arch-specific, see
|
|
`oracle_test.go`), and every solver test carries a residual or an
|
|
exact-reference check. A benchmark result without the gates green is
|
|
not a result.
|
|
|
|
## Reports
|
|
|
|
One live report exists: `docs/benchmarks/release-vs-head.md`, the
|
|
newest release tag against the working tree. Regenerate it with:
|
|
|
|
```sh
|
|
just bench-report
|
|
```
|
|
|
|
The recipe checks out the latest `v*` tag in a scratch worktree, runs
|
|
the representative set `bench_set` names over four interleaved rounds
|
|
on both revisions, and rewrites the file. The set holds one benchmark
|
|
per kernel family; a name either revision lacks is left out rather
|
|
than counted. Run it on an idle machine and commit the file it writes
|
|
alongside the release it describes. Per-change exploration numbers
|
|
belong in the commit's own review, not in a growing pile of report
|
|
files; the repository carries the one comparison that matters, the
|
|
release a reader has against the tree as it stands.
|