Files

77 lines
4.3 KiB
Markdown

# Benchmark report: v1.0.0 against head 045ef28
One live comparison, regenerated by `just bench-report` and never
accumulated: the newest release tag against the working tree, taken in
one interleaved session on an idle machine. Re-run it before a release
and commit the file the run writes.
## Environment
- CPU: AMD RYZEN AI MAX+ PRO 395 w/ Radeon 8060S, 32 logical cores
- Memory: 118 GiB
- OS: Linux 7.2.5-200.fc44.x86_64
- Go: go version go1.27.1 linux/amd64
- Build: the portable build, no pinned `GOAMD64` level, no `GOEXPERIMENT`
## Revisions
- Release: `v1.0.0`
- Head: `045ef28`
- The set: the benchmarks `bench_set` names that both revisions carry;
a name either side lacks is left out, never counted as a result.
## Method
Four interleaved rounds, the side order alternating between rounds, two
counts per side per round, `-benchtime=0.5s` with `-benchmem`,
medians over the eight samples a side collects. A factor at or above
1.05 is faster, at or below 0.95 slower, anything between noise. Time
of the release divides time of the head, so above one means the head is
faster.
## Results
| Benchmark | v1.0.0 | head 045ef28 | factor | B/op release → head | allocs/op release → head |
|---|---|---|---|---|---|
| `KernelTransposeTiled256` | 87.58 µs | 72.89 µs | 1.20x | 524752 → 524752 | 4 → 4 |
| `Prod2D` | 49.68 µs | 42.41 µs | 1.17x | 5416 → 5416 | 37 → 37 |
| `MatMulSmall64` | 43.30 µs | 41.00 µs | 1.06x | 35040 → 35040 | 69 → 69 |
| `Add1M` | 366.14 µs | 349.53 µs | 1.05x | 8005866 → 8005863 | 69 → 69 |
| `CumSum2D` | 232.91 µs | 222.16 µs | 1.05x | 2098495 → 2098498 | 37 → 37 |
| `FFT4096` | 26.88 µs | 25.52 µs | 1.05x | 65899 → 65899 | 3 → 3 |
| `MatVec` | 29.53 µs | 28.21 µs | 1.05x | 5416 → 5416 | 37 → 37 |
| `IntegrateHeat2D` | 936.53 µs | 909.54 µs | 1.03x | 92464 → 92464 | 159 → 159 |
| `MinimiseLBFGSBounded` | 9.02 µs | 8.76 µs | 1.03x | 19016 → 19016 | 53 → 53 |
| `MatMul1024` | 26.00 ms | 25.46 ms | 1.02x | 8392200 → 8391652 | 70 → 69 |
| `MatMulOdd130` | 162.07 µs | 158.28 µs | 1.02x | 141205 → 141205 | 57 → 57 |
| `MeanAxis2D` | 25.85 µs | 25.44 µs | 1.02x | 5464 → 5464 | 37 → 37 |
| `Median` | 2.64 ms | 2.60 ms | 1.02x | 401414 → 401410 | 1 → 1 |
| `Dot` | 84.49 µs | 83.68 µs | 1.01x | 1171 → 1174 | 36 → 36 |
| `EinsumBatchedMatMul` | 314.97 µs | 312.57 µs | 1.01x | 528746 → 528745 | 78 → 78 |
| `FFT2_512x512` | 959.31 µs | 948.74 µs | 1.01x | 4461272 → 4461249 | 168 → 168 |
| `IntegrateRK4` | 887.97 µs | 880.62 µs | 1.01x | 3906104 → 3906102 | 24015 → 24015 |
| `SavitzkyGolay` | 331.05 µs | 328.79 µs | 1.01x | 2099428 → 2099429 | 69 → 69 |
| `ArgSort1M` | 10.03 ms | 10.07 ms | 1.00x | 42213251 → 42213264 | 266 → 266 |
| `Conv1D` | 166.71 µs | 167.04 µs | 1.00x | 281028 → 280800 | 106 → 106 |
| `CovarianceMatrix` | 101.98 µs | 102.44 µs | 1.00x | 136176 → 136176 | 23 → 23 |
| `IntegrateBackwardEuler` | 89.86 µs | 89.96 µs | 1.00x | 352929 → 352929 | 2806 → 2806 |
| `KernelDensity` | 241.38 µs | 241.85 µs | 1.00x | 4467 → 4462 | 69 → 69 |
| `LogisticRegression` | 235.50 µs | 236.46 µs | 1.00x | 10656 → 10656 | 61 → 61 |
| `Exp100k` | 182.19 µs | 183.11 µs | 0.99x | 805029 → 805029 | 69 → 69 |
| `MatMulTall512x64x2048` | 2.09 ms | 2.10 ms | 0.99x | 8390997 → 8390963 | 69 → 69 |
| `SumAxis3DDim1` | 2.60 µs | 2.62 µs | 0.99x | 2768 → 2768 | 4 → 4 |
| `WelchPSD` | 197.78 µs | 198.78 µs | 0.99x | 28436 → 28603 | 75 → 75 |
| `Norm2D` | 54.87 µs | 55.99 µs | 0.98x | 5432 → 5432 | 37 → 37 |
| `SumAxis2D` | 26.27 µs | 26.80 µs | 0.98x | 5473 → 5473 | 37 → 37 |
| `CWTMorlet` | 768.00 µs | 793.52 µs | 0.97x | 6427142 → 6427651 | 103 → 103 |
| `LinearRegression` | 63.01 µs | 64.84 µs | 0.97x | 36152 → 36152 | 58 → 58 |
| `MinimiseDifferentialEvolution` | 164.09 µs | 170.09 µs | 0.96x | 498250 → 498250 | 3845 → 3845 |
| `SVD128` | 17.36 ms | 18.08 ms | 0.96x | 2182742 → 2181705 | 5070 → 5072 |
| `BackwardTwoLayer` | 82.50 µs | 89.10 µs | 0.93x | 104759 → 104757 | 65 → 65 |
| `Cholesky256` | 736.43 µs | 818.51 µs | 0.90x | 1583248 → 1583351 | 342 → 342 |
| `Solve256` | 1.79 ms | 2.25 ms | 0.80x | 537635 → 537639 | 10 → 10 |
## Summary
4 of 37 benchmarks sit above the noise band (faster), 3 below it (slower), 30 inside it.