# Benchmark report: v1.0.0 against head 045ef28 One live comparison, regenerated by `just bench-report` and never accumulated: the newest release tag against the working tree, taken in one interleaved session on an idle machine. Re-run it before a release and commit the file the run writes. ## Environment - CPU: AMD RYZEN AI MAX+ PRO 395 w/ Radeon 8060S, 32 logical cores - Memory: 118 GiB - OS: Linux 7.2.5-200.fc44.x86_64 - Go: go version go1.27.1 linux/amd64 - Build: the portable build, no pinned `GOAMD64` level, no `GOEXPERIMENT` ## Revisions - Release: `v1.0.0` - Head: `045ef28` - The set: the benchmarks `bench_set` names that both revisions carry; a name either side lacks is left out, never counted as a result. ## Method Four interleaved rounds, the side order alternating between rounds, two counts per side per round, `-benchtime=0.5s` with `-benchmem`, medians over the eight samples a side collects. A factor at or above 1.05 is faster, at or below 0.95 slower, anything between noise. Time of the release divides time of the head, so above one means the head is faster. ## Results | Benchmark | v1.0.0 | head 045ef28 | factor | B/op release → head | allocs/op release → head | |---|---|---|---|---|---| | `KernelTransposeTiled256` | 87.58 µs | 72.89 µs | 1.20x | 524752 → 524752 | 4 → 4 | | `Prod2D` | 49.68 µs | 42.41 µs | 1.17x | 5416 → 5416 | 37 → 37 | | `MatMulSmall64` | 43.30 µs | 41.00 µs | 1.06x | 35040 → 35040 | 69 → 69 | | `Add1M` | 366.14 µs | 349.53 µs | 1.05x | 8005866 → 8005863 | 69 → 69 | | `CumSum2D` | 232.91 µs | 222.16 µs | 1.05x | 2098495 → 2098498 | 37 → 37 | | `FFT4096` | 26.88 µs | 25.52 µs | 1.05x | 65899 → 65899 | 3 → 3 | | `MatVec` | 29.53 µs | 28.21 µs | 1.05x | 5416 → 5416 | 37 → 37 | | `IntegrateHeat2D` | 936.53 µs | 909.54 µs | 1.03x | 92464 → 92464 | 159 → 159 | | `MinimiseLBFGSBounded` | 9.02 µs | 8.76 µs | 1.03x | 19016 → 19016 | 53 → 53 | | `MatMul1024` | 26.00 ms | 25.46 ms | 1.02x | 8392200 → 8391652 | 70 → 69 | | `MatMulOdd130` | 162.07 µs | 158.28 µs | 1.02x | 141205 → 141205 | 57 → 57 | | `MeanAxis2D` | 25.85 µs | 25.44 µs | 1.02x | 5464 → 5464 | 37 → 37 | | `Median` | 2.64 ms | 2.60 ms | 1.02x | 401414 → 401410 | 1 → 1 | | `Dot` | 84.49 µs | 83.68 µs | 1.01x | 1171 → 1174 | 36 → 36 | | `EinsumBatchedMatMul` | 314.97 µs | 312.57 µs | 1.01x | 528746 → 528745 | 78 → 78 | | `FFT2_512x512` | 959.31 µs | 948.74 µs | 1.01x | 4461272 → 4461249 | 168 → 168 | | `IntegrateRK4` | 887.97 µs | 880.62 µs | 1.01x | 3906104 → 3906102 | 24015 → 24015 | | `SavitzkyGolay` | 331.05 µs | 328.79 µs | 1.01x | 2099428 → 2099429 | 69 → 69 | | `ArgSort1M` | 10.03 ms | 10.07 ms | 1.00x | 42213251 → 42213264 | 266 → 266 | | `Conv1D` | 166.71 µs | 167.04 µs | 1.00x | 281028 → 280800 | 106 → 106 | | `CovarianceMatrix` | 101.98 µs | 102.44 µs | 1.00x | 136176 → 136176 | 23 → 23 | | `IntegrateBackwardEuler` | 89.86 µs | 89.96 µs | 1.00x | 352929 → 352929 | 2806 → 2806 | | `KernelDensity` | 241.38 µs | 241.85 µs | 1.00x | 4467 → 4462 | 69 → 69 | | `LogisticRegression` | 235.50 µs | 236.46 µs | 1.00x | 10656 → 10656 | 61 → 61 | | `Exp100k` | 182.19 µs | 183.11 µs | 0.99x | 805029 → 805029 | 69 → 69 | | `MatMulTall512x64x2048` | 2.09 ms | 2.10 ms | 0.99x | 8390997 → 8390963 | 69 → 69 | | `SumAxis3DDim1` | 2.60 µs | 2.62 µs | 0.99x | 2768 → 2768 | 4 → 4 | | `WelchPSD` | 197.78 µs | 198.78 µs | 0.99x | 28436 → 28603 | 75 → 75 | | `Norm2D` | 54.87 µs | 55.99 µs | 0.98x | 5432 → 5432 | 37 → 37 | | `SumAxis2D` | 26.27 µs | 26.80 µs | 0.98x | 5473 → 5473 | 37 → 37 | | `CWTMorlet` | 768.00 µs | 793.52 µs | 0.97x | 6427142 → 6427651 | 103 → 103 | | `LinearRegression` | 63.01 µs | 64.84 µs | 0.97x | 36152 → 36152 | 58 → 58 | | `MinimiseDifferentialEvolution` | 164.09 µs | 170.09 µs | 0.96x | 498250 → 498250 | 3845 → 3845 | | `SVD128` | 17.36 ms | 18.08 ms | 0.96x | 2182742 → 2181705 | 5070 → 5072 | | `BackwardTwoLayer` | 82.50 µs | 89.10 µs | 0.93x | 104759 → 104757 | 65 → 65 | | `Cholesky256` | 736.43 µs | 818.51 µs | 0.90x | 1583248 → 1583351 | 342 → 342 | | `Solve256` | 1.79 ms | 2.25 ms | 0.80x | 537635 → 537639 | 10 → 10 | ## Summary 4 of 37 benchmarks sit above the noise band (faster), 3 below it (slower), 30 inside it.