205 lines
7.9 KiB
Markdown
205 lines
7.9 KiB
Markdown
# Development
|
|||
|
|
|
||
|
|
How to work on Tensor.
|
||
|
|
|
||
|
|
## Prerequisites
|
||
|
|
|
||
|
|
- Go 1.27.1, the exact version `go.mod` declares and the newest
|
||
|
|
stable release at the time of writing. Verify the installed version
|
||
|
|
against the release list rather than memory: `go version`.
|
||
|
|
- [just](https://github.com/casey/just) for the recipes. Three of
|
||
|
|
them (`test`, `fmt-check`, `fuzz-all`) are Perl scripts.
|
||
|
|
- Perl, for those recipes and for the CI steps that carry logic. Only
|
||
|
|
the interpreter's own builtins are used, so no module installation
|
||
|
|
is needed.
|
||
|
|
- A C compiler (`gcc`) for the race detector, which `just race` and
|
||
|
|
`just gates` run; `-race` requires cgo.
|
||
|
|
|
||
|
|
Nothing else: Tensor has zero third-party dependencies.
|
||
|
|
|
||
|
|
## Setup
|
||
|
|
|
||
|
|
```sh
|
||
|
|
git clone https://sourcedock.dev/petrbalvin/tensor.git
|
||
|
|
cd tensor
|
||
|
|
just build
|
||
|
|
just test
|
||
|
|
```
|
||
|
|
|
||
|
|
## Recipes
|
||
|
|
|
||
|
|
Every recipe in the project's file, and what it does. Taken from the
|
||
|
|
file itself, so the names and the list match it exactly.
|
||
|
|
|
||
|
|
| Recipe | What it does |
|
||
|
|
|---|---|
|
||
|
|
| `just` | lists the recipes |
|
||
|
|
| `just build` | compiles everything, the example programs included |
|
||
|
|
| `just test` | the test gate: the full suite with no cache, the coverage profile and the 80 percent floor |
|
||
|
|
| `just race` | the same suite under the race detector; the expensive one, run once per task by `just gates` |
|
||
|
|
| `just unit ./internal/core/ TestName` | a fast scoped run for iterating: cached, no race, no coverage |
|
||
|
|
| `just fuzz FuzzName ./io 60s` | a time-boxed fuzz of one target in one package; never a gate |
|
||
|
|
| `just bench` | benchmarks with `-benchmem`, five counts; on an idle machine only |
|
||
|
|
| `just fmt` | formats all Go sources in place with `gofmt` |
|
||
|
|
| `just fmt-check` | verifies that `gofmt` produces no diff; prints nothing on success |
|
||
|
|
| `just vet` | `go vet` and `go fix -diff` |
|
||
|
|
| `just gates` | the definition of done in one command: `build`, `fmt-check`, `vet`, `test`, `race`, in that order |
|
||
|
|
| `just clean` | removes the build artefacts (`bin/`, `coverage.out`) |
|
||
|
|
| `just docs-check` | runs every Go program in `README.md` from a temporary module, so the documentation cannot claim what the code no longer does |
|
||
|
|
| `just fuzz-all 5s` | fuzzes every target for the budget each; exploration, never a gate |
|
||
|
|
|
||
|
|
`docs-check` and `fuzz-all` are the project extensions; none of
|
||
|
|
them is a gate.
|
||
|
|
The `packages` value behind `test`, `race`, `unit` and `bench` names
|
||
|
|
the logic packages and leaves `examples/` out: those are main
|
||
|
|
programs with no tests, and the build is what compiles them. Tensor
|
||
|
|
is a library, so the binary recipes (`install`, `run`, `dev`) have no
|
||
|
|
referent here and are absent from the file.
|
||
|
|
|
||
|
|
The scripted recipes keep their logic in Perl rather than in the
|
||
|
|
shell, which is the repository rule for every non-product script: the
|
||
|
|
shell starts commands, and anything with a branch or a loop is Perl
|
||
|
|
using the interpreter's own builtins.
|
||
|
|
|
||
|
|
`gofmt` is the single formatting authority: there is no configuration
|
||
|
|
beyond it, `just fmt-check` is the gate and `just fmt` the fix.
|
||
|
|
|
||
|
|
## Running a single test
|
||
|
|
|
||
|
|
```sh
|
||
|
|
just unit ./internal/core/ TestQuo
|
||
|
|
```
|
||
|
|
|
||
|
|
`unit` is the scoped, cached run for iterating; the second argument
|
||
|
|
is a regular expression matched against test names. Combine with
|
||
|
|
`-v` for the sub-test names, or call `go test` directly:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go test -run TestQuo -v -count=1 ./internal/core/
|
||
|
|
```
|
||
|
|
|
||
|
|
`-count=1` defeats the test cache when a result looks stale.
|
||
|
|
|
||
|
|
The runnable documentation is part of the suite, so it is exercised
|
||
|
|
the same way. Each package carries its examples beside its tests:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go test ./linalg/ -run Example -count=1 -v
|
||
|
|
```
|
||
|
|
|
||
|
|
A godoc example that stops compiling, or whose printed output drifts
|
||
|
|
from its `// Output:` comment, fails the suite rather than the
|
||
|
|
reader. The programs in `README.md` are checked the same way, though
|
||
|
|
outside the suite, because they are whole `main` programs:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
just docs-check
|
||
|
|
```
|
||
|
|
|
||
|
|
which extracts every `go` block into a temporary module against the
|
||
|
|
working tree, runs it, and reports the block that failed.
|
||
|
|
|
||
|
|
## Coverage
|
||
|
|
|
||
|
|
```sh
|
||
|
|
just test
|
||
|
|
go tool cover -func=coverage.out
|
||
|
|
```
|
||
|
|
|
||
|
|
The `total:` line is the number that matters, and `just test` fails
|
||
|
|
below 80 percent. The sweep names the logic packages, so every
|
||
|
|
library package is measured while the examples stay out of the
|
||
|
|
denominator. For an HTML report:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go tool cover -html=coverage.out -o coverage.html
|
||
|
|
```
|
||
|
|
|
||
|
|
Two harnesses inside the suite guard properties that coverage
|
||
|
|
percentages do not describe, and both live at the root:
|
||
|
|
|
||
|
|
- **`TestOracle`** pins a raw-bit digest of one fixed workload per
|
||
|
|
domain, per platform and per build. A digest that moves is either a
|
||
|
|
deliberate arithmetic change or a regression, and the difference is
|
||
|
|
decided by the person who moved it, not by the test.
|
||
|
|
- **`TestNoResourceLeaks`** measures the heap across three blocks of
|
||
|
|
ten rounds and fails on a net rise above 256 KiB, which is how a
|
||
|
|
buffer that stops being released is caught before it becomes an
|
||
|
|
outage.
|
||
|
|
|
||
|
|
## Benchmarks
|
||
|
|
|
||
|
|
```sh
|
||
|
|
just bench
|
||
|
|
```
|
||
|
|
|
||
|
|
One package at a time, with a fixed budget:
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go test ./internal/core/ -bench 'BenchmarkMatMul$' -benchtime 2s -run xxx
|
||
|
|
```
|
||
|
|
|
||
|
|
Benchmark on an idle machine, compare only runs made in one process
|
||
|
|
against each other, and treat a few percent as noise. The packages
|
||
|
|
carry 147 benchmarks, and the weight sits where the time is: 84 in
|
||
|
|
`internal/core`, 20 in `signal`, 13 in `stats`, 11 in `integrate`, 8
|
||
|
|
in `optim`, 7 in `linalg`, 3 in `grad` and 1 in `internal/engine`.
|
||
|
|
The binding measurement method, the report template and the measured
|
||
|
|
reports live in [docs/BENCHMARKING.md](BENCHMARKING.md) and
|
||
|
|
[docs/benchmarks/](benchmarks/).
|
||
|
|
|
||
|
|
## Debugging the build
|
||
|
|
|
||
|
|
```sh
|
||
|
|
go build -gcflags='-m' ./internal/core/ # inlining decisions
|
||
|
|
go build -gcflags='-S' ./internal/core/ # what the compiler generated
|
||
|
|
```
|
||
|
|
|
||
|
|
There is exactly one build, and it is the product:
|
||
|
|
|
||
|
|
| Build | Command | Assumes |
|
||
|
|
|---|---|---|
|
||
|
|
| portable | `go build ./...` | the toolchain default code generation, no pinned `GOAMD64` |
|
||
|
|
|
||
|
|
The portable build pins no `GOAMD64` level: the compiler has no
|
||
|
|
auto-vectoriser, so a pinned higher level would buy only scalar FMA
|
||
|
|
contraction, which the bit-pinned kernels suppress by spelling anyway
|
||
|
|
(`float64(a*b) + c`). A build pinned to a level the oracle has no
|
||
|
|
digest block for skips loudly, so a quiet mismatch cannot happen.
|
||
|
|
|
||
|
|
## Continuous integration
|
||
|
|
|
||
|
|
Gitea Actions workflows live in `.gitea/workflows/`, are written by
|
||
|
|
hand, and enforce the same gate set as `just gates`, with scripted
|
||
|
|
steps in Perl and parallelism bounded to the shared runner box:
|
||
|
|
|
||
|
|
- **`test.yml`**, on every push and pull request to `development`:
|
||
|
|
build, format check, vet, the full suite with the coverage floor,
|
||
|
|
and the oracle digests for the platform. Race is absent on purpose:
|
||
|
|
the shared box cannot afford it on every push. The one-iteration
|
||
|
|
benchmark smoke that once rode along is retired outright: the
|
||
|
|
minimum degree battery's 3-D mesh scan alone runs for minutes on one
|
||
|
|
core and allocates terabytes cumulatively, so no form of it fits the
|
||
|
|
shared box, and benchmarking is deliberate work on a developer
|
||
|
|
machine.
|
||
|
|
- **`race.yml`**, dispatched by hand: the suite under the race
|
||
|
|
detector, with the oracle digests across `fedora`, `alpine` and
|
||
|
|
`openeuler`, which is the glibc against musl check the
|
||
|
|
floating-point kernels need.
|
||
|
|
- **`release.yml`**, on a `v*` tag: the gate set minus race once at
|
||
|
|
the tag, then the Gitea release created from the matching
|
||
|
|
`CHANGELOG.md` section.
|
||
|
|
|
||
|
|
A green `just gates` locally is the fastest way to a green pipeline.
|
||
|
|
|
||
|
|
## Releases
|
||
|
|
|
||
|
|
Releases are cut by merging `development` into `main` and tagging
|
||
|
|
`vX.Y.Z`. The tag pipeline runs the gates at the tag and publishes
|
||
|
|
the release with the CHANGELOG section as its notes: the pipeline
|
||
|
|
reads the section that begins at `## [X.Y.Z]` and stops at the next
|
||
|
|
`## [`, and refuses a tag whose section is missing or empty. Nothing
|
||
|
|
is injected into the build; the toolchain records the tag because the
|
||
|
|
build simply happens there. Before cutting a tag, run `just gates`
|
||
|
|
locally: the local gate is the one that races the tree.
|