feat: initial release
Release / gates (push) Successful in 4m38s
Test / test (push) Successful in 5m16s
Release / release (push) Successful in 35s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-03 10:00:00 +02:00
commit af4ee19703
617 changed files with 191195 additions and 0 deletions
+3583
View File
File diff suppressed because it is too large Load Diff
+207
View File
@@ -0,0 +1,207 @@
# Architecture
How Tensor is put together. Every node, package and arrow below
exists in the source tree; nothing is aspirational.
## Overview
```mermaid
flowchart TD
X["examples, thirteen main programs"]
F["tensor, the root facade"]
G["grad"]
I["integrate"]
L["linalg"]
S["signal"]
ST["stats"]
O["optim"]
IO["io"]
C["internal/core"]
B["internal/base"]
E["internal/engine"]
X --> F
F --> G
F --> I
F --> L
F --> S
F --> ST
F --> O
F --> IO
G --> S
G --> I
I --> L
I --> O
O --> L
G --> C
I --> C
L --> C
S --> C
ST --> C
O --> C
IO --> C
C --> B
C --> E
B --> E
```
Tensor is a scientific computing library in pure Go: no cgo, no GPU
stack, no third-party dependencies. The module splits into one core
package and one package per domain, with a strict dependency
direction: domains depend on the core, never the other way round, and
nothing below the root imports the root.
The root package is a facade, and its mechanism is a type alias plus a
forward: `type Array = core.Array` makes the root's array and the
core's one the same type rather than a wrapper, and
`facade_generated.go` declares the rest as `var SVD = linalg.SVD` and
its neighbours, so `tensor.SVD` and `linalg.SVD` are one function
value. A new domain export needs a line in that file. Nothing else in
the root package holds logic.
Every arrow above is one import. Two groups are elided for
readability: each domain also imports `internal/base` beside
`internal/core`, and the three domains that drive the parallel
fan-out themselves (`linalg`, `signal`, `grad`) also import
`internal/engine` directly.
## Packages
| Package | Responsibility |
|---|---|
| `tensor` (root) | the facade: re-exports every domain symbol through `facade_generated.go` and owns no logic beyond the alias definitions |
| `internal/core` | the `Array` type and everything that treats it as an n-dimensional value: constructors, element-wise arithmetic, reductions, shape moves, indexing and views, sorting, Einsum, interpolation, sparse COO, special functions, quasirandom sequences, the reproducible generator; deliberately no domain knowledge |
| `internal/base` | the shared primitives: the generic LU (`Factor`, `SolveSystem`), error construction with the `tensor: ` prefix, shape formatting, machine epsilon, so no domain imports another for plumbing; deliberately no array knowledge |
| `internal/engine` | the parallel scheduler and the pooled scratch buffers every kernel fans out through; the only place that starts workers |
| `linalg` | dense and sparse linear algebra: factorisations, eigenproblems, matrix functions, regularised and truncated solves, iterative sparse solvers and eigensolvers, polynomial fitting and cubic splines |
| `signal` | transforms and stencils: Fourier, cosine and sine transforms, the NUFFT, spectral estimation, filter design, wavelets, convolutions, pooling, and the spectral Poisson solves |
| `integrate` | differential equations and quadrature: ODE steppers with events, boundary-value shooting, Gauss rules, cubature, turnkey heat and wave evolution, finite elements |
| `stats` | distributions and inference: CDFs, quantiles, draws, descriptives, tests, regression models, multivariate normals, kernel density |
| `optim` | fitting and root finding: local, bounded, constrained and global minimisation |
| `plot` | deterministic SVG line charts of computed series: linear axes, legends, byte-identical figures |
| `spmd` | explicit SPMD worlds over TCP or in process: rank-0 routed links, the movement collectives and the sharded reductions whose answers are the single-array fold's exact bits at any world size; imported explicitly, the root facade does not re-export it |
| `io` | data formats: CSV, FITS images and tables, HDF5 datasets, the NetCDF classic model, memory-mapped arrays |
| `grad` | the reverse-mode differentiable core over the shared surface, second-order tools, Newton-CG, Hamiltonian Monte Carlo and adjoint sensitivities |
| `examples/*` | thirteen `main` packages, one workflow each, compiled by `just build`; they hold no library code and no tests |
The boundaries are as deliberate as the responsibilities. Five
domain-to-domain edges exist, each one-directional and each earning
its keep: `grad` reads `signal` for the spectral autograd nodes,
`grad` reads `integrate` for the adjoint ODE (imported under the name
`ode`), `integrate` reads `optim` for the shooting solver and the
implicit midpoint stage, `integrate` reads `linalg` for the
tridiagonal solve the Crank-Nicolson heat evolution rides and for the
sparse Cholesky and its orderings behind the FEM Poisson solver, and
`optim` reads `linalg` for `ArrayFromFloatsSafe`, the copying
constructor its callbacks are handed. The linear algebra `optim`
itself stands on is the generic LU in `internal/base`, not a domain
package. Beyond those edges, a domain never imports another domain,
the core never imports a domain, and a new cross-domain edge needs a
reason of the same kind. `spmd` is a domain in the dependency sense, it
reads `internal/core` and `internal/base` and nothing else of the
library, and it is deliberately absent from the root facade: a
distributed program imports it explicitly, the one place the
distributed surface is named. The canonical reduction partition and
the block folds it rides on live in `internal/core` beside the folds
themselves, so the single-array and the sharded reduction are one
computation by construction, not by a test.
The test-level harnesses sit at the root rather than in a package:
`oracle_test.go` pins a raw-bit digest per domain and platform,
`leak_test.go` measures the heap across repeated blocks, and
`example_test.go` carries the runnable godoc examples. Each domain
package carries its own `example_test.go` beside its tests, so the
documented call sequences are compiled and executed by `go test`.
## Data flow
```mermaid
flowchart TD
A1["constructors<br/>FromFloats, Zeros, Grid"]
A2["io loaders<br/>CSV, FITS, HDF5, NetCDF, mmap"]
C1["element-wise ops"]
C2["reductions"]
C3["domain kernels<br/>linalg, signal, integrate,<br/>stats, optim"]
C4["autograd graph<br/>grad.Backward"]
S1["spmd collectives<br/>Broadcast, Scatter, Gather,<br/>AllReduceShards"]
A1 --> C1
A1 --> C2
A1 --> C3
A1 --> C4
A1 --> S1
A2 --> C3
C3 --> C4
C1 --> C2
S1 --> C2
```
One kernel call, start to finish, as the layers see it:
```mermaid
sequenceDiagram
participant Caller
participant Facade as root facade
participant Domain as domain kernel
participant Core as internal/core
participant Eng as internal/engine
Caller->>Facade: tensor.SVD(a)
Facade->>Domain: linalg.SVD(a)
Domain->>Core: read payloads, allocate the output
Domain->>Eng: split the worker ranges
Eng-->>Domain: disjoint chunks, fixed order
Domain-->>Facade: fresh Array, never the input
Facade-->>Caller: the factorisation or an error
```
Arrays are dense and contiguous by construction: element i of an
array is payload index i. `Slice` preserves the invariant by a
rebased-pointer view where the selection keeps the trailing elements
contiguous (a slice along the leading axis, or one covering the whole
extent) and by a materialised copy everywhere else, and a strided
source is materialised before either path runs, because both assume
`payload[i]` is element i. A kernel that meets a non-contiguous input
through another route receives it materialised at the boundary, so
the audit has one rule: every kernel reads payloads assuming density.
That is what lets domain kernels read `RawFloats()` directly with no
per-element dispatch.
Views are read-only: nothing in the library writes through an array
it did not allocate, and optimiser updates route through
materialised parameters, so a view can never alias a buffer a later
step rewrites.
Errors are produced at the layer that detects them, prefixed
`tensor: ` by the shared error constructor in `internal/base`, and
returned unwrapped to the caller; no layer logs another layer's
error, swallows one, or turns one into a silently wrong number.
## State and lifetime
- **Long-lived.** The engine's worker pool, sized once by `SetNumCPU`
(the machine's core count by default) and re-pinnable at any time,
and the cached constant tables (Fourier twiddles, quadrature nodes)
held at package level with fixed contents.
- **Per-call.** Every kernel's output arrays and its index scratch;
nothing survives the call except the pool below.
- **Pooled.** Float64 scratch buffers flow through a typed pool whose
borrow path zeroes the window, closing the stale-buffer bug class
at the source; retention is capped, so one large table cannot pin
memory across the machine.
- **Concurrency.** Arrays are immutable and safe for concurrent use;
the `Generator` is not safe for concurrent use and is meant to be
owned by one goroutine. Parallel kernels keep their reduction order
fixed, so parallel results are bit-identical to serial ones, and
the determinism oracle at the root holds that contract by digest.
## Dependencies
There are no third-party dependencies: `go.mod` requires the standard
library alone, which is a property of the project, not an accident.
The two internal support packages exist to keep it that way and to
keep the dependency arrow one-directional: `internal/base` holds the
generic LU and the shared error and formatting helpers so no domain
imports another for plumbing, and `internal/engine` is the only place
that schedules workers, so every kernel's parallelism is decided in
one file. The five domain-to-domain imports named above are the whole
graph beyond that; a new one needs the same kind of reason.
+82
View File
@@ -0,0 +1,82 @@
# Benchmarking
How Tensor's performance is measured. The numbers a reader quotes must
be reproducible by following this document; anything else is an
impression, not a result.
## The tool
The benchmarks live beside the code they measure as Go benchmark
functions (`func BenchmarkXxx(b *testing.B)`), four hundred and
seventy-one of them across the packages: the array kernels and
parallel scheduler in `internal/core` and `internal/engine`, the dense
and sparse solvers in `linalg`, the transforms and filters in
`signal`, the differentiable core in `grad`, the integrators in
`integrate`, the optimisers in `optim`, the statistics in `stats`, the
collectives in `spmd` and the reader and writer round trips in `io`.
Run the whole suite with:
```sh
just bench
```
which runs `go test -run '^$' -bench=. -benchmem -count=5` over every
logic package. One package at a time:
```sh
go test ./linalg/ -bench 'BenchmarkSolve' -benchmem -count=5 -run xxx
```
`-benchmem` is not optional: allocations per operation are part of the
result. A kernel whose allocations grow has regressed even when its
time did not.
## The discipline
- **One process, A or B.** Two runs of two different binaries differ
by more than the effect being measured. When comparing a change
inside one revision, run both variants inside one process, or
interleave the sub-benchmarks behind a package-level switch.
- **Across revisions, interleave the rounds.** A release against the
head tree is necessarily two binaries; `just bench-report` runs the
two sides in alternating order over four rounds, so a host that
penalises the first run of a pair penalises both sides equally.
- **Idle machine.** A loaded machine profiles and times whatever ran
last. Close everything; treat any run sharing the box with other
work as void.
- **Five counts, median.** `just bench` takes five counts; report the
median and the spread. One to two percent is noise.
- **Deterministic inputs.** Every benchmark builds its inputs from the
seeded generator or fixed literals, so a number is tied to a
revision, not to a dice roll.
- **Complexity, not folklore.** A claim that a kernel is O(n log n)
belongs next to the measurements that show the scaling (two or
three sizes), not as an adjective.
## What is exact, what is fast
Performance numbers say nothing about correctness. The correctness
dossier lives in the test suite: `TestOracle` pins the raw-bit digest
of one fixed workload per domain (arch-specific, see
`oracle_test.go`), and every solver test carries a residual or an
exact-reference check. A benchmark result without the gates green is
not a result.
## Reports
One live report exists: `docs/benchmarks/release-vs-head.md`, the
newest release tag against the working tree. Regenerate it with:
```sh
just bench-report
```
The recipe checks out the latest `v*` tag in a scratch worktree, runs
the representative set `bench_set` names over four interleaved rounds
on both revisions, and rewrites the file. The set holds one benchmark
per kernel family; a name either revision lacks is left out rather
than counted. Run it on an idle machine and commit the file it writes
alongside the release it describes. Per-change exploration numbers
belong in the commit's own review, not in a growing pile of report
files; the repository carries the one comparison that matters, the
release a reader has against the tree as it stands.
+204
View File
@@ -0,0 +1,204 @@
# Development
How to work on Tensor.
## Prerequisites
- Go 1.27.1, the exact version `go.mod` declares and the newest
stable release at the time of writing. Verify the installed version
against the release list rather than memory: `go version`.
- [just](https://github.com/casey/just) for the recipes. Three of
them (`test`, `fmt-check`, `fuzz-all`) are Perl scripts.
- Perl, for those recipes and for the CI steps that carry logic. Only
the interpreter's own builtins are used, so no module installation
is needed.
- A C compiler (`gcc`) for the race detector, which `just race` and
`just gates` run; `-race` requires cgo.
Nothing else: Tensor has zero third-party dependencies.
## Setup
```sh
git clone https://sourcedock.dev/petrbalvin/tensor.git
cd tensor
just build
just test
```
## Recipes
Every recipe in the project's file, and what it does. Taken from the
file itself, so the names and the list match it exactly.
| Recipe | What it does |
|---|---|
| `just` | lists the recipes |
| `just build` | compiles everything, the example programs included |
| `just test` | the test gate: the full suite with no cache, the coverage profile and the 80 percent floor |
| `just race` | the same suite under the race detector; the expensive one, run once per task by `just gates` |
| `just unit ./internal/core/ TestName` | a fast scoped run for iterating: cached, no race, no coverage |
| `just fuzz FuzzName ./io 60s` | a time-boxed fuzz of one target in one package; never a gate |
| `just bench` | benchmarks with `-benchmem`, five counts; on an idle machine only |
| `just fmt` | formats all Go sources in place with `gofmt` |
| `just fmt-check` | verifies that `gofmt` produces no diff; prints nothing on success |
| `just vet` | `go vet` and `go fix -diff` |
| `just gates` | the definition of done in one command: `build`, `fmt-check`, `vet`, `test`, `race`, in that order |
| `just clean` | removes the build artefacts (`bin/`, `coverage.out`) |
| `just docs-check` | runs every Go program in `README.md` from a temporary module, so the documentation cannot claim what the code no longer does |
| `just fuzz-all 5s` | fuzzes every target for the budget each; exploration, never a gate |
`docs-check` and `fuzz-all` are the project extensions; none of
them is a gate.
The `packages` value behind `test`, `race`, `unit` and `bench` names
the logic packages and leaves `examples/` out: those are main
programs with no tests, and the build is what compiles them. Tensor
is a library, so the binary recipes (`install`, `run`, `dev`) have no
referent here and are absent from the file.
The scripted recipes keep their logic in Perl rather than in the
shell, which is the repository rule for every non-product script: the
shell starts commands, and anything with a branch or a loop is Perl
using the interpreter's own builtins.
`gofmt` is the single formatting authority: there is no configuration
beyond it, `just fmt-check` is the gate and `just fmt` the fix.
## Running a single test
```sh
just unit ./internal/core/ TestQuo
```
`unit` is the scoped, cached run for iterating; the second argument
is a regular expression matched against test names. Combine with
`-v` for the sub-test names, or call `go test` directly:
```sh
go test -run TestQuo -v -count=1 ./internal/core/
```
`-count=1` defeats the test cache when a result looks stale.
The runnable documentation is part of the suite, so it is exercised
the same way. Each package carries its examples beside its tests:
```sh
go test ./linalg/ -run Example -count=1 -v
```
A godoc example that stops compiling, or whose printed output drifts
from its `// Output:` comment, fails the suite rather than the
reader. The programs in `README.md` are checked the same way, though
outside the suite, because they are whole `main` programs:
```sh
just docs-check
```
which extracts every `go` block into a temporary module against the
working tree, runs it, and reports the block that failed.
## Coverage
```sh
just test
go tool cover -func=coverage.out
```
The `total:` line is the number that matters, and `just test` fails
below 80 percent. The sweep names the logic packages, so every
library package is measured while the examples stay out of the
denominator. For an HTML report:
```sh
go tool cover -html=coverage.out -o coverage.html
```
Two harnesses inside the suite guard properties that coverage
percentages do not describe, and both live at the root:
- **`TestOracle`** pins a raw-bit digest of one fixed workload per
domain, per platform and per build. A digest that moves is either a
deliberate arithmetic change or a regression, and the difference is
decided by the person who moved it, not by the test.
- **`TestNoResourceLeaks`** measures the heap across three blocks of
ten rounds and fails on a net rise above 256 KiB, which is how a
buffer that stops being released is caught before it becomes an
outage.
## Benchmarks
```sh
just bench
```
One package at a time, with a fixed budget:
```sh
go test ./internal/core/ -bench 'BenchmarkMatMul$' -benchtime 2s -run xxx
```
Benchmark on an idle machine, compare only runs made in one process
against each other, and treat a few percent as noise. The packages
carry 147 benchmarks, and the weight sits where the time is: 84 in
`internal/core`, 20 in `signal`, 13 in `stats`, 11 in `integrate`, 8
in `optim`, 7 in `linalg`, 3 in `grad` and 1 in `internal/engine`.
The binding measurement method, the report template and the measured
reports live in [docs/BENCHMARKING.md](BENCHMARKING.md) and
[docs/benchmarks/](benchmarks/).
## Debugging the build
```sh
go build -gcflags='-m' ./internal/core/ # inlining decisions
go build -gcflags='-S' ./internal/core/ # what the compiler generated
```
There is exactly one build, and it is the product:
| Build | Command | Assumes |
|---|---|---|
| portable | `go build ./...` | the toolchain default code generation, no pinned `GOAMD64` |
The portable build pins no `GOAMD64` level: the compiler has no
auto-vectoriser, so a pinned higher level would buy only scalar FMA
contraction, which the bit-pinned kernels suppress by spelling anyway
(`float64(a*b) + c`). A build pinned to a level the oracle has no
digest block for skips loudly, so a quiet mismatch cannot happen.
## Continuous integration
Gitea Actions workflows live in `.gitea/workflows/`, are written by
hand, and enforce the same gate set as `just gates`, with scripted
steps in Perl and parallelism bounded to the shared runner box:
- **`test.yml`**, on every push and pull request to `development`:
build, format check, vet, the full suite with the coverage floor,
and the oracle digests for the platform. Race is absent on purpose:
the shared box cannot afford it on every push. The one-iteration
benchmark smoke that once rode along is retired outright: the
minimum degree battery's 3-D mesh scan alone runs for minutes on one
core and allocates terabytes cumulatively, so no form of it fits the
shared box, and benchmarking is deliberate work on a developer
machine.
- **`race.yml`**, dispatched by hand: the suite under the race
detector, with the oracle digests across `fedora`, `alpine` and
`openeuler`, which is the glibc against musl check the
floating-point kernels need.
- **`release.yml`**, on a `v*` tag: the gate set minus race once at
the tag, then the Gitea release created from the matching
`CHANGELOG.md` section.
A green `just gates` locally is the fastest way to a green pipeline.
## Releases
Releases are cut by merging `development` into `main` and tagging
`vX.Y.Z`. The tag pipeline runs the gates at the tag and publishes
the release with the CHANGELOG section as its notes: the pipeline
reads the section that begins at `## [X.Y.Z]` and stops at the next
`## [`, and refuses a tag whose section is missing or empty. Nothing
is injected into the build; the toolchain records the tag because the
build simply happens there. Before cutting a tag, run `just gates`
locally: the local gate is the one that races the tree.