# Architecture
How Tensor is put together. Every node, package and arrow below
exists in the source tree; nothing is aspirational.
## Overview
```mermaid
flowchart TD
X["examples, thirteen main programs"]
F["tensor, the root facade"]
G["grad"]
I["integrate"]
L["linalg"]
S["signal"]
ST["stats"]
O["optim"]
IO["io"]
C["internal/core"]
B["internal/base"]
E["internal/engine"]
X --> F
F --> G
F --> I
F --> L
F --> S
F --> ST
F --> O
F --> IO
G --> S
G --> I
I --> L
I --> O
O --> L
G --> C
I --> C
L --> C
S --> C
ST --> C
O --> C
IO --> C
C --> B
C --> E
B --> E
```
Tensor is a scientific computing library in pure Go: no cgo, no GPU
stack, no third-party dependencies. The module splits into one core
package and one package per domain, with a strict dependency
direction: domains depend on the core, never the other way round, and
nothing below the root imports the root.
The root package is a facade, and its mechanism is a type alias plus a
forward: `type Array = core.Array` makes the root's array and the
core's one the same type rather than a wrapper, and
`facade_generated.go` declares the rest as `var SVD = linalg.SVD` and
its neighbours, so `tensor.SVD` and `linalg.SVD` are one function
value. A new domain export needs a line in that file. Nothing else in
the root package holds logic.
Every arrow above is one import. Two groups are elided for
readability: each domain also imports `internal/base` beside
`internal/core`, and the three domains that drive the parallel
fan-out themselves (`linalg`, `signal`, `grad`) also import
`internal/engine` directly.
## Packages
| Package | Responsibility |
|---|---|
| `tensor` (root) | the facade: re-exports every domain symbol through `facade_generated.go` and owns no logic beyond the alias definitions |
| `internal/core` | the `Array` type and everything that treats it as an n-dimensional value: constructors, element-wise arithmetic, reductions, shape moves, indexing and views, sorting, Einsum, interpolation, sparse COO, special functions, quasirandom sequences, the reproducible generator; deliberately no domain knowledge |
| `internal/base` | the shared primitives: the generic LU (`Factor`, `SolveSystem`), error construction with the `tensor: ` prefix, shape formatting, machine epsilon, so no domain imports another for plumbing; deliberately no array knowledge |
| `internal/engine` | the parallel scheduler and the pooled scratch buffers every kernel fans out through; the only place that starts workers |
| `linalg` | dense and sparse linear algebra: factorisations, eigenproblems, matrix functions, regularised and truncated solves, iterative sparse solvers and eigensolvers, polynomial fitting and cubic splines |
| `signal` | transforms and stencils: Fourier, cosine and sine transforms, the NUFFT, spectral estimation, filter design, wavelets, convolutions, pooling, and the spectral Poisson solves |
| `integrate` | differential equations and quadrature: ODE steppers with events, boundary-value shooting, Gauss rules, cubature, turnkey heat and wave evolution, finite elements |
| `stats` | distributions and inference: CDFs, quantiles, draws, descriptives, tests, regression models, multivariate normals, kernel density |
| `optim` | fitting and root finding: local, bounded, constrained and global minimisation |
| `plot` | deterministic SVG line charts of computed series: linear axes, legends, byte-identical figures |
| `spmd` | explicit SPMD worlds over TCP or in process: rank-0 routed links, the movement collectives and the sharded reductions whose answers are the single-array fold's exact bits at any world size; imported explicitly, the root facade does not re-export it |
| `io` | data formats: CSV, FITS images and tables, HDF5 datasets, the NetCDF classic model, memory-mapped arrays |
| `grad` | the reverse-mode differentiable core over the shared surface, second-order tools, Newton-CG, Hamiltonian Monte Carlo and adjoint sensitivities |
| `examples/*` | thirteen `main` packages, one workflow each, compiled by `just build`; they hold no library code and no tests |
The boundaries are as deliberate as the responsibilities. Five
domain-to-domain edges exist, each one-directional and each earning
its keep: `grad` reads `signal` for the spectral autograd nodes,
`grad` reads `integrate` for the adjoint ODE (imported under the name
`ode`), `integrate` reads `optim` for the shooting solver and the
implicit midpoint stage, `integrate` reads `linalg` for the
tridiagonal solve the Crank-Nicolson heat evolution rides and for the
sparse Cholesky and its orderings behind the FEM Poisson solver, and
`optim` reads `linalg` for `ArrayFromFloatsSafe`, the copying
constructor its callbacks are handed. The linear algebra `optim`
itself stands on is the generic LU in `internal/base`, not a domain
package. Beyond those edges, a domain never imports another domain,
the core never imports a domain, and a new cross-domain edge needs a
reason of the same kind. `spmd` is a domain in the dependency sense, it
reads `internal/core` and `internal/base` and nothing else of the
library, and it is deliberately absent from the root facade: a
distributed program imports it explicitly, the one place the
distributed surface is named. The canonical reduction partition and
the block folds it rides on live in `internal/core` beside the folds
themselves, so the single-array and the sharded reduction are one
computation by construction, not by a test.
The test-level harnesses sit at the root rather than in a package:
`oracle_test.go` pins a raw-bit digest per domain and platform,
`leak_test.go` measures the heap across repeated blocks, and
`example_test.go` carries the runnable godoc examples. Each domain
package carries its own `example_test.go` beside its tests, so the
documented call sequences are compiled and executed by `go test`.
## Data flow
```mermaid
flowchart TD
A1["constructors
FromFloats, Zeros, Grid"]
A2["io loaders
CSV, FITS, HDF5, NetCDF, mmap"]
C1["element-wise ops"]
C2["reductions"]
C3["domain kernels
linalg, signal, integrate,
stats, optim"]
C4["autograd graph
grad.Backward"]
S1["spmd collectives
Broadcast, Scatter, Gather,
AllReduceShards"]
A1 --> C1
A1 --> C2
A1 --> C3
A1 --> C4
A1 --> S1
A2 --> C3
C3 --> C4
C1 --> C2
S1 --> C2
```
One kernel call, start to finish, as the layers see it:
```mermaid
sequenceDiagram
participant Caller
participant Facade as root facade
participant Domain as domain kernel
participant Core as internal/core
participant Eng as internal/engine
Caller->>Facade: tensor.SVD(a)
Facade->>Domain: linalg.SVD(a)
Domain->>Core: read payloads, allocate the output
Domain->>Eng: split the worker ranges
Eng-->>Domain: disjoint chunks, fixed order
Domain-->>Facade: fresh Array, never the input
Facade-->>Caller: the factorisation or an error
```
Arrays are dense and contiguous by construction: element i of an
array is payload index i. `Slice` preserves the invariant by a
rebased-pointer view where the selection keeps the trailing elements
contiguous (a slice along the leading axis, or one covering the whole
extent) and by a materialised copy everywhere else, and a strided
source is materialised before either path runs, because both assume
`payload[i]` is element i. A kernel that meets a non-contiguous input
through another route receives it materialised at the boundary, so
the audit has one rule: every kernel reads payloads assuming density.
That is what lets domain kernels read `RawFloats()` directly with no
per-element dispatch.
Views are read-only: nothing in the library writes through an array
it did not allocate, and optimiser updates route through
materialised parameters, so a view can never alias a buffer a later
step rewrites.
Errors are produced at the layer that detects them, prefixed
`tensor: ` by the shared error constructor in `internal/base`, and
returned unwrapped to the caller; no layer logs another layer's
error, swallows one, or turns one into a silently wrong number.
## State and lifetime
- **Long-lived.** The engine's worker pool, sized once by `SetNumCPU`
(the machine's core count by default) and re-pinnable at any time,
and the cached constant tables (Fourier twiddles, quadrature nodes)
held at package level with fixed contents.
- **Per-call.** Every kernel's output arrays and its index scratch;
nothing survives the call except the pool below.
- **Pooled.** Float64 scratch buffers flow through a typed pool whose
borrow path zeroes the window, closing the stale-buffer bug class
at the source; retention is capped, so one large table cannot pin
memory across the machine.
- **Concurrency.** Arrays are immutable and safe for concurrent use;
the `Generator` is not safe for concurrent use and is meant to be
owned by one goroutine. Parallel kernels keep their reduction order
fixed, so parallel results are bit-identical to serial ones, and
the determinism oracle at the root holds that contract by digest.
## Dependencies
There are no third-party dependencies: `go.mod` requires the standard
library alone, which is a property of the project, not an accident.
The two internal support packages exist to keep it that way and to
keep the dependency arrow one-directional: `internal/base` holds the
generic LU and the shared error and formatting helpers so no domain
imports another for plumbing, and `internal/engine` is the only place
that schedules workers, so every kernel's parallelism is decided in
one file. The five domain-to-domain imports named above are the whole
graph beyond that; a new one needs the same kind of reason.