208 lines
9.6 KiB
Markdown
208 lines
9.6 KiB
Markdown
# Architecture
|
|||
|
|
|
||
|
|
How Tensor is put together. Every node, package and arrow below
|
||
|
|
exists in the source tree; nothing is aspirational.
|
||
|
|
|
||
|
|
## Overview
|
||
|
|
|
||
|
|
```mermaid
|
||
|
|
flowchart TD
|
||
|
|
X["examples, thirteen main programs"]
|
||
|
|
F["tensor, the root facade"]
|
||
|
|
G["grad"]
|
||
|
|
I["integrate"]
|
||
|
|
L["linalg"]
|
||
|
|
S["signal"]
|
||
|
|
ST["stats"]
|
||
|
|
O["optim"]
|
||
|
|
IO["io"]
|
||
|
|
C["internal/core"]
|
||
|
|
B["internal/base"]
|
||
|
|
E["internal/engine"]
|
||
|
|
|
||
|
|
X --> F
|
||
|
|
F --> G
|
||
|
|
F --> I
|
||
|
|
F --> L
|
||
|
|
F --> S
|
||
|
|
F --> ST
|
||
|
|
F --> O
|
||
|
|
F --> IO
|
||
|
|
G --> S
|
||
|
|
G --> I
|
||
|
|
I --> L
|
||
|
|
I --> O
|
||
|
|
O --> L
|
||
|
|
G --> C
|
||
|
|
I --> C
|
||
|
|
L --> C
|
||
|
|
S --> C
|
||
|
|
ST --> C
|
||
|
|
O --> C
|
||
|
|
IO --> C
|
||
|
|
C --> B
|
||
|
|
C --> E
|
||
|
|
B --> E
|
||
|
|
```
|
||
|
|
|
||
|
|
Tensor is a scientific computing library in pure Go: no cgo, no GPU
|
||
|
|
stack, no third-party dependencies. The module splits into one core
|
||
|
|
package and one package per domain, with a strict dependency
|
||
|
|
direction: domains depend on the core, never the other way round, and
|
||
|
|
nothing below the root imports the root.
|
||
|
|
|
||
|
|
The root package is a facade, and its mechanism is a type alias plus a
|
||
|
|
forward: `type Array = core.Array` makes the root's array and the
|
||
|
|
core's one the same type rather than a wrapper, and
|
||
|
|
`facade_generated.go` declares the rest as `var SVD = linalg.SVD` and
|
||
|
|
its neighbours, so `tensor.SVD` and `linalg.SVD` are one function
|
||
|
|
value. A new domain export needs a line in that file. Nothing else in
|
||
|
|
the root package holds logic.
|
||
|
|
|
||
|
|
Every arrow above is one import. Two groups are elided for
|
||
|
|
readability: each domain also imports `internal/base` beside
|
||
|
|
`internal/core`, and the three domains that drive the parallel
|
||
|
|
fan-out themselves (`linalg`, `signal`, `grad`) also import
|
||
|
|
`internal/engine` directly.
|
||
|
|
|
||
|
|
## Packages
|
||
|
|
|
||
|
|
| Package | Responsibility |
|
||
|
|
|---|---|
|
||
|
|
| `tensor` (root) | the facade: re-exports every domain symbol through `facade_generated.go` and owns no logic beyond the alias definitions |
|
||
|
|
| `internal/core` | the `Array` type and everything that treats it as an n-dimensional value: constructors, element-wise arithmetic, reductions, shape moves, indexing and views, sorting, Einsum, interpolation, sparse COO, special functions, quasirandom sequences, the reproducible generator; deliberately no domain knowledge |
|
||
|
|
| `internal/base` | the shared primitives: the generic LU (`Factor`, `SolveSystem`), error construction with the `tensor: ` prefix, shape formatting, machine epsilon, so no domain imports another for plumbing; deliberately no array knowledge |
|
||
|
|
| `internal/engine` | the parallel scheduler and the pooled scratch buffers every kernel fans out through; the only place that starts workers |
|
||
|
|
| `linalg` | dense and sparse linear algebra: factorisations, eigenproblems, matrix functions, regularised and truncated solves, iterative sparse solvers and eigensolvers, polynomial fitting and cubic splines |
|
||
|
|
| `signal` | transforms and stencils: Fourier, cosine and sine transforms, the NUFFT, spectral estimation, filter design, wavelets, convolutions, pooling, and the spectral Poisson solves |
|
||
|
|
| `integrate` | differential equations and quadrature: ODE steppers with events, boundary-value shooting, Gauss rules, cubature, turnkey heat and wave evolution, finite elements |
|
||
|
|
| `stats` | distributions and inference: CDFs, quantiles, draws, descriptives, tests, regression models, multivariate normals, kernel density |
|
||
|
|
| `optim` | fitting and root finding: local, bounded, constrained and global minimisation |
|
||
|
|
| `plot` | deterministic SVG line charts of computed series: linear axes, legends, byte-identical figures |
|
||
|
|
| `spmd` | explicit SPMD worlds over TCP or in process: rank-0 routed links, the movement collectives and the sharded reductions whose answers are the single-array fold's exact bits at any world size; imported explicitly, the root facade does not re-export it |
|
||
|
|
| `io` | data formats: CSV, FITS images and tables, HDF5 datasets, the NetCDF classic model, memory-mapped arrays |
|
||
|
|
| `grad` | the reverse-mode differentiable core over the shared surface, second-order tools, Newton-CG, Hamiltonian Monte Carlo and adjoint sensitivities |
|
||
|
|
| `examples/*` | thirteen `main` packages, one workflow each, compiled by `just build`; they hold no library code and no tests |
|
||
|
|
|
||
|
|
The boundaries are as deliberate as the responsibilities. Five
|
||
|
|
domain-to-domain edges exist, each one-directional and each earning
|
||
|
|
its keep: `grad` reads `signal` for the spectral autograd nodes,
|
||
|
|
`grad` reads `integrate` for the adjoint ODE (imported under the name
|
||
|
|
`ode`), `integrate` reads `optim` for the shooting solver and the
|
||
|
|
implicit midpoint stage, `integrate` reads `linalg` for the
|
||
|
|
tridiagonal solve the Crank-Nicolson heat evolution rides and for the
|
||
|
|
sparse Cholesky and its orderings behind the FEM Poisson solver, and
|
||
|
|
`optim` reads `linalg` for `ArrayFromFloatsSafe`, the copying
|
||
|
|
constructor its callbacks are handed. The linear algebra `optim`
|
||
|
|
itself stands on is the generic LU in `internal/base`, not a domain
|
||
|
|
package. Beyond those edges, a domain never imports another domain,
|
||
|
|
the core never imports a domain, and a new cross-domain edge needs a
|
||
|
|
reason of the same kind. `spmd` is a domain in the dependency sense, it
|
||
|
|
reads `internal/core` and `internal/base` and nothing else of the
|
||
|
|
library, and it is deliberately absent from the root facade: a
|
||
|
|
distributed program imports it explicitly, the one place the
|
||
|
|
distributed surface is named. The canonical reduction partition and
|
||
|
|
the block folds it rides on live in `internal/core` beside the folds
|
||
|
|
themselves, so the single-array and the sharded reduction are one
|
||
|
|
computation by construction, not by a test.
|
||
|
|
|
||
|
|
The test-level harnesses sit at the root rather than in a package:
|
||
|
|
`oracle_test.go` pins a raw-bit digest per domain and platform,
|
||
|
|
`leak_test.go` measures the heap across repeated blocks, and
|
||
|
|
`example_test.go` carries the runnable godoc examples. Each domain
|
||
|
|
package carries its own `example_test.go` beside its tests, so the
|
||
|
|
documented call sequences are compiled and executed by `go test`.
|
||
|
|
|
||
|
|
## Data flow
|
||
|
|
|
||
|
|
```mermaid
|
||
|
|
flowchart TD
|
||
|
|
A1["constructors<br/>FromFloats, Zeros, Grid"]
|
||
|
|
A2["io loaders<br/>CSV, FITS, HDF5, NetCDF, mmap"]
|
||
|
|
C1["element-wise ops"]
|
||
|
|
C2["reductions"]
|
||
|
|
C3["domain kernels<br/>linalg, signal, integrate,<br/>stats, optim"]
|
||
|
|
C4["autograd graph<br/>grad.Backward"]
|
||
|
|
S1["spmd collectives<br/>Broadcast, Scatter, Gather,<br/>AllReduceShards"]
|
||
|
|
A1 --> C1
|
||
|
|
A1 --> C2
|
||
|
|
A1 --> C3
|
||
|
|
A1 --> C4
|
||
|
|
A1 --> S1
|
||
|
|
A2 --> C3
|
||
|
|
C3 --> C4
|
||
|
|
C1 --> C2
|
||
|
|
S1 --> C2
|
||
|
|
```
|
||
|
|
|
||
|
|
One kernel call, start to finish, as the layers see it:
|
||
|
|
|
||
|
|
```mermaid
|
||
|
|
sequenceDiagram
|
||
|
|
participant Caller
|
||
|
|
participant Facade as root facade
|
||
|
|
participant Domain as domain kernel
|
||
|
|
participant Core as internal/core
|
||
|
|
participant Eng as internal/engine
|
||
|
|
Caller->>Facade: tensor.SVD(a)
|
||
|
|
Facade->>Domain: linalg.SVD(a)
|
||
|
|
Domain->>Core: read payloads, allocate the output
|
||
|
|
Domain->>Eng: split the worker ranges
|
||
|
|
Eng-->>Domain: disjoint chunks, fixed order
|
||
|
|
Domain-->>Facade: fresh Array, never the input
|
||
|
|
Facade-->>Caller: the factorisation or an error
|
||
|
|
```
|
||
|
|
|
||
|
|
Arrays are dense and contiguous by construction: element i of an
|
||
|
|
array is payload index i. `Slice` preserves the invariant by a
|
||
|
|
rebased-pointer view where the selection keeps the trailing elements
|
||
|
|
contiguous (a slice along the leading axis, or one covering the whole
|
||
|
|
extent) and by a materialised copy everywhere else, and a strided
|
||
|
|
source is materialised before either path runs, because both assume
|
||
|
|
`payload[i]` is element i. A kernel that meets a non-contiguous input
|
||
|
|
through another route receives it materialised at the boundary, so
|
||
|
|
the audit has one rule: every kernel reads payloads assuming density.
|
||
|
|
That is what lets domain kernels read `RawFloats()` directly with no
|
||
|
|
per-element dispatch.
|
||
|
|
|
||
|
|
Views are read-only: nothing in the library writes through an array
|
||
|
|
it did not allocate, and optimiser updates route through
|
||
|
|
materialised parameters, so a view can never alias a buffer a later
|
||
|
|
step rewrites.
|
||
|
|
|
||
|
|
Errors are produced at the layer that detects them, prefixed
|
||
|
|
`tensor: ` by the shared error constructor in `internal/base`, and
|
||
|
|
returned unwrapped to the caller; no layer logs another layer's
|
||
|
|
error, swallows one, or turns one into a silently wrong number.
|
||
|
|
|
||
|
|
## State and lifetime
|
||
|
|
|
||
|
|
- **Long-lived.** The engine's worker pool, sized once by `SetNumCPU`
|
||
|
|
(the machine's core count by default) and re-pinnable at any time,
|
||
|
|
and the cached constant tables (Fourier twiddles, quadrature nodes)
|
||
|
|
held at package level with fixed contents.
|
||
|
|
- **Per-call.** Every kernel's output arrays and its index scratch;
|
||
|
|
nothing survives the call except the pool below.
|
||
|
|
- **Pooled.** Float64 scratch buffers flow through a typed pool whose
|
||
|
|
borrow path zeroes the window, closing the stale-buffer bug class
|
||
|
|
at the source; retention is capped, so one large table cannot pin
|
||
|
|
memory across the machine.
|
||
|
|
- **Concurrency.** Arrays are immutable and safe for concurrent use;
|
||
|
|
the `Generator` is not safe for concurrent use and is meant to be
|
||
|
|
owned by one goroutine. Parallel kernels keep their reduction order
|
||
|
|
fixed, so parallel results are bit-identical to serial ones, and
|
||
|
|
the determinism oracle at the root holds that contract by digest.
|
||
|
|
|
||
|
|
## Dependencies
|
||
|
|
|
||
|
|
There are no third-party dependencies: `go.mod` requires the standard
|
||
|
|
library alone, which is a property of the project, not an accident.
|
||
|
|
|
||
|
|
The two internal support packages exist to keep it that way and to
|
||
|
|
keep the dependency arrow one-directional: `internal/base` holds the
|
||
|
|
generic LU and the shared error and formatting helpers so no domain
|
||
|
|
imports another for plumbing, and `internal/engine` is the only place
|
||
|
|
that schedules workers, so every kernel's parallelism is decided in
|
||
|
|
one file. The five domain-to-domain imports named above are the whole
|
||
|
|
graph beyond that; a new one needs the same kind of reason.
|