Files
tensor/docs/ARCHITECTURE.md
T
petrbalvin af4ee19703
Release / gates (push) Successful in 4m38s
Test / test (push) Successful in 5m16s
Release / release (push) Successful in 35s
feat: initial release
Assisted-by: GLM 5.3 Flash
2026-09-03 10:00:00 +02:00

9.6 KiB

Architecture

How Tensor is put together. Every node, package and arrow below exists in the source tree; nothing is aspirational.

Overview

flowchart TD
    X["examples, thirteen main programs"]
    F["tensor, the root facade"]
    G["grad"]
    I["integrate"]
    L["linalg"]
    S["signal"]
    ST["stats"]
    O["optim"]
    IO["io"]
    C["internal/core"]
    B["internal/base"]
    E["internal/engine"]

    X --> F
    F --> G
    F --> I
    F --> L
    F --> S
    F --> ST
    F --> O
    F --> IO
    G --> S
    G --> I
    I --> L
    I --> O
    O --> L
    G --> C
    I --> C
    L --> C
    S --> C
    ST --> C
    O --> C
    IO --> C
    C --> B
    C --> E
    B --> E

Tensor is a scientific computing library in pure Go: no cgo, no GPU stack, no third-party dependencies. The module splits into one core package and one package per domain, with a strict dependency direction: domains depend on the core, never the other way round, and nothing below the root imports the root.

The root package is a facade, and its mechanism is a type alias plus a forward: type Array = core.Array makes the root's array and the core's one the same type rather than a wrapper, and facade_generated.go declares the rest as var SVD = linalg.SVD and its neighbours, so tensor.SVD and linalg.SVD are one function value. A new domain export needs a line in that file. Nothing else in the root package holds logic.

Every arrow above is one import. Two groups are elided for readability: each domain also imports internal/base beside internal/core, and the three domains that drive the parallel fan-out themselves (linalg, signal, grad) also import internal/engine directly.

Packages

Package Responsibility
tensor (root) the facade: re-exports every domain symbol through facade_generated.go and owns no logic beyond the alias definitions
internal/core the Array type and everything that treats it as an n-dimensional value: constructors, element-wise arithmetic, reductions, shape moves, indexing and views, sorting, Einsum, interpolation, sparse COO, special functions, quasirandom sequences, the reproducible generator; deliberately no domain knowledge
internal/base the shared primitives: the generic LU (Factor, SolveSystem), error construction with the tensor: prefix, shape formatting, machine epsilon, so no domain imports another for plumbing; deliberately no array knowledge
internal/engine the parallel scheduler and the pooled scratch buffers every kernel fans out through; the only place that starts workers
linalg dense and sparse linear algebra: factorisations, eigenproblems, matrix functions, regularised and truncated solves, iterative sparse solvers and eigensolvers, polynomial fitting and cubic splines
signal transforms and stencils: Fourier, cosine and sine transforms, the NUFFT, spectral estimation, filter design, wavelets, convolutions, pooling, and the spectral Poisson solves
integrate differential equations and quadrature: ODE steppers with events, boundary-value shooting, Gauss rules, cubature, turnkey heat and wave evolution, finite elements
stats distributions and inference: CDFs, quantiles, draws, descriptives, tests, regression models, multivariate normals, kernel density
optim fitting and root finding: local, bounded, constrained and global minimisation
plot deterministic SVG line charts of computed series: linear axes, legends, byte-identical figures
spmd explicit SPMD worlds over TCP or in process: rank-0 routed links, the movement collectives and the sharded reductions whose answers are the single-array fold's exact bits at any world size; imported explicitly, the root facade does not re-export it
io data formats: CSV, FITS images and tables, HDF5 datasets, the NetCDF classic model, memory-mapped arrays
grad the reverse-mode differentiable core over the shared surface, second-order tools, Newton-CG, Hamiltonian Monte Carlo and adjoint sensitivities
examples/* thirteen main packages, one workflow each, compiled by just build; they hold no library code and no tests

The boundaries are as deliberate as the responsibilities. Five domain-to-domain edges exist, each one-directional and each earning its keep: grad reads signal for the spectral autograd nodes, grad reads integrate for the adjoint ODE (imported under the name ode), integrate reads optim for the shooting solver and the implicit midpoint stage, integrate reads linalg for the tridiagonal solve the Crank-Nicolson heat evolution rides and for the sparse Cholesky and its orderings behind the FEM Poisson solver, and optim reads linalg for ArrayFromFloatsSafe, the copying constructor its callbacks are handed. The linear algebra optim itself stands on is the generic LU in internal/base, not a domain package. Beyond those edges, a domain never imports another domain, the core never imports a domain, and a new cross-domain edge needs a reason of the same kind. spmd is a domain in the dependency sense, it reads internal/core and internal/base and nothing else of the library, and it is deliberately absent from the root facade: a distributed program imports it explicitly, the one place the distributed surface is named. The canonical reduction partition and the block folds it rides on live in internal/core beside the folds themselves, so the single-array and the sharded reduction are one computation by construction, not by a test.

The test-level harnesses sit at the root rather than in a package: oracle_test.go pins a raw-bit digest per domain and platform, leak_test.go measures the heap across repeated blocks, and example_test.go carries the runnable godoc examples. Each domain package carries its own example_test.go beside its tests, so the documented call sequences are compiled and executed by go test.

Data flow

flowchart TD
    A1["constructors<br/>FromFloats, Zeros, Grid"]
    A2["io loaders<br/>CSV, FITS, HDF5, NetCDF, mmap"]
    C1["element-wise ops"]
    C2["reductions"]
    C3["domain kernels<br/>linalg, signal, integrate,<br/>stats, optim"]
    C4["autograd graph<br/>grad.Backward"]
    S1["spmd collectives<br/>Broadcast, Scatter, Gather,<br/>AllReduceShards"]
    A1 --> C1
    A1 --> C2
    A1 --> C3
    A1 --> C4
    A1 --> S1
    A2 --> C3
    C3 --> C4
    C1 --> C2
    S1 --> C2

One kernel call, start to finish, as the layers see it:

sequenceDiagram
    participant Caller
    participant Facade as root facade
    participant Domain as domain kernel
    participant Core as internal/core
    participant Eng as internal/engine
    Caller->>Facade: tensor.SVD(a)
    Facade->>Domain: linalg.SVD(a)
    Domain->>Core: read payloads, allocate the output
    Domain->>Eng: split the worker ranges
    Eng-->>Domain: disjoint chunks, fixed order
    Domain-->>Facade: fresh Array, never the input
    Facade-->>Caller: the factorisation or an error

Arrays are dense and contiguous by construction: element i of an array is payload index i. Slice preserves the invariant by a rebased-pointer view where the selection keeps the trailing elements contiguous (a slice along the leading axis, or one covering the whole extent) and by a materialised copy everywhere else, and a strided source is materialised before either path runs, because both assume payload[i] is element i. A kernel that meets a non-contiguous input through another route receives it materialised at the boundary, so the audit has one rule: every kernel reads payloads assuming density. That is what lets domain kernels read RawFloats() directly with no per-element dispatch.

Views are read-only: nothing in the library writes through an array it did not allocate, and optimiser updates route through materialised parameters, so a view can never alias a buffer a later step rewrites.

Errors are produced at the layer that detects them, prefixed tensor: by the shared error constructor in internal/base, and returned unwrapped to the caller; no layer logs another layer's error, swallows one, or turns one into a silently wrong number.

State and lifetime

  • Long-lived. The engine's worker pool, sized once by SetNumCPU (the machine's core count by default) and re-pinnable at any time, and the cached constant tables (Fourier twiddles, quadrature nodes) held at package level with fixed contents.
  • Per-call. Every kernel's output arrays and its index scratch; nothing survives the call except the pool below.
  • Pooled. Float64 scratch buffers flow through a typed pool whose borrow path zeroes the window, closing the stale-buffer bug class at the source; retention is capped, so one large table cannot pin memory across the machine.
  • Concurrency. Arrays are immutable and safe for concurrent use; the Generator is not safe for concurrent use and is meant to be owned by one goroutine. Parallel kernels keep their reduction order fixed, so parallel results are bit-identical to serial ones, and the determinism oracle at the root holds that contract by digest.

Dependencies

There are no third-party dependencies: go.mod requires the standard library alone, which is a property of the project, not an accident.

The two internal support packages exist to keep it that way and to keep the dependency arrow one-directional: internal/base holds the generic LU and the shared error and formatting helpers so no domain imports another for plumbing, and internal/engine is the only place that schedules workers, so every kernel's parallelism is decided in one file. The five domain-to-domain imports named above are the whole graph beyond that; a new one needs the same kind of reason.