# Architecture How Tensor is put together. Every node, package and arrow below exists in the source tree; nothing is aspirational. ## Overview ```mermaid flowchart TD X["examples, thirteen main programs"] F["tensor, the root facade"] G["grad"] I["integrate"] L["linalg"] S["signal"] ST["stats"] O["optim"] IO["io"] C["internal/core"] B["internal/base"] E["internal/engine"] X --> F F --> G F --> I F --> L F --> S F --> ST F --> O F --> IO G --> S G --> I I --> L I --> O O --> L G --> C I --> C L --> C S --> C ST --> C O --> C IO --> C C --> B C --> E B --> E ``` Tensor is a scientific computing library in pure Go: no cgo, no GPU stack, no third-party dependencies. The module splits into one core package and one package per domain, with a strict dependency direction: domains depend on the core, never the other way round, and nothing below the root imports the root. The root package is a facade, and its mechanism is a type alias plus a forward: `type Array = core.Array` makes the root's array and the core's one the same type rather than a wrapper, and `facade_generated.go` declares the rest as `var SVD = linalg.SVD` and its neighbours, so `tensor.SVD` and `linalg.SVD` are one function value. A new domain export needs a line in that file. Nothing else in the root package holds logic. Every arrow above is one import. Two groups are elided for readability: each domain also imports `internal/base` beside `internal/core`, and the three domains that drive the parallel fan-out themselves (`linalg`, `signal`, `grad`) also import `internal/engine` directly. ## Packages | Package | Responsibility | |---|---| | `tensor` (root) | the facade: re-exports every domain symbol through `facade_generated.go` and owns no logic beyond the alias definitions | | `internal/core` | the `Array` type and everything that treats it as an n-dimensional value: constructors, element-wise arithmetic, reductions, shape moves, indexing and views, sorting, Einsum, interpolation, sparse COO, special functions, quasirandom sequences, the reproducible generator; deliberately no domain knowledge | | `internal/base` | the shared primitives: the generic LU (`Factor`, `SolveSystem`), error construction with the `tensor: ` prefix, shape formatting, machine epsilon, so no domain imports another for plumbing; deliberately no array knowledge | | `internal/engine` | the parallel scheduler and the pooled scratch buffers every kernel fans out through; the only place that starts workers | | `linalg` | dense and sparse linear algebra: factorisations, eigenproblems, matrix functions, regularised and truncated solves, iterative sparse solvers and eigensolvers, polynomial fitting and cubic splines | | `signal` | transforms and stencils: Fourier, cosine and sine transforms, the NUFFT, spectral estimation, filter design, wavelets, convolutions, pooling, and the spectral Poisson solves | | `integrate` | differential equations and quadrature: ODE steppers with events, boundary-value shooting, Gauss rules, cubature, turnkey heat and wave evolution, finite elements | | `stats` | distributions and inference: CDFs, quantiles, draws, descriptives, tests, regression models, multivariate normals, kernel density | | `optim` | fitting and root finding: local, bounded, constrained and global minimisation | | `plot` | deterministic SVG line charts of computed series: linear axes, legends, byte-identical figures | | `spmd` | explicit SPMD worlds over TCP or in process: rank-0 routed links, the movement collectives and the sharded reductions whose answers are the single-array fold's exact bits at any world size; imported explicitly, the root facade does not re-export it | | `io` | data formats: CSV, FITS images and tables, HDF5 datasets, the NetCDF classic model, memory-mapped arrays | | `grad` | the reverse-mode differentiable core over the shared surface, second-order tools, Newton-CG, Hamiltonian Monte Carlo and adjoint sensitivities | | `examples/*` | thirteen `main` packages, one workflow each, compiled by `just build`; they hold no library code and no tests | The boundaries are as deliberate as the responsibilities. Five domain-to-domain edges exist, each one-directional and each earning its keep: `grad` reads `signal` for the spectral autograd nodes, `grad` reads `integrate` for the adjoint ODE (imported under the name `ode`), `integrate` reads `optim` for the shooting solver and the implicit midpoint stage, `integrate` reads `linalg` for the tridiagonal solve the Crank-Nicolson heat evolution rides and for the sparse Cholesky and its orderings behind the FEM Poisson solver, and `optim` reads `linalg` for `ArrayFromFloatsSafe`, the copying constructor its callbacks are handed. The linear algebra `optim` itself stands on is the generic LU in `internal/base`, not a domain package. Beyond those edges, a domain never imports another domain, the core never imports a domain, and a new cross-domain edge needs a reason of the same kind. `spmd` is a domain in the dependency sense, it reads `internal/core` and `internal/base` and nothing else of the library, and it is deliberately absent from the root facade: a distributed program imports it explicitly, the one place the distributed surface is named. The canonical reduction partition and the block folds it rides on live in `internal/core` beside the folds themselves, so the single-array and the sharded reduction are one computation by construction, not by a test. The test-level harnesses sit at the root rather than in a package: `oracle_test.go` pins a raw-bit digest per domain and platform, `leak_test.go` measures the heap across repeated blocks, and `example_test.go` carries the runnable godoc examples. Each domain package carries its own `example_test.go` beside its tests, so the documented call sequences are compiled and executed by `go test`. ## Data flow ```mermaid flowchart TD A1["constructors
FromFloats, Zeros, Grid"] A2["io loaders
CSV, FITS, HDF5, NetCDF, mmap"] C1["element-wise ops"] C2["reductions"] C3["domain kernels
linalg, signal, integrate,
stats, optim"] C4["autograd graph
grad.Backward"] S1["spmd collectives
Broadcast, Scatter, Gather,
AllReduceShards"] A1 --> C1 A1 --> C2 A1 --> C3 A1 --> C4 A1 --> S1 A2 --> C3 C3 --> C4 C1 --> C2 S1 --> C2 ``` One kernel call, start to finish, as the layers see it: ```mermaid sequenceDiagram participant Caller participant Facade as root facade participant Domain as domain kernel participant Core as internal/core participant Eng as internal/engine Caller->>Facade: tensor.SVD(a) Facade->>Domain: linalg.SVD(a) Domain->>Core: read payloads, allocate the output Domain->>Eng: split the worker ranges Eng-->>Domain: disjoint chunks, fixed order Domain-->>Facade: fresh Array, never the input Facade-->>Caller: the factorisation or an error ``` Arrays are dense and contiguous by construction: element i of an array is payload index i. `Slice` preserves the invariant by a rebased-pointer view where the selection keeps the trailing elements contiguous (a slice along the leading axis, or one covering the whole extent) and by a materialised copy everywhere else, and a strided source is materialised before either path runs, because both assume `payload[i]` is element i. A kernel that meets a non-contiguous input through another route receives it materialised at the boundary, so the audit has one rule: every kernel reads payloads assuming density. That is what lets domain kernels read `RawFloats()` directly with no per-element dispatch. Views are read-only: nothing in the library writes through an array it did not allocate, and optimiser updates route through materialised parameters, so a view can never alias a buffer a later step rewrites. Errors are produced at the layer that detects them, prefixed `tensor: ` by the shared error constructor in `internal/base`, and returned unwrapped to the caller; no layer logs another layer's error, swallows one, or turns one into a silently wrong number. ## State and lifetime - **Long-lived.** The engine's worker pool, sized once by `SetNumCPU` (the machine's core count by default) and re-pinnable at any time, and the cached constant tables (Fourier twiddles, quadrature nodes) held at package level with fixed contents. - **Per-call.** Every kernel's output arrays and its index scratch; nothing survives the call except the pool below. - **Pooled.** Float64 scratch buffers flow through a typed pool whose borrow path zeroes the window, closing the stale-buffer bug class at the source; retention is capped, so one large table cannot pin memory across the machine. - **Concurrency.** Arrays are immutable and safe for concurrent use; the `Generator` is not safe for concurrent use and is meant to be owned by one goroutine. Parallel kernels keep their reduction order fixed, so parallel results are bit-identical to serial ones, and the determinism oracle at the root holds that contract by digest. ## Dependencies There are no third-party dependencies: `go.mod` requires the standard library alone, which is a property of the project, not an accident. The two internal support packages exist to keep it that way and to keep the dependency arrow one-directional: `internal/base` holds the generic LU and the shared error and formatting helpers so no domain imports another for plumbing, and `internal/engine` is the only place that schedules workers, so every kernel's parallelism is decided in one file. The five domain-to-domain imports named above are the whole graph beyond that; a new one needs the same kind of reason.