Files
interpres/docs/ARCHITECTURE.md
2026-09-22 21:15:07 +02:00

7.8 KiB

Architecture

How interpres is put together. Every file, package and arrow below exists in the source tree; nothing is aspirational.

Overview

interpres is one public library package, one command, and two examples. The library implements the whole of TOML 1.1, decoding and encoding, in the standard library alone; the command wraps the parser and the encoder for the toml-test compliance harness, against which it stands at 214 valid, 467 invalid and 214 encoder cases with zero failures; the examples demonstrate the API: one the document round trip, one the statement iterator.

flowchart TD
    CLI[cmd/interpres-decode<br/>toml-test adapter] --> API
    EX[examples/basic<br/>usage demo] --> API
    EX2[examples/statements<br/>statement iterator demo] --> API
    subgraph Lib [package interpres]
        API[interpres.go<br/>public API and types]
        API --> P[parser.go<br/>recursive-descent parser]
        API --> DEC[decode.go<br/>tree onto Go values]
        API --> ENC[encode.go<br/>Go values to TOML]
        P --> N[number.go<br/>numeric tokens]
        P --> DT[datetime.go<br/>date-time atoms and types]
    end

The public API in interpres.go is a thin facade: every entry point funnels into ParseContext for parsing and into the unexported decoder and encoder for the reflection work. parser.go owns the grammar; it leans on number.go and datetime.go for the two token families that need their own strict validation.

Packages

Path Responsibility
. (package interpres) The whole library. interpres.go declares the exported surface (Parse, Unmarshal, Marshal, the *Context variants, the option constructors, Marshaler, Unmarshaler, SyntaxError, the error and option types); everything below it is unexported.
cmd/interpres-decode The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with --encode it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic.
examples/basic A runnable tour of the API. Documentation in executable form, not part of the library.
examples/statements The Statements iterator over a document, the shape a configuration tool reads. Documentation in executable form.

Inside the library package, one file owns one concern:

File Responsibility
parser.go The recursive-descent parser. Produces the map[string]any tree, records the nodes a Document is built from, and enforces the structural rules of TOML 1.1 (table redefinitions, dotted keys, arrays of tables, multi-line inline tables). Reports a 1-based line on failure.
document.go The parsed-document types: Document, Table and Entry, which carry the key order, whether a table was written inline, and the comments. The values they expose are the parser's own tree, not a copy.
number.go Strict numeric tokens: integers in the four radixes with _ separators, and floats including inf and nan. Rejects leading zeros, misplaced underscores and malformed fractions.
datetime.go The four date-time types (OffsetDateTime and the three local wrappers) and parseDateTime, which classifies a token into the four date-time kinds under the strict TOML grammar.
orderedmap.go OrderedMap, the table that keeps its key order, and the node index the decoder reads the written order from.
target.go The targeted parse: the struct skeleton resolved against the document while it scans, no intermediate tree. Falls back to the tree path for every shape it does not model.
docwrite.go The write side of the document pipeline: UnmarshalDocument and the writer that renders a Document back with its order and comments.
decode.go Maps the parsed tree onto Go values by reflection: struct fields, maps, slices, scalar conversion with overflow checks, Unmarshaler dispatch.
encode.go The reverse walk: builds an intermediate tomlDoc per table (which is what preserves declaration order and enables the group-by-kind partition) and then emits it as TOML.

The boundary that matters: parser.go produces only untyped trees (map[string]any, []any, []map[string]any, scalars); the reflection work lives in decode.go, encode.go, target.go and orderedmap.go; the command consumes ParseMap, Parse and Marshal, and owns no parsing or emission logic of its own.

Data flow

Decoding has two paths. The direct one parses straight into a struct destination: target.go resolves the table skeleton against the struct schema while the document scans, and values assign through the ordinary decoder rules, so no intermediate tree exists; that is the hot path every Unmarshal into a struct takes. A document or destination the direct skeleton cannot model falls back to the tree path: parse the whole document, then one reflection walk over the tree. SyntaxError values are produced inside parser.go and returned as-is; conversion errors are produced inside decode.go and wrapped with the key path as they unwind.

sequenceDiagram
    participant Caller
    participant API as interpres.go
    participant P as parser.go
    participant T as target.go
    participant D as decode.go
    Caller->>API: Unmarshal(data, v)
    API->>T: targeted parse into the struct
    T->>P: scanner, grammar, atoms
    T-->>API: result, error or fallback
    API->>P: on fallback, ParseContext(ctx, data)
    P-->>API: map tree or *SyntaxError
    API->>D: decode(tree, reflect value)
    D-->>API: nil or wrapped field error
    API-->>Caller: error

Encoding walks the other way. encode.go first builds a tomlDoc from the value, then emits it; the two phases are why Layout can reorder entries without a second reflection pass, and why cancellation is checked during both.

sequenceDiagram
    participant Caller
    participant API as interpres.go
    participant B as encode.go build
    participant E as encode.go emit
    Caller->>API: Marshal(v)
    API->>B: build tomlDoc from struct or map
    B-->>API: tomlDoc or error
    API->>E: emitDoc(doc)
    E-->>API: bytes or error
    API-->>Caller: bytes, error

State and lifetime

  • The option values are stateless: every Unmarshal, Marshal and their variants apply their own options into a per-call unexported worker, so the entries are safe for concurrent use.
  • The parser is allocated per ParseContext call; the parser itself caches nothing between documents.
  • The shared state is a set of caches and pools whose entries are immutable once published, each growing with the number of distinct types rather than with document size: the struct-schema cache in decode.go (a sync.Map keyed on reflect.Type, holding the flattened field layout the decoder and the encoder both consult), the per-type interface flag caches in decode.go and encode.go (recording where Marshaler, Unmarshaler and the text interfaces can be found, so a walk builds an interface value only where the assertion can succeed), each fronted by a monomorphic hint holding the type resolved last, and the encoder's output-buffer pool in encode.go (sync.Pool, buffers returned to it only within a 1 MiB retention cap). A published schema or flag set never mutates, so concurrent callers only race to build an identical value, the same trade-off encoding/json's field cache makes.
  • The date-time wrappers are values, not pointers, and are immutable in use.
  • Nothing in the library starts goroutines; apart from the caches and the pool above, which never mutate a published entry, there is no shared mutable state.

Dependencies

None. go.mod declares the module and the Go version and carries no requires; the library imports the standard library only, which is the point of the project. The toml-test binary is a development and CI tool, never a module dependency.