diff --git a/.gitignore b/.gitignore
index 0d56e65..e4db533 100644
--- a/.gitignore
+++ b/.gitignore
@@ -1,7 +1,7 @@
.idea/
.zcode/
-# Build artifacts
+# Build artefacts
bin/
*.test
*.out
diff --git a/CHANGELOG.md b/CHANGELOG.md
index 4a49340..f1d54e2 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -155,7 +155,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
interface, or a nil or empty slice, array or map. In 1.x the option covered
only the collections.
- The `toml` tag gained the `inline` option: a struct or map field tagged
- `toml:"retry,inline"` emits as `retry = {…}` instead of a header section,
+ `toml:"retry,inline"` emits as `retry = {â¦}` instead of a header section,
whatever its size, a named embedded struct included. Forcing it on an array
of tables is an error, because the inline form would re-parse as a value
array and change the value's Go type.
@@ -239,6 +239,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- A top-level value the encoder could not normalise reported its path with
a leading dot, `interpres: .port: ...`; the message now reads
`interpres: port: ...`, the shape `EncodeError.Path` already used.
+- A token shaped like a date-time with a component out of range, such as an
+ hour of 24 or a February the 30th, fell through to the number decoder and
+ failed with the number complaint `invalid character "-" in number`; it now
+ fails as the date-time it visibly is, `invalid date-time "..."`.
+- A `time.Time` or `OffsetDateTime` whose zone offset is not a whole number
+ of minutes wrote only the minutes, silently shifting the instant by the
+ seconds dropped; the encoder now refuses such an offset, which TOML has no
+ form for, instead of corrupting the value.
+- An empty array of tables over pointer elements, `[]*T{}`, emitted as
+ `key = []` while its value form was omitted; it is omitted too now, the
+ rule TOML forces, because an empty `[[a]]` has no valid form.
+- Two lenient grammar edges are closed: a sign in a `\u` or `\U` escape,
+ which is not a hex digit, is rejected instead of evaluating, and a bare
+ carriage return right after a multi-line string's opening delimiter is
+ the bare-CR error instead of a newline trimmed silently.
### Migration from 1.x
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index 3119429..ed01415 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -115,7 +115,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | What it does |
|---|---|---|
| Test | push or pull request to `development` | format check, vet, modernisation, build, the test suite with the coverage floor, the toml-test compliance suite |
-| Race | `workflow_dispatch`, by hand | the suite under the race detector, the same race gate the local `just gates` runs |
+| Race | `workflow_dispatch`, plus a nightly schedule | the suite under the race detector, the same race gate the local `just gates` runs |
+| Fuzz | `workflow_dispatch`, plus a nightly schedule | a 30 second fuzz smoke per target over the seeds and the gathered corpus |
| Release | a `v*` tag | tag validation, format, vet, modernisation, build and the test suite with the coverage floor, then the Gitea release created from the `CHANGELOG.md` section; no race detector |
The local equivalent is `just gates`, which is the same set plus the race
diff --git a/README.md b/README.md
index 2d87477..d578358 100644
--- a/README.md
+++ b/README.md
@@ -13,16 +13,17 @@ the entire official [toml-test](https://github.com/toml-lang/toml-test) suite:
`\xHH` escapes; integers in the four radixes with `_` separators; floats with
exponents, `inf` and `nan`; booleans; the four date-time kinds, seconds
optional as of 1.1; arrays and inline tables, multi-line as of 1.1.
-- **Decoding and encoding**: `Parse` for an untyped tree, `Unmarshal` and
- `Marshal` for structs and maps, mirroring `encoding/json`.
+- **Decoding and encoding**: `Unmarshal` and `Marshal` for structs and maps,
+ mirroring `encoding/json`; `Parse` and `ParseMap` for the document with its
+ key order and the plain untyped tree.
- **Strict decoding**: `RejectUnknownFields(true)` rejects keys that
match no destination field, at every struct depth.
- **Custom types**: `Marshaler` and `Unmarshaler` let a type control its own
TOML representation in both directions, and `encoding.TextMarshaler` and
`TextUnmarshaler` are honoured by default, so `net.IP`, `time.Duration` and
user types with text methods need no configuration.
-- **Cancellation**: every entry point has a `*Context` sibling that honours a
- `context.Context`.
+- **Cancellation**: the parse, decode and marshal entries have `*Context`
+ siblings that honour a `context.Context`, checked while the work runs.
- **Ordered documents**: `Parse` gives a `*Document` that keeps the key order,
tells an inline table from a header one, and carries the comments; `ParseMap`
gives the plain `map[string]any` tree.
@@ -173,7 +174,8 @@ See [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the full workflow, and
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/API.md](docs/API.md): the API reference, decoding and encoding rules
-- [docs/CLI.md](docs/CLI.md): the interpres-decode adapter and validator
+- [docs/CLI.md](docs/CLI.md): the interpres-decode adapter and validator,
+ also shipped as the manual page `man/interpres-decode.1`
## Licence
diff --git a/docs/API.md b/docs/API.md
index 6d40003..45d7cb6 100644
--- a/docs/API.md
+++ b/docs/API.md
@@ -105,7 +105,7 @@ Encode options:
| `InlineTables(threshold int)` | `0` | write a sub-table inline when its single-line form is at most `threshold` bytes |
| `EmitFieldComments(v bool)` | off | print the `comment=` tag option of a field above its line or header |
-### `func ParseAs[T any](data []byte) (T, error)`
+### `func ParseAs[T any](data []byte, opts ...UnmarshalOption) (T, error)`
The generic shorthand for `Unmarshal` with a destination variable:
@@ -125,7 +125,8 @@ that is not a struct warms nothing.
### `func Statements(r io.Reader) iter.Seq2[Statement, error]`
Iterates the top-level statements of the document r carries, in written
-order: key/value statements, a `[table]` header as one statement carrying
+order: key/value statements, a value array or an inline table among them as
+one statement whatever it holds, a `[table]` header as one statement carrying
its `Table` node, and an `[[array of tables]]` as one statement per element
with the element's node and its `Index`. Iteration stops at the first error
and at a false yield, so a caller looking for one section reads no further.
@@ -223,11 +224,7 @@ introduced:
and no surrounding space, so `# note` is stored as `note` and a bare `#` as
`""`.
-A `Document` is not a value to marshal: `Marshal` writes values, so it refuses
-one and points at `doc.Map()`. Writing a document back, with its order and its
-comments, belongs with the editing API.
-
-### `func Unmarshal(data []byte, v any) error`
+### `func Unmarshal(data []byte, v any, opts ...UnmarshalOption) error`
Parses `data` and stores the result in the value pointed to by `v`, typically a
pointer to a struct or to `map[string]any`. Equivalent to
@@ -240,11 +237,11 @@ if err := interpres.Unmarshal(data, &cfg); err != nil {
}
```
-### `func UnmarshalContext(ctx context.Context, data []byte, v any) error`
+### `func UnmarshalContext(ctx context.Context, data []byte, v any, opts ...UnmarshalOption) error`
The cancellable variant of `Unmarshal`.
-### `func Marshal(v any) ([]byte, error)`
+### `func Marshal(v any, opts ...MarshalOption) ([]byte, error)`
Encodes a `struct` or `map[string]V` value, or a non-nil pointer to one, into a
TOML document. The emission rules are in the [Encoding](#encoding) section
@@ -254,12 +251,12 @@ below. Equivalent to `MarshalContext(context.Background(), v)`.
out, err := interpres.Marshal(cfg)
```
-### `func MarshalContext(ctx context.Context, v any) ([]byte, error)`
+### `func MarshalContext(ctx context.Context, v any, opts ...MarshalOption) ([]byte, error)`
The cancellable variant of `Marshal`. The context is checked before any work
and every 64 fields during the reflection walk.
-### `func MarshalAppend(buf []byte, v any) ([]byte, error)`
+### `func MarshalAppend(buf []byte, v any, opts ...MarshalOption) ([]byte, error)`
Appends the TOML encoding of `v` to `buf` and returns the extended buffer, the
shape `json.MarshalAppend` has. A failed encoding leaves `buf` untouched.
@@ -525,9 +522,9 @@ through the ordinary assignment rules, so every conversion, hook and error the
path is pinned by a differential fuzz target that decodes every generated
document both ways and compares the results.
-A document or destination the direct skeleton cannot model — an unknown table
+A document or destination the direct skeleton cannot model (an unknown table
under strictness it must sink, a hook that needs the whole parsed value, an
-embedded map filler — falls back to the tree path and reruns, so the
+embedded map filler) falls back to the tree path and reruns, so the
observable behaviour is always the tree path's, exactly. Nothing changes for
`Parse`, `ParseMap` or the document API: the tree remains theirs.
@@ -561,8 +558,10 @@ sequenceDiagram
### Input constraints
`Marshal` and `MarshalWrite` accept a `struct`, a `map[string]V`, or a
-non-nil pointer to one, where `V` is any value `Marshal` itself understands. A
-different top-level value fails:
+non-nil pointer to one, where `V` is any value `Marshal` itself understands.
+An `OrderedMap` and a `Document` are accepted as themselves: the first in its
+written key order, the second written back as it stands. A different
+top-level value fails:
| Input | Error |
|---|---|
@@ -736,7 +735,7 @@ across a round-trip.
A nil slice is always omitted. An empty (length 0) array of tables is always
omitted, because TOML forbids an empty `[[a]]`. Other empty arrays emit as
-`key = []` by default; `OmitEmptyArrays()` skips them as well, so
+`key = []` by default; `OmitEmptyArrays(true)` skips them as well, so
`[]string{}` is treated like a nil slice.
### Long strings
@@ -863,7 +862,7 @@ encoding/json/v2 made current, with the differences TOML asks for:
| `encoding.TextMarshaler`, `TextUnmarshaler` | honoured, the same | a type that renders itself as text becomes a TOML string, both ways |
| `*json.UnmarshalTypeError` | `*DecodeError` | the path is segments with a `String()` renderer, not a dotted string |
| `*json.SyntaxError` | `*SyntaxError` | the TOML error adds the byte `Offset` and the `Column` to the line |
-| context support | `*Context` variants of every entry point | encoding/json has none |
+| context support | `*Context` variants of the parse, decode and marshal entries | encoding/json has none |
## Types
diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md
index 5f84e04..694727d 100644
--- a/docs/ARCHITECTURE.md
+++ b/docs/ARCHITECTURE.md
@@ -5,16 +5,17 @@ source tree; nothing is aspirational.
## Overview
-interpres is one public library package, one command, and one example. The
+interpres is one public library package, one command, and two examples. The
library implements the whole of TOML 1.1, decoding and encoding, in the
standard library alone; the command wraps the parser and the encoder for the
toml-test compliance harness, against which it stands at 214 valid, 467 invalid
-and 214 encoder cases with zero failures; the example demonstrates the API.
+and 214 encoder cases with zero failures; the examples demonstrate the API: one the document round trip, one the statement iterator.
```mermaid
flowchart TD
CLI[cmd/interpres-decode
toml-test adapter] --> API
EX[examples/basic
usage demo] --> API
+ EX2[examples/statements
statement iterator demo] --> API
subgraph Lib [package interpres]
API[interpres.go
public API and types]
API --> P[parser.go
recursive-descent parser]
@@ -35,9 +36,10 @@ strict validation.
| Path | Responsibility |
|---|---|
-| `.` (package `interpres`) | The whole library. `interpres.go` declares the exported surface (`Parse`, `Unmarshal`, `Marshal`, the `*Context` variants, the option constructors, `Marshaler`, `Unmarshaler`, `SyntaxError`, the local date-time types); everything below it is unexported. |
-| `cmd/interpres-decode` | The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with `-encode` it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic. |
+| `.` (package `interpres`) | The whole library. `interpres.go` declares the exported surface (`Parse`, `Unmarshal`, `Marshal`, the `*Context` variants, the option constructors, `Marshaler`, `Unmarshaler`, `SyntaxError`, the error and option types); everything below it is unexported. |
+| `cmd/interpres-decode` | The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with `--encode` it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic. |
| `examples/basic` | A runnable tour of the API. Documentation in executable form, not part of the library. |
+| `examples/statements` | The `Statements` iterator over a document, the shape a configuration tool reads. Documentation in executable form. |
Inside the library package, one file owns one concern:
@@ -46,18 +48,28 @@ Inside the library package, one file owns one concern:
| `parser.go` | The recursive-descent parser. Produces the `map[string]any` tree, records the nodes a [Document](API.md#documents) is built from, and enforces the structural rules of TOML 1.1 (table redefinitions, dotted keys, arrays of tables, multi-line inline tables). Reports a 1-based line on failure. |
| `document.go` | The parsed-document types: `Document`, `Table` and `Entry`, which carry the key order, whether a table was written inline, and the comments. The values they expose are the parser's own tree, not a copy. |
| `number.go` | Strict numeric tokens: integers in the four radixes with `_` separators, and floats including `inf` and `nan`. Rejects leading zeros, misplaced underscores and malformed fractions. |
-| `datetime.go` | The three local date-time wrapper types and `parseDateTime`, which classifies a token into the four date-time kinds under the strict TOML grammar. |
+| `datetime.go` | The four date-time types (`OffsetDateTime` and the three local wrappers) and `parseDateTime`, which classifies a token into the four date-time kinds under the strict TOML grammar. |
+| `orderedmap.go` | `OrderedMap`, the table that keeps its key order, and the node index the decoder reads the written order from. |
+| `target.go` | The targeted parse: the struct skeleton resolved against the document while it scans, no intermediate tree. Falls back to the tree path for every shape it does not model. |
+| `docwrite.go` | The write side of the document pipeline: `UnmarshalDocument` and the writer that renders a `Document` back with its order and comments. |
| `decode.go` | Maps the parsed tree onto Go values by reflection: struct fields, maps, slices, scalar conversion with overflow checks, `Unmarshaler` dispatch. |
| `encode.go` | The reverse walk: builds an intermediate `tomlDoc` per table (which is what preserves declaration order and enables the group-by-kind partition) and then emits it as TOML. |
The boundary that matters: `parser.go` produces only untyped trees
-(`map[string]any`, `[]any`, `[]map[string]any`, scalars); `decode.go` and
-`encode.go` are the only files that touch `reflect`; the command never touches
-either, it consumes `Parse` alone.
+(`map[string]any`, `[]any`, `[]map[string]any`, scalars); the reflection work
+lives in `decode.go`, `encode.go`, `target.go` and `orderedmap.go`; the command
+consumes `ParseMap`, `Parse` and `Marshal`, and owns no parsing or emission
+logic of its own.
## Data flow
-Decoding is parse, then one reflection walk. `SyntaxError` values are produced
+Decoding has two paths. The direct one parses straight into a struct
+destination: `target.go` resolves the table skeleton against the struct
+schema while the document scans, and values assign through the ordinary
+decoder rules, so no intermediate tree exists; that is the hot path every
+`Unmarshal` into a struct takes. A document or destination the direct
+skeleton cannot model falls back to the tree path: parse the whole document,
+then one reflection walk over the tree. `SyntaxError` values are produced
inside `parser.go` and returned as-is; conversion errors are produced inside
`decode.go` and wrapped with the key path as they unwind.
@@ -66,10 +78,13 @@ sequenceDiagram
participant Caller
participant API as interpres.go
participant P as parser.go
+ participant T as target.go
participant D as decode.go
Caller->>API: Unmarshal(data, v)
- API->>P: ParseContext(ctx, data)
- P->>P: number and datetime atoms
+ API->>T: targeted parse into the struct
+ T->>P: scanner, grammar, atoms
+ T-->>API: result, error or fallback
+ API->>P: on fallback, ParseContext(ctx, data)
P-->>API: map tree or *SyntaxError
API->>D: decode(tree, reflect value)
D-->>API: nil or wrapped field error
diff --git a/docs/BENCHMARKING.md b/docs/BENCHMARKING.md
index 00831f4..10cacfa 100644
--- a/docs/BENCHMARKING.md
+++ b/docs/BENCHMARKING.md
@@ -15,6 +15,8 @@ The benchmarks live in `bench_test.go`, next to the code they measure:
| `BenchmarkParseLong` | `ParseMap` over a generated document with about 2000 array-of-tables entries |
| `BenchmarkStrictDecodeLong` | `Unmarshal` into a typed document under `RejectUnknownFields`, over the same long document |
| `BenchmarkMarshalLong` | `Marshal` of the tree `ParseMap` produced from the long document |
+| `BenchmarkStrictDecodeTree` | the tree-path reference decode of the representative document: parse, then the reflection walk |
+| `BenchmarkStrictDecodeTreeLong` | the tree-path reference decode of the long document, the A/B baseline of the targeted parse |
## Running
diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md
index 2179b68..a8d0132 100644
--- a/docs/DEVELOPMENT.md
+++ b/docs/DEVELOPMENT.md
@@ -46,6 +46,9 @@ prints the same list.
| `just install` | builds, then copies the binary into `~/.local/bin` (`BINDIR` overrides) |
| `just uninstall` | removes the installed binary |
| `just clean` | removes `bin/` and `coverage.out` |
+| `just cross` | cross-compile smoke of the library and the command for arm64, loong64, riscv64 and the browser and edge runtimes; nightly convenience, not a gate |
+| `just release-check X.Y.Z` | the release pre-flight: the branch, a clean tree, a sync with origin, the gates, and a CHANGELOG section ready to release |
+| `just docs-drift` | compares the toml-test counts the documentation quotes with a live suite run |
## Running a single test
@@ -98,7 +101,8 @@ pipeline.
| Workflow | Trigger | What it does |
|---|---|---|
| `test.yml` | push or pull request to `development` | format check, vet, modernisation, build, the test suite with the 80 percent coverage floor, then the toml-test compliance suite |
-| `race.yml` | `workflow_dispatch`, by hand | the suite under the race detector; the same race gate `just gates` runs locally |
+| `race.yml` | `workflow_dispatch`, plus a nightly schedule at 03:00 UTC | the suite under the race detector, the same race gate `just gates` runs locally; the nightly schedule runs on the default branch, the released line |
+| `fuzz.yml` | `workflow_dispatch`, plus a nightly schedule at 03:30 UTC | a 30 second fuzz smoke per target over the seeds and the gathered corpus |
| `release.yml` | a `v*` tag | tag validation, then format, vet, modernisation, build and the test suite with the coverage floor, then the Gitea release from the CHANGELOG section. No race detector: race never runs on a push path, and the local `just gates` raced the tree before the tag was cut |
## Releases