From 2fa075de000128f6aebf0d77f252575708c243e4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Petr=20Balv=C3=ADn?= Date: Tue, 22 Sep 2026 21:15:07 +0200 Subject: [PATCH] docs: align every document with the reviewed behaviour Assisted-by: GLM 5.3 --- .gitignore | 2 +- CHANGELOG.md | 17 ++++++++++++++++- CONTRIBUTING.md | 3 ++- README.md | 12 +++++++----- docs/API.md | 33 ++++++++++++++++----------------- docs/ARCHITECTURE.md | 37 ++++++++++++++++++++++++++----------- docs/BENCHMARKING.md | 2 ++ docs/DEVELOPMENT.md | 6 +++++- 8 files changed, 75 insertions(+), 37 deletions(-) diff --git a/.gitignore b/.gitignore index 0d56e65..e4db533 100644 --- a/.gitignore +++ b/.gitignore @@ -1,7 +1,7 @@ .idea/ .zcode/ -# Build artifacts +# Build artefacts bin/ *.test *.out diff --git a/CHANGELOG.md b/CHANGELOG.md index 4a49340..f1d54e2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -155,7 +155,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 interface, or a nil or empty slice, array or map. In 1.x the option covered only the collections. - The `toml` tag gained the `inline` option: a struct or map field tagged - `toml:"retry,inline"` emits as `retry = {…}` instead of a header section, + `toml:"retry,inline"` emits as `retry = {…}` instead of a header section, whatever its size, a named embedded struct included. Forcing it on an array of tables is an error, because the inline form would re-parse as a value array and change the value's Go type. @@ -239,6 +239,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - A top-level value the encoder could not normalise reported its path with a leading dot, `interpres: .port: ...`; the message now reads `interpres: port: ...`, the shape `EncodeError.Path` already used. +- A token shaped like a date-time with a component out of range, such as an + hour of 24 or a February the 30th, fell through to the number decoder and + failed with the number complaint `invalid character "-" in number`; it now + fails as the date-time it visibly is, `invalid date-time "..."`. +- A `time.Time` or `OffsetDateTime` whose zone offset is not a whole number + of minutes wrote only the minutes, silently shifting the instant by the + seconds dropped; the encoder now refuses such an offset, which TOML has no + form for, instead of corrupting the value. +- An empty array of tables over pointer elements, `[]*T{}`, emitted as + `key = []` while its value form was omitted; it is omitted too now, the + rule TOML forces, because an empty `[[a]]` has no valid form. +- Two lenient grammar edges are closed: a sign in a `\u` or `\U` escape, + which is not a hex digit, is rejected instead of evaluating, and a bare + carriage return right after a multi-line string's opening delimiter is + the bare-CR error instead of a newline trimmed silently. ### Migration from 1.x diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 3119429..ed01415 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -115,7 +115,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners: | Workflow | Trigger | What it does | |---|---|---| | Test | push or pull request to `development` | format check, vet, modernisation, build, the test suite with the coverage floor, the toml-test compliance suite | -| Race | `workflow_dispatch`, by hand | the suite under the race detector, the same race gate the local `just gates` runs | +| Race | `workflow_dispatch`, plus a nightly schedule | the suite under the race detector, the same race gate the local `just gates` runs | +| Fuzz | `workflow_dispatch`, plus a nightly schedule | a 30 second fuzz smoke per target over the seeds and the gathered corpus | | Release | a `v*` tag | tag validation, format, vet, modernisation, build and the test suite with the coverage floor, then the Gitea release created from the `CHANGELOG.md` section; no race detector | The local equivalent is `just gates`, which is the same set plus the race diff --git a/README.md b/README.md index 2d87477..d578358 100644 --- a/README.md +++ b/README.md @@ -13,16 +13,17 @@ the entire official [toml-test](https://github.com/toml-lang/toml-test) suite: `\xHH` escapes; integers in the four radixes with `_` separators; floats with exponents, `inf` and `nan`; booleans; the four date-time kinds, seconds optional as of 1.1; arrays and inline tables, multi-line as of 1.1. -- **Decoding and encoding**: `Parse` for an untyped tree, `Unmarshal` and - `Marshal` for structs and maps, mirroring `encoding/json`. +- **Decoding and encoding**: `Unmarshal` and `Marshal` for structs and maps, + mirroring `encoding/json`; `Parse` and `ParseMap` for the document with its + key order and the plain untyped tree. - **Strict decoding**: `RejectUnknownFields(true)` rejects keys that match no destination field, at every struct depth. - **Custom types**: `Marshaler` and `Unmarshaler` let a type control its own TOML representation in both directions, and `encoding.TextMarshaler` and `TextUnmarshaler` are honoured by default, so `net.IP`, `time.Duration` and user types with text methods need no configuration. -- **Cancellation**: every entry point has a `*Context` sibling that honours a - `context.Context`. +- **Cancellation**: the parse, decode and marshal entries have `*Context` + siblings that honour a `context.Context`, checked while the work runs. - **Ordered documents**: `Parse` gives a `*Document` that keeps the key order, tells an inline table from a header one, and carries the comments; `ParseMap` gives the plain `map[string]any` tree. @@ -173,7 +174,8 @@ See [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the full workflow, and - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow - [docs/API.md](docs/API.md): the API reference, decoding and encoding rules -- [docs/CLI.md](docs/CLI.md): the interpres-decode adapter and validator +- [docs/CLI.md](docs/CLI.md): the interpres-decode adapter and validator, + also shipped as the manual page `man/interpres-decode.1` ## Licence diff --git a/docs/API.md b/docs/API.md index 6d40003..45d7cb6 100644 --- a/docs/API.md +++ b/docs/API.md @@ -105,7 +105,7 @@ Encode options: | `InlineTables(threshold int)` | `0` | write a sub-table inline when its single-line form is at most `threshold` bytes | | `EmitFieldComments(v bool)` | off | print the `comment=` tag option of a field above its line or header | -### `func ParseAs[T any](data []byte) (T, error)` +### `func ParseAs[T any](data []byte, opts ...UnmarshalOption) (T, error)` The generic shorthand for `Unmarshal` with a destination variable: @@ -125,7 +125,8 @@ that is not a struct warms nothing. ### `func Statements(r io.Reader) iter.Seq2[Statement, error]` Iterates the top-level statements of the document r carries, in written -order: key/value statements, a `[table]` header as one statement carrying +order: key/value statements, a value array or an inline table among them as +one statement whatever it holds, a `[table]` header as one statement carrying its `Table` node, and an `[[array of tables]]` as one statement per element with the element's node and its `Index`. Iteration stops at the first error and at a false yield, so a caller looking for one section reads no further. @@ -223,11 +224,7 @@ introduced: and no surrounding space, so `# note` is stored as `note` and a bare `#` as `""`. -A `Document` is not a value to marshal: `Marshal` writes values, so it refuses -one and points at `doc.Map()`. Writing a document back, with its order and its -comments, belongs with the editing API. - -### `func Unmarshal(data []byte, v any) error` +### `func Unmarshal(data []byte, v any, opts ...UnmarshalOption) error` Parses `data` and stores the result in the value pointed to by `v`, typically a pointer to a struct or to `map[string]any`. Equivalent to @@ -240,11 +237,11 @@ if err := interpres.Unmarshal(data, &cfg); err != nil { } ``` -### `func UnmarshalContext(ctx context.Context, data []byte, v any) error` +### `func UnmarshalContext(ctx context.Context, data []byte, v any, opts ...UnmarshalOption) error` The cancellable variant of `Unmarshal`. -### `func Marshal(v any) ([]byte, error)` +### `func Marshal(v any, opts ...MarshalOption) ([]byte, error)` Encodes a `struct` or `map[string]V` value, or a non-nil pointer to one, into a TOML document. The emission rules are in the [Encoding](#encoding) section @@ -254,12 +251,12 @@ below. Equivalent to `MarshalContext(context.Background(), v)`. out, err := interpres.Marshal(cfg) ``` -### `func MarshalContext(ctx context.Context, v any) ([]byte, error)` +### `func MarshalContext(ctx context.Context, v any, opts ...MarshalOption) ([]byte, error)` The cancellable variant of `Marshal`. The context is checked before any work and every 64 fields during the reflection walk. -### `func MarshalAppend(buf []byte, v any) ([]byte, error)` +### `func MarshalAppend(buf []byte, v any, opts ...MarshalOption) ([]byte, error)` Appends the TOML encoding of `v` to `buf` and returns the extended buffer, the shape `json.MarshalAppend` has. A failed encoding leaves `buf` untouched. @@ -525,9 +522,9 @@ through the ordinary assignment rules, so every conversion, hook and error the path is pinned by a differential fuzz target that decodes every generated document both ways and compares the results. -A document or destination the direct skeleton cannot model — an unknown table +A document or destination the direct skeleton cannot model (an unknown table under strictness it must sink, a hook that needs the whole parsed value, an -embedded map filler — falls back to the tree path and reruns, so the +embedded map filler) falls back to the tree path and reruns, so the observable behaviour is always the tree path's, exactly. Nothing changes for `Parse`, `ParseMap` or the document API: the tree remains theirs. @@ -561,8 +558,10 @@ sequenceDiagram ### Input constraints `Marshal` and `MarshalWrite` accept a `struct`, a `map[string]V`, or a -non-nil pointer to one, where `V` is any value `Marshal` itself understands. A -different top-level value fails: +non-nil pointer to one, where `V` is any value `Marshal` itself understands. +An `OrderedMap` and a `Document` are accepted as themselves: the first in its +written key order, the second written back as it stands. A different +top-level value fails: | Input | Error | |---|---| @@ -736,7 +735,7 @@ across a round-trip. A nil slice is always omitted. An empty (length 0) array of tables is always omitted, because TOML forbids an empty `[[a]]`. Other empty arrays emit as -`key = []` by default; `OmitEmptyArrays()` skips them as well, so +`key = []` by default; `OmitEmptyArrays(true)` skips them as well, so `[]string{}` is treated like a nil slice. ### Long strings @@ -863,7 +862,7 @@ encoding/json/v2 made current, with the differences TOML asks for: | `encoding.TextMarshaler`, `TextUnmarshaler` | honoured, the same | a type that renders itself as text becomes a TOML string, both ways | | `*json.UnmarshalTypeError` | `*DecodeError` | the path is segments with a `String()` renderer, not a dotted string | | `*json.SyntaxError` | `*SyntaxError` | the TOML error adds the byte `Offset` and the `Column` to the line | -| context support | `*Context` variants of every entry point | encoding/json has none | +| context support | `*Context` variants of the parse, decode and marshal entries | encoding/json has none | ## Types diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 5f84e04..694727d 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -5,16 +5,17 @@ source tree; nothing is aspirational. ## Overview -interpres is one public library package, one command, and one example. The +interpres is one public library package, one command, and two examples. The library implements the whole of TOML 1.1, decoding and encoding, in the standard library alone; the command wraps the parser and the encoder for the toml-test compliance harness, against which it stands at 214 valid, 467 invalid -and 214 encoder cases with zero failures; the example demonstrates the API. +and 214 encoder cases with zero failures; the examples demonstrate the API: one the document round trip, one the statement iterator. ```mermaid flowchart TD CLI[cmd/interpres-decode
toml-test adapter] --> API EX[examples/basic
usage demo] --> API + EX2[examples/statements
statement iterator demo] --> API subgraph Lib [package interpres] API[interpres.go
public API and types] API --> P[parser.go
recursive-descent parser] @@ -35,9 +36,10 @@ strict validation. | Path | Responsibility | |---|---| -| `.` (package `interpres`) | The whole library. `interpres.go` declares the exported surface (`Parse`, `Unmarshal`, `Marshal`, the `*Context` variants, the option constructors, `Marshaler`, `Unmarshaler`, `SyntaxError`, the local date-time types); everything below it is unexported. | -| `cmd/interpres-decode` | The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with `-encode` it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic. | +| `.` (package `interpres`) | The whole library. `interpres.go` declares the exported surface (`Parse`, `Unmarshal`, `Marshal`, the `*Context` variants, the option constructors, `Marshaler`, `Unmarshaler`, `SyntaxError`, the error and option types); everything below it is unexported. | +| `cmd/interpres-decode` | The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with `--encode` it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic. | | `examples/basic` | A runnable tour of the API. Documentation in executable form, not part of the library. | +| `examples/statements` | The `Statements` iterator over a document, the shape a configuration tool reads. Documentation in executable form. | Inside the library package, one file owns one concern: @@ -46,18 +48,28 @@ Inside the library package, one file owns one concern: | `parser.go` | The recursive-descent parser. Produces the `map[string]any` tree, records the nodes a [Document](API.md#documents) is built from, and enforces the structural rules of TOML 1.1 (table redefinitions, dotted keys, arrays of tables, multi-line inline tables). Reports a 1-based line on failure. | | `document.go` | The parsed-document types: `Document`, `Table` and `Entry`, which carry the key order, whether a table was written inline, and the comments. The values they expose are the parser's own tree, not a copy. | | `number.go` | Strict numeric tokens: integers in the four radixes with `_` separators, and floats including `inf` and `nan`. Rejects leading zeros, misplaced underscores and malformed fractions. | -| `datetime.go` | The three local date-time wrapper types and `parseDateTime`, which classifies a token into the four date-time kinds under the strict TOML grammar. | +| `datetime.go` | The four date-time types (`OffsetDateTime` and the three local wrappers) and `parseDateTime`, which classifies a token into the four date-time kinds under the strict TOML grammar. | +| `orderedmap.go` | `OrderedMap`, the table that keeps its key order, and the node index the decoder reads the written order from. | +| `target.go` | The targeted parse: the struct skeleton resolved against the document while it scans, no intermediate tree. Falls back to the tree path for every shape it does not model. | +| `docwrite.go` | The write side of the document pipeline: `UnmarshalDocument` and the writer that renders a `Document` back with its order and comments. | | `decode.go` | Maps the parsed tree onto Go values by reflection: struct fields, maps, slices, scalar conversion with overflow checks, `Unmarshaler` dispatch. | | `encode.go` | The reverse walk: builds an intermediate `tomlDoc` per table (which is what preserves declaration order and enables the group-by-kind partition) and then emits it as TOML. | The boundary that matters: `parser.go` produces only untyped trees -(`map[string]any`, `[]any`, `[]map[string]any`, scalars); `decode.go` and -`encode.go` are the only files that touch `reflect`; the command never touches -either, it consumes `Parse` alone. +(`map[string]any`, `[]any`, `[]map[string]any`, scalars); the reflection work +lives in `decode.go`, `encode.go`, `target.go` and `orderedmap.go`; the command +consumes `ParseMap`, `Parse` and `Marshal`, and owns no parsing or emission +logic of its own. ## Data flow -Decoding is parse, then one reflection walk. `SyntaxError` values are produced +Decoding has two paths. The direct one parses straight into a struct +destination: `target.go` resolves the table skeleton against the struct +schema while the document scans, and values assign through the ordinary +decoder rules, so no intermediate tree exists; that is the hot path every +`Unmarshal` into a struct takes. A document or destination the direct +skeleton cannot model falls back to the tree path: parse the whole document, +then one reflection walk over the tree. `SyntaxError` values are produced inside `parser.go` and returned as-is; conversion errors are produced inside `decode.go` and wrapped with the key path as they unwind. @@ -66,10 +78,13 @@ sequenceDiagram participant Caller participant API as interpres.go participant P as parser.go + participant T as target.go participant D as decode.go Caller->>API: Unmarshal(data, v) - API->>P: ParseContext(ctx, data) - P->>P: number and datetime atoms + API->>T: targeted parse into the struct + T->>P: scanner, grammar, atoms + T-->>API: result, error or fallback + API->>P: on fallback, ParseContext(ctx, data) P-->>API: map tree or *SyntaxError API->>D: decode(tree, reflect value) D-->>API: nil or wrapped field error diff --git a/docs/BENCHMARKING.md b/docs/BENCHMARKING.md index 00831f4..10cacfa 100644 --- a/docs/BENCHMARKING.md +++ b/docs/BENCHMARKING.md @@ -15,6 +15,8 @@ The benchmarks live in `bench_test.go`, next to the code they measure: | `BenchmarkParseLong` | `ParseMap` over a generated document with about 2000 array-of-tables entries | | `BenchmarkStrictDecodeLong` | `Unmarshal` into a typed document under `RejectUnknownFields`, over the same long document | | `BenchmarkMarshalLong` | `Marshal` of the tree `ParseMap` produced from the long document | +| `BenchmarkStrictDecodeTree` | the tree-path reference decode of the representative document: parse, then the reflection walk | +| `BenchmarkStrictDecodeTreeLong` | the tree-path reference decode of the long document, the A/B baseline of the targeted parse | ## Running diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md index 2179b68..a8d0132 100644 --- a/docs/DEVELOPMENT.md +++ b/docs/DEVELOPMENT.md @@ -46,6 +46,9 @@ prints the same list. | `just install` | builds, then copies the binary into `~/.local/bin` (`BINDIR` overrides) | | `just uninstall` | removes the installed binary | | `just clean` | removes `bin/` and `coverage.out` | +| `just cross` | cross-compile smoke of the library and the command for arm64, loong64, riscv64 and the browser and edge runtimes; nightly convenience, not a gate | +| `just release-check X.Y.Z` | the release pre-flight: the branch, a clean tree, a sync with origin, the gates, and a CHANGELOG section ready to release | +| `just docs-drift` | compares the toml-test counts the documentation quotes with a live suite run | ## Running a single test @@ -98,7 +101,8 @@ pipeline. | Workflow | Trigger | What it does | |---|---|---| | `test.yml` | push or pull request to `development` | format check, vet, modernisation, build, the test suite with the 80 percent coverage floor, then the toml-test compliance suite | -| `race.yml` | `workflow_dispatch`, by hand | the suite under the race detector; the same race gate `just gates` runs locally | +| `race.yml` | `workflow_dispatch`, plus a nightly schedule at 03:00 UTC | the suite under the race detector, the same race gate `just gates` runs locally; the nightly schedule runs on the default branch, the released line | +| `fuzz.yml` | `workflow_dispatch`, plus a nightly schedule at 03:30 UTC | a 30 second fuzz smoke per target over the seeds and the gathered corpus | | `release.yml` | a `v*` tag | tag validation, then format, vet, modernisation, build and the test suite with the coverage floor, then the Gitea release from the CHANGELOG section. No race detector: race never runs on a push path, and the local `just gates` raced the tree before the tag was cut | ## Releases