Files
interpres/CHANGELOG.md
T
petrbalvin d92bb56853
Test / test (push) Successful in 1m35s
refactor: rename the encoder layout options
Assisted-by: GLM 5.3 Flash
2026-09-22 01:09:00 +02:00

389 lines
23 KiB
Markdown

# Changelog
All notable changes to **interpres** are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [development]
### Added
- `encoding.TextMarshaler` and `encoding.TextUnmarshaler` are honoured by
default, with no option to switch them off. A type that implements them is
encoded as a TOML string and decoded from one: `net.IP` becomes
`"192.0.2.1"`, and a user type with `MarshalText` or `UnmarshalText` follows.
`MarshalTOML` and `UnmarshalTOML` still win over the text methods, and the
four date-time types keep their bare timestamp form instead of becoming a
quoted string. A struct type that implements the interface now encodes as a
string where it was a table before, which is the breaking part of the change.
- `time.Duration` is encoded in its canonical Go form as a TOML string,
`1h30m0s`, because TOML has no duration type; the decoder reads that string
back and still accepts a bare integer as the nanosecond count.
- `interpres-decode -encode`, the adapter's other direction: it reads the
toml-test tagged JSON from stdin and writes the TOML document it describes.
The compliance suite now runs the encoder as well as the decoder, 214
encoder cases against the tagged JSON of the valid corpus.
- `Encoder.InlineTables(threshold)`: a sub-table whose single-line rendering is
at most `threshold` bytes is written as an inline table instead of a header
section, which shortens a document of small tables. An array of tables keeps
its header form, because its inline form would re-parse as a value array.
- `Document` and `ParseMap`: `Parse` now returns a `*Document`, which holds the
values together with the key order, whether a table was written as an inline
table or under a header, and the comments, with `Keys`, `Entries`, `Get`,
`Comments` and `SetComments` to read and write them. `ParseMap` returns the
plain `map[string]any` tree, the shape `Parse` used to give. `Marshal` does
not accept a `Document`; it writes values, so `doc.Map()` is the way through.
- `OffsetDateTime`, the Go type of the offset date-time kind, so that all four
TOML date-time kinds have one of their own. `Parse` and `Unmarshal` hand it
back where they produced a bare `time.Time` before, and `Marshal` accepts it.
Unmarshalling into a struct field of type `time.Time` keeps working, because
the plain type takes an offset date-time as it always did; code that asserts
the tree's type, and `UnmarshalTOML` implementations that expect a
`time.Time`, need the new type.
- `Decoder.MaxDepth(depth)` and `Decoder.MaxInputSize(size)` bound the parse a
`Decode` performs, and every parse carries a nesting limit in any case
(10000 levels, which no hand-written document approaches): a document that
nests arrays or inline tables deeper used to run the stack out and is now
rejected with a `SyntaxError` naming the limit.
- `ParseAs[T](data)`, the generic one-line decode, and `NewSchema[T]()`,
which precompiles the struct schema and the interface flags for a hot path
before the first document arrives.
- `Encoder.EmitFieldComments()` prints the comment a field's `toml` tag
carries in a `comment=` option above the field's line or header, the
comments a round trip through the Go type would otherwise drop. Go doc
comments are not visible to reflection, so the tag is the channel that
carries the text.
- `Decoder.LocalTimeLocation(loc)` lets a local date-time fill a plain
`time.Time` destination in the location given, relabelled rather than
shifted: `07:32` in the document is `07:32` in the zone. Without the
option the wrapper types remain the only destinations a local kind fills.
- The parse checks its context inside a value as well as between statements:
an array, an inline table and a multi-line string check every 64 elements
or lines, so one huge value cannot hold the parse past its cancellation.
- `OrderedMap`, the string-keyed table that remembers the order its keys were
set in: decoding into one fills it in the order the document wrote the
keys, and `Marshal` writes one back in that order, where a map carries no
order on decode and sorts on encode. It works as a decode target on its
own, in a struct field, and as the element of an array of tables; its
values are untyped, so a nested table stays a `map[string]any`.
- `UnmarshalWithOptions(data, v, opts)` decodes with a `DecodeOptions` struct
in one call, the options a `Decoder` sets without building one: unknown
keys, `Number` literals, and the parse limits.
- `Marshal` carries a nesting limit of 10000 levels, the parser's own figure:
cyclic data, which used to run the stack out, is now rejected with an error
that names the limit and the path it was met at.
- The `toml` tag gained the `required` option: a field tagged
`toml:"host,required"` makes the decode fail with
`missing required key "host"` when the document carries no key that
resolves to it. The option shapes decoding only, and the encoder ignores
it.
- `UnmarshalerContext`, the custom-decode interface that hands the decode's
context to the method, `UnmarshalTOMLContext(ctx, data)`. It wins over
`UnmarshalTOML` when a type implements both, so a long custom decode can
abort on cancellation; a non-cancellable entry point hands in
`context.Background`, never nil.
- A TOML array decodes into a Go fixed-size array, `[N]T`, where only a slice
was accepted before; the encoder could already encode one. A length mismatch
is an error wrapped with the key path.
- `MarshalAppend(buf, v)` appends the TOML encoding of v to buf and returns
the extended buffer, the shape `json.MarshalAppend` has.
- `interpres-decode -version` prints the binary's version, the module version
the toolchain recorded, so a release-built binary names its own tag.
- `interpres-decode -json` prints plain indented JSON instead of the tagged
form, the shape for people and diffs, with the date-time wrappers in their
TOML form.
- `interpres-decode -validate` walks a named directory for `.toml` files and
closes the sweep with a summary naming how many documents were checked and
how many were invalid; single files stay quiet on success as before.
- `interpres-decode -struct` infers a Go struct definition from a document:
one field per key in written order, nested tables as nested struct types,
an array of tables as a slice. The printed type compiles and decodes the
document it came from.
- `interpres-decode -schema TYPE file.go` writes a TOML template for the
named struct type of a Go source, the `comment=` tag option printed as a
comment and the `default=` option as the value. It is the inverse of
`-struct`.
- `ParseFile(path)` reads the file and parses it into a `Document`, with the
file name at the front of every error it returns, read failure and parse
failure alike. `Valid(data)` reports whether a document parses, nil on
success and the parse error on failure, the library call the `-validate`
mode of interpres-decode is built on.
- `SyntaxError` carries the byte `Offset` the scan stopped at and the 1-based
`Column` on the line, beside the line it always had, and `SourceLine(src)`
renders that line with a caret under the position, for messages shown under
the input. An input that is not valid UTF-8 names the offset of the first
invalid byte in its message. The new fields are additive: a `SyntaxError`
built from a line and a message alone is unchanged.
- `Decoder.UseNumber()` decodes the integers and floats of the document into
`Number`, which carries the literal the document wrote, so `0x1f`, `1_000`,
`+1.0` and `inf` survive a round trip with their spelling instead of the
normalised `31`, `1000` and `1.0`. Typed destinations take the evaluated
value as before, a `Number` field takes the literal, and `Marshal` writes a
`Number` back as its bare literal, rejecting one that is not a valid TOML
number.
### Changed
- `Encoder.GroupByKind(bool)` is renamed to `Encoder.Layout(kind)` and takes
a `LayoutKind`: `LayoutKindGrouped`, the default, or
`LayoutKindDeclaration` for the declaration order. `UseLiteralMultiline`
is renamed to `LiteralMultiline`. The behaviour is unchanged; 2.0 is the
only chance a rename has, and the migrator updates the calls mechanically.
- `DecodeError` and `EncodeError` carry one `Path` type, a list of segments
(`"items"`, `"[0]"`, `"weight"`) with a `String()` rendering the TOML
notation, `items[0].weight`. The decode error used to hold a bare
`[]string`, the encode error a plain string. Both messages render the same
way now, `interpres: items[0].weight: ...`, with one `interpres:` prefix
where the composition used to double it.
- `omitempty` follows the encoding/json semantics: the field is skipped when
it holds an empty string, a zero number, `false`, a nil pointer or
interface, or a nil or empty slice, array or map. In 1.x the option covered
only the collections.
- The `toml` tag gained the `inline` option: a struct or map field tagged
`toml:"retry,inline"` emits as `retry = {…}` instead of a header section,
whatever its size, a named embedded struct included. Forcing it on an array
of tables is an error, because the inline form would re-parse as a value
array and change the value's Go type.
- The output takes the TOML 1.1 form. A date-time writes its seconds only when
the value carries them and drops the trailing zeros of a fractional second,
so `07:32:00` is written `07:32` and half a second as `00.5`. Both are the
same value, and a document written without seconds now comes back without
them. `LocalDateTime.String()`, `LocalTime.String()` and the offset date-time
rendering follow the same rule.
- An inline table that would pass the hundredth column is written across lines
with a trailing comma and one tab of indentation per nesting level, the shape
TOML 1.1 allows an inline table to take.
- `MarshalTOML` reaches every array element and every field, whatever the Go
kind, and its result is normalised like any other value: an element rendering
itself as a table keeps the `[[header]]` form, one rendering itself as a
scalar turns the array into a value array, and the method runs once per
element. It is found on the addressable pointer as well, so a
pointer-receiver `MarshalTOML` is called for a field or an element, exactly
as `MarshalText` is.
- TOML 1.1 is the acceptance contract, and TOML 1.0 is not. The compliance
suite runs the 1.1 corpus alone, and the promise that every 1.0 document
parses exactly as before is withdrawn. Nothing that parses today stops
parsing: the 1.0 valid corpus still passes in full. The documents whose
verdict changes are the ones 1.1 relaxed, such as the `\xHH` escape
sequences 1.0 rejected.
- The module path carries the /v2 suffix the Go toolchain requires of
every major version 2 module: imports change to
`sourcedock.dev/petrbalvin/interpres/v2`.
- Input that is not valid UTF-8 is now rejected where the parser's scan
meets the invalid byte, with a `SyntaxError` naming that line, instead of
a whole-input check that always reported line 1. Invalid input is still
rejected; the reported location is now the byte's own.
**Performance**
- Parsing is faster than in 1.1.0 while carrying the new document layer:
the suite's representative document decodes at about 79 MB/s with 104
allocations per call, and the long array-of-tables document at about
106 MB/s against 56 MB/s in 1.1.0, with allocations on that document
halved from 67 664 to 31 765. Date-time tokens are validated by a byte
scan instead of regular expressions, repeated keys share one string
across array-of-tables elements, and per-statement buffers are reused.
- Typed decoding is 12 percent faster than in 1.1.0 on the representative
document (9792 ns against 11 147 ns) with 24 percent fewer allocations
(167 against 220); interface lookups resolve through a cached per-type
flag set instead of boxing every value into an interface to ask.
- `Marshal` runs at the 1.1.0 speed while emitting the new TOML 1.1 output
form, at half the bytes per operation (6170 against 11 348 on the
representative document), and writes through a pooled output buffer with
a 1 MiB retention cap; repeated marshals keep the live heap flat.
- Two benchmarks measure the shapes that drove the work:
`BenchmarkStrictDecodeLong` and `BenchmarkMarshalLong` run the 2000-entry
document at about 3.8 ms and 3.4 ms per call, at 63 772 and 63 660
allocations.
### Fixed
- An offset date-time written with the `+00:00` offset kept the anonymous
location `time.Parse` invents for it, so a round trip through the tree and
`Marshal`, which writes a zero offset as `Z`, changed the value's
reflection-visible location. The zero offset normalises to `time.UTC` at
parse, and the tree is stable across the round trip.
- Decoding into a defined type whose underlying kind is string or bool, such
as `type Name string`, panicked instead of storing the value, because a
value of the predeclared type is not assignable to a defined type and the
decoder assigned it without a conversion.
- A top-level value the encoder could not normalise reported its path with
a leading dot, `interpres: .port: ...`; the message now reads
`interpres: port: ...`, the shape `EncodeError.Path` already used.
## [1.1.0] - 2026-09-18
### Added
- TOML 1.1 support, on by default: date-times and times without seconds
(`07:32`, `1979-05-27T07:32`, normalised to full seconds on output), the
`\e` and `\xHH` escape sequences, and multi-line inline tables with
comments and trailing commas. The compliance suite runs in TOML 1.1 mode:
214 valid and 467 invalid cases, zero failures. Every TOML 1.0 document
parses exactly as before.
- `interpres-decode -validate [file ...]`: a validate mode beside the
toml-test adapter. It parses each named file, or stdin when none are named,
prints one line per invalid document to stderr, and exits 0 when all are
valid, 1 when one is not, and 2 on a usage or read failure. Install it with
`go install .../cmd/interpres-decode@latest`; releases still ship no
binaries.
- `DecodeError` and `EncodeError`: decode and encode failures are wrapped in
typed errors carrying the key path, read with `errors.AsType` instead of
parsing the message text. The rendered messages keep their shape; the only
visible change is that an encode failure on a top-level field no longer
gains a meaningless leading dot in its path.
- `omitzero` and `omitempty` tag options on encode: `toml:"name,omitzero"`
skips a field whose value is the zero value of its type (a type with an
`IsZero() bool` method decides through the method), and
`toml:"name,omitempty"` skips a nil or empty slice, array, or map. The
decoder ignores both options.
### Changed
- The compliance suite is [toml-test](https://github.com/toml-lang/toml-test)
v2.2.0, up from v1.6.0. Its TOML 1.0 corpus holds 205 valid and 474 invalid
cases (185 and 371 before), and it caught the two documents the parser
still accepted, fixed below.
- The flattened struct layout the decoder consults is cached per struct type
and shared with the encoder, which now resolves duplicate field keys with
it. Strict decoding of an array of tables of structs runs about a quarter
faster; marshalling structs gained the same layout without measurable cost.
- The parser scans the input bytes in place instead of building a `[]rune`
copy of the document: every character that drives the grammar is ASCII and
the input is validated UTF-8 up front, so the conversion pass and its four
bytes per rune were pure overhead. Parsing a large array-of-tables document
runs about a fifth faster and allocates about half the memory.
- Numeric tokens without underscores skip the normalising rebuild: digits are
validated in place in `joinDigits`, and a float whose token is already
clean goes to `strconv.ParseFloat` directly. One allocation per integer
atom and two per float atom disappear.
### Fixed
- A `MarshalTOML` result of `nil` with a nil error fails the marshal with
`MarshalTOML returned a nil value`. The field silently vanished before, and
inside a value array the nil result reached reflection as a zero value and
panicked.
- Strict decoding reports the smallest unknown key. Several unknown keys in
one table made the message depend on Go's random map iteration order, so
the same document reported different keys across runs.
- Decoding into a struct that embeds a pointer to itself terminates. The
schema walk recursed through the embedded type forever, so such a
`Unmarshal` call hung the process; the walk now tracks the struct types on
the current path and stops when one repeats.
- An array-of-tables header whose path runs through an inline table
(`a = {b = {}}` followed by `[[a.b.c]]`) is rejected. The frozen-inline-table
check covered `[table]` headers and dotted keys but not the intermediate
steps of an array-of-tables header, so such a document silently extended the
inline table.
- A new element of an array of tables starts a fresh scope for dotted-key paths
and nested arrays of tables: `[[a]]`, `b.c = 1`, `[[a]]`, `[a.b]` parses, as
the TOML examples in the spec shape it. The records of the previous element
falsely rejected the same paths in the next one.
- `Marshal` emits exactly one key when two struct fields resolve to the same
TOML name, picking the field the decoder would fill (the shallower one, the
later declaration at equal depth). Such a struct previously marshalled into
a duplicate key, and the output never re-parsed, breaking the round-trip
guarantee.
- `Marshal` returns an error for a table header key or an inline-table key that
is not valid UTF-8, the way scalar keys already did, instead of silently
emitting corrupt TOML (a header that lost its key, an inline table with a
missing key).
- `UseLiteralMultiline` falls back to the escaped basic string when the value
cannot be carried verbatim by the literal form: a run of three single quotes,
a control character, or a lone carriage return. Such values previously
produced output that did not re-parse.
- A `[]any` holding only tables marshals in the value-array form with inline
tables, keeping the type `Parse` produces for such an array. It previously
took the `[[header]]` form, so a round-trip changed the value's type from
`[]any` to `[]map[string]any`.
- Decoding into a `uint` destination checks the type's platform width instead
of only the fixed widths, so a 32-bit `uint` no longer truncates silently;
decoding a finite float beyond the `float32` range is an overflow error
instead of a silent infinity.
- Struct fields that resolve to one key at equal depth decode through the
field declared later, matching the documented rule; the first one won before.
- A float with an exponent marker but no digits (`1e`, `0.0E`) is rejected;
the exponent requires at least one digit.
- A date-time offset outside 00:00 through 23:59 is rejected; such offsets
were accepted and silently rolled over (`+00:60` decoded as `+01:00`).
- Untagged embedded fields now decode symmetrically with encode: an embedded
struct receives its keys inline (a nil embedded pointer struct is
allocated), an embedded map catches the keys no field claims, and a name
clash resolves in favour of the shallower field. A struct with an untagged
embedded field previously decoded with all inline keys dropped and did not
round-trip.
- `Marshal` re-emits arrays that mix tables with scalars: the table elements
render as inline tables inside the value array. A tree that `Parse` accepts
from such a document previously failed with
`cannot encode map[string]interface {}`.
## [1.0.0] - 2026-08-20
First stable release: a dependency-free TOML 1.0 parser and encoder for Go that
uses only the standard library and passes the entire
[toml-test](https://github.com/toml-lang/toml-test) suite
(**185 valid and 371 invalid cases, 0 failures**).
### Added
**Parsing**
- `Parse(data []byte) (map[string]any, error)`: decode a TOML document into an
untyped tree.
- Full TOML 1.0 syntax: comments; bare, quoted, and dotted keys; tables
(`[a.b]`) and arrays of tables (`[[a]]`); basic and literal strings, including
multiline (`"""` / `'''`) with escape sequences and line-ending backslash
trimming; integers in decimal, hex (`0x`), octal (`0o`), and binary (`0b`)
with `_` separators; floats with exponents and `inf` / `nan`; booleans;
offset/local date-times, dates, and times; arrays; and inline tables.
- Distinct date-time types: offset date-times decode to `time.Time`, while
`LocalDateTime`, `LocalDate`, and `LocalTime` represent the local variants,
each with a `String()` method returning the TOML-canonical rendering.
- Strict, spec-conformant validation that rejects invalid documents: leading
zeros, misplaced underscores, malformed floats and radix literals, control
characters in strings and comments, bare carriage returns, non-UTF-8 input,
out-of-range Unicode escapes, multiline strings used as keys, single-digit
date-time components, duplicate/overwriting inline-table keys, and the full
family of table redefinitions (header vs. header, header vs. dotted key,
array of tables vs. table, and dotted-key appends to defined tables).
- `SyntaxError` carrying the 1-based line of a malformed document.
**Decoding**
- `Unmarshal(data []byte, v any)`: parse and map onto a struct or
`map[string]any` via reflection, with overflow-checked numeric conversion,
nested structs, slices, and `map[string]T`.
- `toml:"name"` field tags, case-insensitive name fallback, and `toml:"-"` to
skip a field.
- `Decoder` with `DisallowUnknownFields` for strict decoding that rejects keys
without a destination field, at every struct depth.
- `Unmarshaler` interface (`UnmarshalTOML(data any) error`) for types that take
full control of their decode.
**Encoding**
- `Marshal(v any) ([]byte, error)`: encode a struct or `map[string]V` value to
a TOML 1.0 document that re-parses to an equivalent value tree.
- `Encoder` with chainable policy options: `GroupByKind` (group-by-kind layout
versus declaration order), `OmitEmptyArrays`, and `UseLiteralMultiline`.
- `Marshaler` interface (`MarshalTOML() (any, error)`) for types that need a
custom TOML shape; the returned value is encoded normally.
**Cancellation**
- `ParseContext`, `UnmarshalContext`, `MarshalContext`,
`(*Decoder).DecodeContext` and `(*Encoder).MarshalContext` honour a
`context.Context`, checked up front and every 64 statements or fields.
**Project**
- Standard library only: zero third-party modules.
- `cmd/interpres-decode`, a toml-test harness adapter (TOML on stdin, tagged
JSON on stdout).
- A runnable example, a table-driven Go test suite, and a canonical `just`
recipe set whose `gates` recipe is the definition of done.
- Hand-written CI pipelines for the test, race and release gates.
- `SECURITY.md` for private vulnerability reports.