Files
interpres/CHANGELOG.md
T

277 lines
15 KiB
Markdown

# Changelog
All notable changes to **interpres** are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [development]
### Added
- `encoding.TextMarshaler` and `encoding.TextUnmarshaler` are honoured by
default, with no option to switch them off. A type that implements them is
encoded as a TOML string and decoded from one: `net.IP` becomes
`"192.0.2.1"`, and a user type with `MarshalText` or `UnmarshalText` follows.
`MarshalTOML` and `UnmarshalTOML` still win over the text methods, and the
four date-time types keep their bare timestamp form instead of becoming a
quoted string. A struct type that implements the interface now encodes as a
string where it was a table before, which is the breaking part of the change.
- `time.Duration` is encoded in its canonical Go form as a TOML string,
`1h30m0s`, because TOML has no duration type; the decoder reads that string
back and still accepts a bare integer as the nanosecond count.
- `interpres-decode -encode`, the adapter's other direction: it reads the
toml-test tagged JSON from stdin and writes the TOML document it describes.
The compliance suite now runs the encoder as well as the decoder, 214
encoder cases against the tagged JSON of the valid corpus.
- `Encoder.InlineTables(threshold)`: a sub-table whose single-line rendering is
at most `threshold` bytes is written as an inline table instead of a header
section, which shortens a document of small tables. An array of tables keeps
its header form, because its inline form would re-parse as a value array.
- `Document` and `ParseMap`: `Parse` now returns a `*Document`, which holds the
values together with the key order, whether a table was written as an inline
table or under a header, and the comments, with `Keys`, `Entries`, `Get`,
`Comments` and `SetComments` to read and write them. `ParseMap` returns the
plain `map[string]any` tree, the shape `Parse` used to give. `Marshal` does
not accept a `Document`; it writes values, so `doc.Map()` is the way through.
- `OffsetDateTime`, the Go type of the offset date-time kind, so that all four
TOML date-time kinds have one of their own. `Parse` and `Unmarshal` hand it
back where they produced a bare `time.Time` before, and `Marshal` accepts it.
Unmarshalling into a struct field of type `time.Time` keeps working, because
the plain type takes an offset date-time as it always did; code that asserts
the tree's type, and `UnmarshalTOML` implementations that expect a
`time.Time`, need the new type.
- `Decoder.MaxDepth(depth)` and `Decoder.MaxInputSize(size)` bound the parse a
`Decode` performs, and every parse carries a nesting limit in any case
(10000 levels, which no hand-written document approaches): a document that
nests arrays or inline tables deeper used to run the stack out and is now
rejected with a `SyntaxError` naming the limit.
### Changed
- The output takes the TOML 1.1 form. A date-time writes its seconds only when
the value carries them and drops the trailing zeros of a fractional second,
so `07:32:00` is written `07:32` and half a second as `00.5`. Both are the
same value, and a document written without seconds now comes back without
them. `LocalDateTime.String()`, `LocalTime.String()` and the offset date-time
rendering follow the same rule.
- An inline table that would pass the hundredth column is written across lines
with a trailing comma and one tab of indentation per nesting level, the shape
TOML 1.1 allows an inline table to take.
- `MarshalTOML` reaches every array element and every field, whatever the Go
kind, and its result is normalised like any other value: an element rendering
itself as a table keeps the `[[header]]` form, one rendering itself as a
scalar turns the array into a value array, and the method runs once per
element. It is found on the addressable pointer as well, so a
pointer-receiver `MarshalTOML` is called for a field or an element, exactly
as `MarshalText` is.
- TOML 1.1 is the acceptance contract, and TOML 1.0 is not. The compliance
suite runs the 1.1 corpus alone, and the promise that every 1.0 document
parses exactly as before is withdrawn. Nothing that parses today stops
parsing: the 1.0 valid corpus still passes in full. The documents whose
verdict changes are the ones 1.1 relaxed, such as the `\xHH` escape
sequences 1.0 rejected.
- The module path carries the /v2 suffix the Go toolchain requires of
every major version 2 module: imports change to
`sourcedock.dev/petrbalvin/interpres/v2`.
- Input that is not valid UTF-8 is now rejected where the parser's scan
meets the invalid byte, with a `SyntaxError` naming that line, instead of
a whole-input check that always reported line 1. Invalid input is still
rejected; the reported location is now the byte's own.
**Performance**
- Parsing is faster than in 1.1.0 while carrying the new document layer:
the suite's representative document decodes at about 79 MB/s with 104
allocations per call, and the long array-of-tables document at about
106 MB/s against 56 MB/s in 1.1.0, with allocations on that document
halved from 67 664 to 31 765. Date-time tokens are validated by a byte
scan instead of regular expressions, repeated keys share one string
across array-of-tables elements, and per-statement buffers are reused.
- Typed decoding is 12 percent faster than in 1.1.0 on the representative
document (9792 ns against 11 147 ns) with 24 percent fewer allocations
(167 against 220); interface lookups resolve through a cached per-type
flag set instead of boxing every value into an interface to ask.
### Fixed
- Decoding into a defined type whose underlying kind is string or bool, such
as `type Name string`, panicked instead of storing the value, because a
value of the predeclared type is not assignable to a defined type and the
decoder assigned it without a conversion.
## [1.1.0] - 2026-09-18
### Added
- TOML 1.1 support, on by default: date-times and times without seconds
(`07:32`, `1979-05-27T07:32`, normalised to full seconds on output), the
`\e` and `\xHH` escape sequences, and multi-line inline tables with
comments and trailing commas. The compliance suite runs in TOML 1.1 mode:
214 valid and 467 invalid cases, zero failures. Every TOML 1.0 document
parses exactly as before.
- `interpres-decode -validate [file ...]`: a validate mode beside the
toml-test adapter. It parses each named file, or stdin when none are named,
prints one line per invalid document to stderr, and exits 0 when all are
valid, 1 when one is not, and 2 on a usage or read failure. Install it with
`go install .../cmd/interpres-decode@latest`; releases still ship no
binaries.
- `DecodeError` and `EncodeError`: decode and encode failures are wrapped in
typed errors carrying the key path, read with `errors.AsType` instead of
parsing the message text. The rendered messages keep their shape; the only
visible change is that an encode failure on a top-level field no longer
gains a meaningless leading dot in its path.
- `omitzero` and `omitempty` tag options on encode: `toml:"name,omitzero"`
skips a field whose value is the zero value of its type (a type with an
`IsZero() bool` method decides through the method), and
`toml:"name,omitempty"` skips a nil or empty slice, array, or map. The
decoder ignores both options.
### Changed
- The compliance suite is [toml-test](https://github.com/toml-lang/toml-test)
v2.2.0, up from v1.6.0. Its TOML 1.0 corpus holds 205 valid and 474 invalid
cases (185 and 371 before), and it caught the two documents the parser
still accepted, fixed below.
- The flattened struct layout the decoder consults is cached per struct type
and shared with the encoder, which now resolves duplicate field keys with
it. Strict decoding of an array of tables of structs runs about a quarter
faster; marshalling structs gained the same layout without measurable cost.
- The parser scans the input bytes in place instead of building a `[]rune`
copy of the document: every character that drives the grammar is ASCII and
the input is validated UTF-8 up front, so the conversion pass and its four
bytes per rune were pure overhead. Parsing a large array-of-tables document
runs about a fifth faster and allocates about half the memory.
- Numeric tokens without underscores skip the normalising rebuild: digits are
validated in place in `joinDigits`, and a float whose token is already
clean goes to `strconv.ParseFloat` directly. One allocation per integer
atom and two per float atom disappear.
### Fixed
- A `MarshalTOML` result of `nil` with a nil error fails the marshal with
`MarshalTOML returned a nil value`. The field silently vanished before, and
inside a value array the nil result reached reflection as a zero value and
panicked.
- Strict decoding reports the smallest unknown key. Several unknown keys in
one table made the message depend on Go's random map iteration order, so
the same document reported different keys across runs.
- Decoding into a struct that embeds a pointer to itself terminates. The
schema walk recursed through the embedded type forever, so such a
`Unmarshal` call hung the process; the walk now tracks the struct types on
the current path and stops when one repeats.
- An array-of-tables header whose path runs through an inline table
(`a = {b = {}}` followed by `[[a.b.c]]`) is rejected. The frozen-inline-table
check covered `[table]` headers and dotted keys but not the intermediate
steps of an array-of-tables header, so such a document silently extended the
inline table.
- A new element of an array of tables starts a fresh scope for dotted-key paths
and nested arrays of tables: `[[a]]`, `b.c = 1`, `[[a]]`, `[a.b]` parses, as
the TOML examples in the spec shape it. The records of the previous element
falsely rejected the same paths in the next one.
- `Marshal` emits exactly one key when two struct fields resolve to the same
TOML name, picking the field the decoder would fill (the shallower one, the
later declaration at equal depth). Such a struct previously marshalled into
a duplicate key, and the output never re-parsed, breaking the round-trip
guarantee.
- `Marshal` returns an error for a table header key or an inline-table key that
is not valid UTF-8, the way scalar keys already did, instead of silently
emitting corrupt TOML (a header that lost its key, an inline table with a
missing key).
- `UseLiteralMultiline` falls back to the escaped basic string when the value
cannot be carried verbatim by the literal form: a run of three single quotes,
a control character, or a lone carriage return. Such values previously
produced output that did not re-parse.
- A `[]any` holding only tables marshals in the value-array form with inline
tables, keeping the type `Parse` produces for such an array. It previously
took the `[[header]]` form, so a round-trip changed the value's type from
`[]any` to `[]map[string]any`.
- Decoding into a `uint` destination checks the type's platform width instead
of only the fixed widths, so a 32-bit `uint` no longer truncates silently;
decoding a finite float beyond the `float32` range is an overflow error
instead of a silent infinity.
- Struct fields that resolve to one key at equal depth decode through the
field declared later, matching the documented rule; the first one won before.
- A float with an exponent marker but no digits (`1e`, `0.0E`) is rejected;
the exponent requires at least one digit.
- A date-time offset outside 00:00 through 23:59 is rejected; such offsets
were accepted and silently rolled over (`+00:60` decoded as `+01:00`).
- Untagged embedded fields now decode symmetrically with encode: an embedded
struct receives its keys inline (a nil embedded pointer struct is
allocated), an embedded map catches the keys no field claims, and a name
clash resolves in favour of the shallower field. A struct with an untagged
embedded field previously decoded with all inline keys dropped and did not
round-trip.
- `Marshal` re-emits arrays that mix tables with scalars: the table elements
render as inline tables inside the value array. A tree that `Parse` accepts
from such a document previously failed with
`cannot encode map[string]interface {}`.
## [1.0.0] - 2026-08-20
First stable release: a dependency-free TOML 1.0 parser and encoder for Go that
uses only the standard library and passes the entire
[toml-test](https://github.com/toml-lang/toml-test) suite
(**185 valid and 371 invalid cases, 0 failures**).
### Added
**Parsing**
- `Parse(data []byte) (map[string]any, error)`: decode a TOML document into an
untyped tree.
- Full TOML 1.0 syntax: comments; bare, quoted, and dotted keys; tables
(`[a.b]`) and arrays of tables (`[[a]]`); basic and literal strings, including
multiline (`"""` / `'''`) with escape sequences and line-ending backslash
trimming; integers in decimal, hex (`0x`), octal (`0o`), and binary (`0b`)
with `_` separators; floats with exponents and `inf` / `nan`; booleans;
offset/local date-times, dates, and times; arrays; and inline tables.
- Distinct date-time types: offset date-times decode to `time.Time`, while
`LocalDateTime`, `LocalDate`, and `LocalTime` represent the local variants,
each with a `String()` method returning the TOML-canonical rendering.
- Strict, spec-conformant validation that rejects invalid documents: leading
zeros, misplaced underscores, malformed floats and radix literals, control
characters in strings and comments, bare carriage returns, non-UTF-8 input,
out-of-range Unicode escapes, multiline strings used as keys, single-digit
date-time components, duplicate/overwriting inline-table keys, and the full
family of table redefinitions (header vs. header, header vs. dotted key,
array of tables vs. table, and dotted-key appends to defined tables).
- `SyntaxError` carrying the 1-based line of a malformed document.
**Decoding**
- `Unmarshal(data []byte, v any)`: parse and map onto a struct or
`map[string]any` via reflection, with overflow-checked numeric conversion,
nested structs, slices, and `map[string]T`.
- `toml:"name"` field tags, case-insensitive name fallback, and `toml:"-"` to
skip a field.
- `Decoder` with `DisallowUnknownFields` for strict decoding that rejects keys
without a destination field, at every struct depth.
- `Unmarshaler` interface (`UnmarshalTOML(data any) error`) for types that take
full control of their decode.
**Encoding**
- `Marshal(v any) ([]byte, error)`: encode a struct or `map[string]V` value to
a TOML 1.0 document that re-parses to an equivalent value tree.
- `Encoder` with chainable policy options: `GroupByKind` (group-by-kind layout
versus declaration order), `OmitEmptyArrays`, and `UseLiteralMultiline`.
- `Marshaler` interface (`MarshalTOML() (any, error)`) for types that need a
custom TOML shape; the returned value is encoded normally.
**Cancellation**
- `ParseContext`, `UnmarshalContext`, `MarshalContext`,
`(*Decoder).DecodeContext` and `(*Encoder).MarshalContext` honour a
`context.Context`, checked up front and every 64 statements or fields.
**Project**
- Standard library only: zero third-party modules.
- `cmd/interpres-decode`, a toml-test harness adapter (TOML on stdin, tagged
JSON on stdout).
- A runnable example, a table-driven Go test suite, and a canonical `just`
recipe set whose `gates` recipe is the definition of done.
- Hand-written CI pipelines for the test, race and release gates.
- `SECURITY.md` for private vulnerability reports.