• v2.0.0 e3dda5e115

    v2.0.0
    Test / test (push) Successful in 1m46s
    Release / gates (push) Successful in 1m36s
    Release / release (push) Successful in 39s
    Stable

    petrbalvin released this 2026-09-22 19:52:42 +00:00 | 0 commits to main since this release

    Added

    • encoding.TextMarshaler and encoding.TextUnmarshaler are honoured by
      default, with no option to switch them off. A type that implements them is
      encoded as a TOML string and decoded from one: net.IP becomes
      "192.0.2.1", and a user type with MarshalText or UnmarshalText follows.
      MarshalTOML and UnmarshalTOML still win over the text methods, and the
      four date-time types keep their bare timestamp form instead of becoming a
      quoted string. A struct type that implements the interface now encodes as a
      string where it was a table before, which is the breaking part of the change.
    • time.Duration is encoded in its canonical Go form as a TOML string,
      1h30m0s, because TOML has no duration type; the decoder reads that string
      back and still accepts a bare integer as the nanosecond count.
    • interpres-decode --encode, the adapter's other direction: it reads the
      toml-test tagged JSON from stdin and writes the TOML document it describes.
      The compliance suite now runs the encoder as well as the decoder, 214
      encoder cases against the tagged JSON of the valid corpus.
    • InlineTables(threshold): a sub-table whose single-line rendering is
      at most threshold bytes is written as an inline table instead of a header
      section, which shortens a document of small tables. An array of tables keeps
      its header form, because its inline form would re-parse as a value array.
    • Document and ParseMap: Parse now returns a *Document, which holds the
      values together with the key order, whether a table was written as an inline
      table or under a header, and the comments, with Keys, Entries, Get,
      Comments and SetComments to read and write them. ParseMap returns the
      plain map[string]any tree, the shape Parse used to give.
    • The edit pipeline on a Document: typed getters on Document and Table
      (GetString, GetInt, GetFloat, GetBool, GetArray, GetTable),
      Set and Delete that keep the surviving keys' positions and comments,
      UnmarshalDocument, which decodes the document into a typed destination
      without parsing again, and Marshal of a Document, which writes the keys
      in written order, the comments above the lines and headers they belonged
      to, and the inline tables inline.
    • OffsetDateTime, the Go type of the offset date-time kind, so that all four
      TOML date-time kinds have one of their own. Parse and Unmarshal hand it
      back where they produced a bare time.Time before, and Marshal accepts it.
      Unmarshalling into a struct field of type time.Time keeps working, because
      the plain type takes an offset date-time as it always did; code that asserts
      the tree's type, and UnmarshalTOML implementations that expect a
      time.Time, need the new type.
    • MaxNestingDepth(depth) and MaxInputSize(size) options bound the parse an
      Unmarshal performs, and every parse carries a nesting limit in any case
      (10000 levels, which no hand-written document approaches): a document that
      nests arrays or inline tables deeper used to run the stack out and is now
      rejected with a SyntaxError naming the limit.
    • Statements(r), an iterator over the top-level statements of the document
      the reader carries, in written order: key/value statements, a [table]
      header as one statement with its node, an [[array of tables]] as one
      statement per element. A caller that breaks after the statement it wanted
      reads no further ones. examples/statements shows the walk.
    • ParseAs[T](data, opts...), the generic one-line decode, and NewSchema[T](),
      which precompiles the struct schema and the interface flags for a hot path
      before the first document arrives.
    • EmitFieldComments(true) prints the comment a field's toml tag
      carries in a comment= option above the field's line or header, the
      comments a round trip through the Go type would otherwise drop. Go doc
      comments are not visible to reflection, so the tag is the channel that
      carries the text.
    • LocalTimeLocation(loc) lets a local date-time fill a plain
      time.Time destination in the location given, relabelled rather than
      shifted: 07:32 in the document is 07:32 in the zone. Without the
      option the wrapper types remain the only destinations a local kind fills.
    • The parse checks its context inside a value as well as between statements:
      an array, an inline table and a multi-line string check every 64 elements
      or lines, so one huge value cannot hold the parse past its cancellation.
    • OrderedMap, the string-keyed table that remembers the order its keys were
      set in: decoding into one fills it in the order the document wrote the
      keys, and Marshal writes one back in that order, where a map carries no
      order on decode and sorts on encode. It works as a decode target on its
      own, in a struct field, and as the element of an array of tables; its
      values are untyped, so a nested table stays a map[string]any.
    • Unmarshal(data, v, opts...) and the other entries take variadic options,
      the shape encoding/json/v2 uses: RejectUnknownFields,
      NumbersAsLiterals, MaxNestingDepth, MaxInputSize,
      LocalTimeLocation. MarshalWrite(w, v, opts...) and
      UnmarshalRead(r, v, opts...) are the streaming forms.
    • Marshal carries a nesting limit of 10000 levels, the parser's own figure:
      cyclic data, which used to run the stack out, is now rejected with an error
      that names the limit and the path it was met at.
    • The toml tag gained the required option: a field tagged
      toml:"host,required" makes the decode fail with
      missing required key "host" when the document carries no key that
      resolves to it. The option shapes decoding only, and the encoder ignores
      it.
    • UnmarshalerContext, the custom-decode interface that hands the decode's
      context to the method, UnmarshalTOMLContext(ctx, data). It wins over
      UnmarshalTOML when a type implements both, so a long custom decode can
      abort on cancellation; a non-cancellable entry point hands in
      context.Background, never nil.
    • A TOML array decodes into a Go fixed-size array, [N]T, where only a slice
      was accepted before; the encoder could already encode one. A length mismatch
      is an error wrapped with the key path.
    • MarshalAppend(buf, v) appends the TOML encoding of v to buf and returns
      the extended buffer, the shape json.MarshalAppend has.
    • interpres-decode --version prints the binary's version, the module version
      the toolchain recorded, so a release-built binary names its own tag.
    • interpres-decode --json prints plain indented JSON instead of the tagged
      form, the shape for people and diffs, with the date-time wrappers in their
      TOML form.
    • interpres-decode --validate walks a named directory for .toml files and
      closes the sweep with a summary naming how many documents were checked and
      how many were invalid; single files stay quiet on success as before.
    • interpres-decode --struct infers a Go struct definition from a document:
      one field per key in written order, nested tables as nested struct types,
      an array of tables as a slice. The printed type compiles and decodes the
      document it came from.
    • interpres-decode --schema TYPE file.go writes a TOML template for the
      named struct type of a Go source, the comment= tag option printed as a
      comment and the default= option as the value. It is the inverse of
      --struct.
    • ParseFile(path) reads the file and parses it into a Document, with the
      file name at the front of every error it returns, read failure and parse
      failure alike. Valid(data) reports whether a document parses, nil on
      success and the parse error on failure, the library call the --validate
      mode of interpres-decode is built on.
    • SyntaxError carries the byte Offset the scan stopped at and the 1-based
      Column on the line, beside the line it always had, and SourceLine(src)
      renders that line with a caret under the position, for messages shown under
      the input. An input that is not valid UTF-8 names the offset of the first
      invalid byte in its message. The new fields are additive: a SyntaxError
      built from a line and a message alone is unchanged.
    • NumbersAsLiterals(true) decodes the integers and floats of the document into
      Number, which carries the literal the document wrote, so 0x1f, 1_000,
      +1.0 and inf survive a round trip with their spelling instead of the
      normalised 31, 1000 and 1.0. Typed destinations take the evaluated
      value as before, a Number field takes the literal, and Marshal writes a
      Number back as its bare literal, rejecting one that is not a valid TOML
      number.

    Changed

    • The stateful Decoder and Encoder of 1.x are replaced by variadic
      options on the entries, the shape encoding/json/v2 uses: Layout(kind)
      with LayoutKindGrouped or LayoutKindDeclaration, OmitEmptyArrays,
      LiteralMultiline(threshold), InlineTables(threshold),
      EmitFieldComments, RejectUnknownFields, NumbersAsLiterals,
      MaxNestingDepth, MaxInputSize, LocalTimeLocation.
    • DecodeError and EncodeError carry one Path type, a list of segments
      ("items", "[0]", "weight") with a String() rendering the TOML
      notation, items[0].weight. The decode error used to hold a bare
      []string, the encode error a plain string. Both messages render the same
      way now, interpres: items[0].weight: ..., with one interpres: prefix
      where the composition used to double it.
    • omitempty follows the encoding/json semantics: the field is skipped when
      it holds an empty string, a zero number, false, a nil pointer or
      interface, or a nil or empty slice, array or map. In 1.x the option covered
      only the collections.
    • The toml tag gained the inline option: a struct or map field tagged
      toml:"retry,inline" emits as retry = {…} instead of a header section,
      whatever its size, a named embedded struct included. Forcing it on an array
      of tables is an error, because the inline form would re-parse as a value
      array and change the value's Go type.
    • The output takes the TOML 1.1 form. A date-time writes its seconds only when
      the value carries them and drops the trailing zeros of a fractional second,
      so 07:32:00 is written 07:32 and half a second as 00.5. Both are the
      same value, and a document written without seconds now comes back without
      them. LocalDateTime.String(), LocalTime.String() and the offset date-time
      rendering follow the same rule.
    • An inline table that would pass the hundredth column is written across lines
      with a trailing comma and one tab of indentation per nesting level, the shape
      TOML 1.1 allows an inline table to take.
    • MarshalTOML reaches every array element and every field, whatever the Go
      kind, and its result is normalised like any other value: an element rendering
      itself as a table keeps the [[header]] form, one rendering itself as a
      scalar turns the array into a value array, and the method runs once per
      element. It is found on the addressable pointer as well, so a
      pointer-receiver MarshalTOML is called for a field or an element, exactly
      as MarshalText is.
    • TOML 1.1 is the acceptance contract, and TOML 1.0 is not. The compliance
      suite runs the 1.1 corpus alone, and the promise that every 1.0 document
      parses exactly as before is withdrawn. Nothing that parses today stops
      parsing: the 1.0 valid corpus still passes in full. The documents whose
      verdict changes are the ones 1.1 relaxed, such as the \xHH escape
      sequences 1.0 rejected.
    • The module path carries the /v2 suffix the Go toolchain requires of
      every major version 2 module: imports change to
      sourcedock.dev/petrbalvin/interpres/v2.
    • Input that is not valid UTF-8 is now rejected where the parser's scan
      meets the invalid byte, with a SyntaxError naming that line, instead of
      a whole-input check that always reported line 1. Invalid input is still
      rejected; the reported location is now the byte's own.
      Performance
    • Struct destinations decode directly: for a type the direct skeleton can
      model, the parser resolves tables and keys against the struct schema while
      the document scans and no intermediate value tree is kept. The strict
      decode of the representative document drops from 168 to 160 allocations
      per call against the tree path in the same process, and the 2000-element
      document reaches allocation parity; every document the skeleton cannot
      model falls back to the tree path and its exact error contracts. A
      differential fuzz target decodes every generated document both ways.
    • Marshal writes plain scalars and typed scalar arrays straight from their
      reflect cells instead of boxing them into interface values first, and skips
      the per-element resolution for arrays that can never take the [[header]]
      form. The representative document now costs 130 allocations per call
      instead of 141, the long array-of-tables document 55 915 instead of
      63 660, with byte-identical output.
    • Parsing is faster than in 1.1.0 while carrying the new document layer:
      the suite's representative document decodes at about 79 MB/s with 104
      allocations per call, and the long array-of-tables document at about
      106 MB/s against 56 MB/s in 1.1.0, with allocations on that document
      halved from 67 664 to 31 765. Date-time tokens are validated by a byte
      scan instead of regular expressions, repeated keys share one string
      across array-of-tables elements, and per-statement buffers are reused.
    • Typed decoding is 12 percent faster than in 1.1.0 on the representative
      document (9792 ns against 11 147 ns) with 24 percent fewer allocations
      (167 against 220); interface lookups resolve through a cached per-type
      flag set instead of boxing every value into an interface to ask.
    • Marshal runs at the 1.1.0 speed while emitting the new TOML 1.1 output
      form, at half the bytes per operation (6170 against 11 348 on the
      representative document), and writes through a pooled output buffer with
      a 1 MiB retention cap; repeated marshals keep the live heap flat.
    • Two benchmarks measure the shapes that drove the work:
      BenchmarkStrictDecodeLong and BenchmarkMarshalLong run the 2000-entry
      document at about 3.8 ms and 3.4 ms per call, at 63 772 and 63 660
      allocations.

    Fixed

    • An offset date-time written with the +00:00 offset kept the anonymous
      location time.Parse invents for it, so a round trip through the tree and
      Marshal, which writes a zero offset as Z, changed the value's
      reflection-visible location. The zero offset normalises to time.UTC at
      parse, and the tree is stable across the round trip.
    • Decoding into a defined type whose underlying kind is string or bool, such
      as type Name string, panicked instead of storing the value, because a
      value of the predeclared type is not assignable to a defined type and the
      decoder assigned it without a conversion.
    • A top-level value the encoder could not normalise reported its path with
      a leading dot, interpres: .port: ...; the message now reads
      interpres: port: ..., the shape EncodeError.Path already used.
    • A token shaped like a date-time with a component out of range, such as an
      hour of 24 or a February the 30th, fell through to the number decoder and
      failed with the number complaint invalid character "-" in number; it now
      fails as the date-time it visibly is, invalid date-time "...".
    • A time.Time or OffsetDateTime whose zone offset is not a whole number
      of minutes wrote only the minutes, silently shifting the instant by the
      seconds dropped; the encoder now refuses such an offset, which TOML has no
      form for, instead of corrupting the value.
    • An empty array of tables over pointer elements, []*T{}, emitted as
      key = [] while its value form was omitted; it is omitted too now, the
      rule TOML forces, because an empty [[a]] has no valid form.
    • Two lenient grammar edges are closed: a sign in a \u or \U escape,
      which is not a hex digit, is rejected instead of evaluating, and a bare
      carriage return right after a multi-line string's opening delimiter is
      the bare-CR error instead of a newline trimmed silently.

    Migration from 1.x

    The module path. 2.0 lives at sourcedock.dev/petrbalvin/interpres/v2,
    the suffix the Go toolchain requires of every major version 2 module. Change
    every import and go get line:

    go get sourcedock.dev/petrbalvin/interpres/v2
    

    TOML 1.1 only. The acceptance contract is the TOML 1.1 corpus, and the
    promise that every 1.0 document parses exactly as before is withdrawn.
    Documents whose verdict changes are the ones 1.1 relaxed: \e and \xHH
    escapes, times without seconds, multi-line inline tables with comments and a
    trailing comma. Nothing that parsed in 1.x stops parsing, because the 1.1
    grammar contains the 1.0 one.
    The output takes the 1.1 form. A date-time writes seconds only when the
    value carries them, a fraction drops its trailing zeros, and a long inline
    table breaks across lines. A document written from the same value can come
    out shorter; it re-parses to the same value.
    Text methods on by default. A type implementing
    encoding.TextMarshaler or encoding.TextUnmarshaler now takes the text
    path with no option to switch it off. A struct that implemented the
    interface encodes as a string where it was a table before. MarshalTOML and
    UnmarshalTOML still win.
    One Go type per date-time kind. Offset date-times hand back
    OffsetDateTime, not a bare time.Time. Code that type-asserts the tree or
    expects time.Time inside UnmarshalTOML needs the new wrapper; a
    destination field of type time.Time keeps working.
    The document carries what the map could not. Parse returns a
    *Document with the key order, the inline distinction and the comments;
    ParseMap gives the plain map[string]any tree the old Parse returned.
    The document is writable, and Marshal writes it back with its comments.
    Options instead of Decoder and Encoder. The stateful types of 1.x are
    gone; the entries take variadic options, the shape encoding/json/v2 uses.
    NewDecoder().DisallowUnknownFields().Decode(data, &cfg) becomes
    Unmarshal(data, &cfg, RejectUnknownFields(true)), and the encoder
    methods become options: Layout(LayoutKindDeclaration) replaces
    GroupByKind(false), LiteralMultiline replaces UseLiteralMultiline.
    Tag options. omitempty follows encoding/json: it now also drops empty
    strings, zero numbers, false, nil pointers and nil interfaces. required
    demands a key at decode. inline forces the inline table form at encode.
    comment=text carries a comment EmitFieldComments prints.
    Errors. DecodeError.Path is a Path (segments with a String()
    renderer), EncodeError.Path the same type instead of a plain string, and
    both messages render interpres: server.ports[2]: ... with one prefix.
    SyntaxError gained Offset, Column and SourceLine. Decode errors into
    a Number-carrying tree and the fixed-size array decode are new shapes a
    match on the old messages would not see.
    Decoding shapes. map[string]any values merge into a non-empty map
    destination; untagged embedded maps beyond the first stay empty; numbers can
    stay literals under NumbersAsLiterals; local date-times can decode into
    time.Time under LocalTimeLocation. All three are opt-in or additive
    except where noted above.

    Downloads