12 Commits
Author SHA1 Message Date
petrbalvin 72f8b21ac4 feat: parse into a Document that keeps order and comments
Test / test (push) Successful in 1m34s
Assisted-by: DeepSeek V4.1 Flash
2026-09-20 10:40:57 +02:00
petrbalvin a6e3e3fe31 feat: add OffsetDateTime, nesting limits and uniform Marshaler dispatch
Test / test (push) Successful in 2m18s
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 19:36:24 +02:00
petrbalvin b695b69768 docs(api): use an inline table sample that really breaks
Test / test (push) Successful in 1m33s
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 12:18:57 +02:00
petrbalvin 959eaba4b0 feat(encode): add the InlineTables option
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 12:18:30 +02:00
petrbalvin 8f85bb68fa feat(encode): write the TOML 1.1 output form
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 12:18:18 +02:00
petrbalvin bccaf087c8 feat(cmd): add the encoder mode to the toml-test adapter
Test / test (push) Successful in 1m33s
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 11:38:19 +02:00
petrbalvin 0149a5b4d1 docs(encoder): correct the multi-line string claim
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 11:38:17 +02:00
petrbalvin 815141440e feat: honour TextMarshaler and TextUnmarshaler by default
Test / test (push) Successful in 1m35s
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 02:41:09 +02:00
petrbalvin 9023784da3 fix(decode): decode into a defined string or bool type
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 02:40:51 +02:00
petrbalvin 942c4b1489 docs(security): list the newest release as supported
Test / test (push) Successful in 1m33s
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 02:24:24 +02:00
petrbalvin 8f0eae6604 docs: drop the TOML 1.0 compatibility promise
Assisted-by: DeepSeek V4.1 Flash
2026-09-19 02:24:24 +02:00
petrbalvin 1c7329aeea build: move the module path to /v2
Test / test (push) Successful in 1m32s
Assisted-by: GLM 5.3 Flash
2026-09-19 00:14:39 +02:00
26 changed files with 3230 additions and 310 deletions
+5 -3
View File
@@ -105,6 +105,8 @@ jobs:
run: go build -o bin/interpres-decode ./cmd/interpres-decode
- name: Compliance suite
# interpres implements TOML 1.0 and 1.1; the mode is pinned so an upstream
# default change cannot silently move the corpus.
run: bin/toml-test test -decoder=bin/interpres-decode -toml=1.1
# interpres implements TOML 1.1, and the suite runs both directions: the decoder
# on the valid and invalid corpora, the encoder on the tagged JSON of the valid
# one. The mode is pinned so an upstream default change cannot silently move the
# corpus.
run: bin/toml-test test -decoder=bin/interpres-decode -encoder='bin/interpres-decode -encode' -toml=1.1
+72 -1
View File
@@ -9,7 +9,78 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
-
- `encoding.TextMarshaler` and `encoding.TextUnmarshaler` are honoured by
default, with no option to switch them off. A type that implements them is
encoded as a TOML string and decoded from one: `net.IP` becomes
`"192.0.2.1"`, and a user type with `MarshalText` or `UnmarshalText` follows.
`MarshalTOML` and `UnmarshalTOML` still win over the text methods, and the
four date-time types keep their bare timestamp form instead of becoming a
quoted string. A struct type that implements the interface now encodes as a
string where it was a table before, which is the breaking part of the change.
- `time.Duration` is encoded in its canonical Go form as a TOML string,
`1h30m0s`, because TOML has no duration type; the decoder reads that string
back and still accepts a bare integer as the nanosecond count.
- `interpres-decode -encode`, the adapter's other direction: it reads the
toml-test tagged JSON from stdin and writes the TOML document it describes.
The compliance suite now runs the encoder as well as the decoder, 214
encoder cases against the tagged JSON of the valid corpus.
- `Encoder.InlineTables(threshold)`: a sub-table whose single-line rendering is
at most `threshold` bytes is written as an inline table instead of a header
section, which shortens a document of small tables. An array of tables keeps
its header form, because its inline form would re-parse as a value array.
- `Document` and `ParseMap`: `Parse` now returns a `*Document`, which holds the
values together with the key order, whether a table was written as an inline
table or under a header, and the comments, with `Keys`, `Entries`, `Get`,
`Comments` and `SetComments` to read and write them. `ParseMap` returns the
plain `map[string]any` tree, the shape `Parse` used to give. `Marshal` does
not accept a `Document`; it writes values, so `doc.Map()` is the way through.
- `OffsetDateTime`, the Go type of the offset date-time kind, so that all four
TOML date-time kinds have one of their own. `Parse` and `Unmarshal` hand it
back where they produced a bare `time.Time` before, and `Marshal` accepts it.
Unmarshalling into a struct field of type `time.Time` keeps working, because
the plain type takes an offset date-time as it always did; code that asserts
the tree's type, and `UnmarshalTOML` implementations that expect a
`time.Time`, need the new type.
- `Decoder.MaxDepth(depth)` and `Decoder.MaxInputSize(size)` bound the parse a
`Decode` performs, and every parse carries a nesting limit in any case
(10000 levels, which no hand-written document approaches): a document that
nests arrays or inline tables deeper used to run the stack out and is now
rejected with a `SyntaxError` naming the limit.
### Changed
- The output takes the TOML 1.1 form. A date-time writes its seconds only when
the value carries them and drops the trailing zeros of a fractional second,
so `07:32:00` is written `07:32` and half a second as `00.5`. Both are the
same value, and a document written without seconds now comes back without
them. `LocalDateTime.String()`, `LocalTime.String()` and the offset date-time
rendering follow the same rule.
- An inline table that would pass the hundredth column is written across lines
with a trailing comma and one tab of indentation per nesting level, the shape
TOML 1.1 allows an inline table to take.
- `MarshalTOML` reaches every array element and every field, whatever the Go
kind, and its result is normalised like any other value: an element rendering
itself as a table keeps the `[[header]]` form, one rendering itself as a
scalar turns the array into a value array, and the method runs once per
element. It is found on the addressable pointer as well, so a
pointer-receiver `MarshalTOML` is called for a field or an element, exactly
as `MarshalText` is.
- TOML 1.1 is the acceptance contract, and TOML 1.0 is not. The compliance
suite runs the 1.1 corpus alone, and the promise that every 1.0 document
parses exactly as before is withdrawn. Nothing that parses today stops
parsing: the 1.0 valid corpus still passes in full. The documents whose
verdict changes are the ones 1.1 relaxed, such as the `\xHH` escape
sequences 1.0 rejected.
- The module path carries the /v2 suffix the Go toolchain requires of
every major version 2 module: imports change to
`sourcedock.dev/petrbalvin/interpres/v2`.
### Fixed
- Decoding into a defined type whose underlying kind is string or bool, such
as `type Name string`, panicked instead of storing the value, because a
value of the predeclared type is not assignable to a defined type and the
decoder assigned it without a conversion.
## [1.1.0] - 2026-09-18
+2 -2
View File
@@ -45,8 +45,8 @@ just test
formatting pass are three commits, never one.
4. Record every user-visible change in `CHANGELOG.md` under `## [development]`.
5. Add or update tests. Coverage stays at 80 percent or more; it is a hard
gate. Parser and decoder changes must also keep the toml-test suite at zero
failures, checked with `just toml-test`.
gate. Parser, decoder and encoder changes must also keep both directions of
the toml-test suite at zero failures, checked with `just toml-test`.
6. Update the documentation when the public API, the configuration or the
behaviour changes; the documents move in the same commit as the behaviour
they describe.
+17 -10
View File
@@ -1,14 +1,14 @@
# interpres
A TOML 1.0 and 1.1 parser and encoder for Go, written with the standard
library alone. `interpres` (Latin for *interpreter*) gives zero-dependency
A TOML 1.1 parser and encoder for Go, written with the standard library
alone. `interpres` (Latin for *interpreter*) gives zero-dependency
programs an `encoding/json`-style API for reading and writing TOML, and passes
the entire official [toml-test](https://github.com/toml-lang/toml-test) suite:
214 valid and 467 invalid cases, zero failures.
214 valid, 467 invalid and 214 encoder cases, zero failures.
## Features
- **Full TOML 1.0 and 1.1**: bare, quoted and dotted keys; tables and arrays of
- **Full TOML 1.1**: bare, quoted and dotted keys; tables and arrays of
tables; basic and literal strings including multiline, with the 1.1 `\e` and
`\xHH` escapes; integers in the four radixes with `_` separators; floats with
exponents, `inf` and `nan`; booleans; the four date-time kinds, seconds
@@ -18,18 +18,24 @@ the entire official [toml-test](https://github.com/toml-lang/toml-test) suite:
- **Strict decoding**: `NewDecoder().DisallowUnknownFields()` rejects keys that
match no destination field, at every struct depth.
- **Custom types**: `Marshaler` and `Unmarshaler` let a type control its own
TOML representation in both directions.
TOML representation in both directions, and `encoding.TextMarshaler` and
`TextUnmarshaler` are honoured by default, so `net.IP`, `time.Duration` and
user types with text methods need no configuration.
- **Cancellation**: every entry point has a `*Context` sibling that honours a
`context.Context`.
- **Ordered documents**: `Parse` gives a `*Document` that keeps the key order,
tells an inline table from a header one, and carries the comments; `ParseMap`
gives the plain `map[string]any` tree.
- **Configurable emission**: `Encoder` options for declaration-order output,
omitting empty arrays, and literal multiline strings.
omitting empty arrays, literal multiline strings, and inlining small
sub-tables.
## Install
As a library:
```sh
go get sourcedock.dev/petrbalvin/interpres
go get sourcedock.dev/petrbalvin/interpres/v2
```
Requires Go 1.27.1 or newer. The module imports only the standard library.
@@ -64,8 +70,9 @@ err := interpres.Unmarshal(data, &cfg)
```
Fields match by the `toml:"name"` tag, or by the lower-cased field name when no
tag is present; `toml:"-"` skips a field. `Parse` returns the untyped
`map[string]any` tree instead, and `UnmarshalContext` accepts a context.
tag is present; `toml:"-"` skips a field. `Parse` returns a `*Document` that
also carries the key order and the comments, `ParseMap` returns the plain
`map[string]any` tree, and `UnmarshalContext` accepts a context.
### Encode from a struct
@@ -161,7 +168,7 @@ See [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the full workflow, and
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/API.md](docs/API.md): the API reference, decoding and encoding rules
- [docs/CLI.md](docs/CLI.md): the interpres-decode toml-test adapter and validator
- [docs/CLI.md](docs/CLI.md): the interpres-decode adapter and validator
## Licence
+1 -1
View File
@@ -7,7 +7,7 @@ releases do not receive them.
| Version | Supported |
|---|---|
| 1.0.0 | yes |
| 1.1.0 | yes |
| older releases | no |
## Reporting a vulnerability
+3 -3
View File
@@ -87,14 +87,14 @@ func BenchmarkParse(b *testing.B) {
b.ReportAllocs()
b.SetBytes(int64(len(benchDoc)))
for b.Loop() {
if _, err := Parse(benchDoc); err != nil {
if _, err := ParseMap(benchDoc); err != nil {
b.Fatal(err)
}
}
}
func BenchmarkMarshal(b *testing.B) {
tree, err := Parse(benchDoc)
tree, err := ParseMap(benchDoc)
if err != nil {
b.Fatal(err)
}
@@ -122,7 +122,7 @@ func BenchmarkParseLong(b *testing.B) {
b.ReportAllocs()
b.SetBytes(int64(len(longDoc)))
for b.Loop() {
if _, err := Parse(longDoc); err != nil {
if _, err := ParseMap(longDoc); err != nil {
b.Fatal(err)
}
}
+203 -7
View File
@@ -4,14 +4,16 @@
// Command interpres-decode is the toml-test harness adapter and a TOML
// validator. Without flags it reads a TOML document from standard input and
// writes the toml-test "tagged JSON" representation to standard output. With
// -validate it checks the named documents, or standard input when none are
// named, and exits non-zero on the first invalid one:
// -encode it is the reverse: it reads tagged JSON and writes the TOML document
// it describes. With -validate it checks the named documents, or standard
// input when none are named, and exits non-zero on the first invalid one:
//
// interpres-decode -validate config.toml
// interpres-decode -encode < case.json
//
// Run the official suite against the adapter with:
// Run the official suite in both directions against the adapter with:
//
// toml-test ./interpres-decode
// toml-test test -decoder=./interpres-decode -encoder='./interpres-decode -encode'
package main
import (
@@ -25,7 +27,7 @@ import (
"strconv"
"time"
"sourcedock.dev/petrbalvin/interpres"
"sourcedock.dev/petrbalvin/interpres/v2"
)
func main() {
@@ -39,12 +41,17 @@ func Run(args []string, stdin io.Reader, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("interpres-decode", flag.ContinueOnError)
fs.SetOutput(stderr)
validate := fs.Bool("validate", false, "validate the documents instead of emitting tagged JSON")
encode := fs.Bool("encode", false, "read tagged JSON from stdin and write TOML instead")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
if *validate && *encode {
fmt.Fprintln(stderr, "interpres-decode: -validate and -encode cannot be combined")
return 2
}
if *validate {
return validatePaths(fs.Args(), stdin, stderr)
}
@@ -52,12 +59,15 @@ func Run(args []string, stdin io.Reader, stdout, stderr io.Writer) int {
fmt.Fprintln(stderr, "interpres-decode: the adapter mode takes no arguments; name files with -validate")
return 2
}
if *encode {
return encodeJSON(stdin, stdout, stderr)
}
data, err := io.ReadAll(stdin)
if err != nil {
fmt.Fprintln(stderr, "read stdin:", err)
return 2
}
tree, err := interpres.Parse(data)
tree, err := interpres.ParseMap(data)
if err != nil {
fmt.Fprintln(stderr, err)
return 1
@@ -98,7 +108,7 @@ func validatePaths(paths []string, stdin io.Reader, stderr io.Writer) int {
fmt.Fprintf(stderr, "interpres-decode: %s: %v\n", name, err)
return 2
}
if _, err := interpres.Parse(data); err != nil {
if _, err := interpres.ParseMap(data); err != nil {
fmt.Fprintf(stderr, "%s: %v\n", name, err)
valid = false
}
@@ -109,6 +119,190 @@ func validatePaths(paths []string, stdin io.Reader, stderr io.Writer) int {
return 0
}
// encodeJSON reads a toml-test tagged JSON description from standard input and
// writes the TOML document it describes to standard output.
func encodeJSON(stdin io.Reader, stdout, stderr io.Writer) int {
data, err := io.ReadAll(stdin)
if err != nil {
fmt.Fprintln(stderr, "read stdin:", err)
return 2
}
var desc any
if err := json.Unmarshal(data, &desc); err != nil {
fmt.Fprintln(stderr, "decode JSON:", err)
return 2
}
tree, err := untag(desc)
if err != nil {
fmt.Fprintln(stderr, err)
return 2
}
doc, ok := tree.(map[string]any)
if !ok {
fmt.Fprintln(stderr, "interpres-decode: the description must be a JSON object at the top level")
return 2
}
out, err := interpres.Marshal(doc)
if err != nil {
fmt.Fprintln(stderr, err)
return 2
}
if _, err := stdout.Write(out); err != nil {
fmt.Fprintln(stderr, "write stdout:", err)
return 2
}
return 0
}
// untag converts a toml-test JSON description into the value tree Marshal
// expects: a JSON object becomes a map[string]any, a JSON array becomes a
// []any, and an object carrying exactly the keys "type" and "value" becomes
// the Go value for that TOML type.
func untag(v any) (any, error) {
switch x := v.(type) {
case map[string]any:
if typ, val, ok := taggedValue(x); ok {
return decodeTagged(typ, val)
}
out := make(map[string]any, len(x))
for k, e := range x {
u, err := untag(e)
if err != nil {
return nil, fmt.Errorf("%s: %w", k, err)
}
out[k] = u
}
return out, nil
case []any:
out := make([]any, len(x))
for i, e := range x {
u, err := untag(e)
if err != nil {
return nil, fmt.Errorf("[%d]: %w", i, err)
}
out[i] = u
}
return asTables(out), nil
default:
return nil, fmt.Errorf("unsupported JSON value %T", v)
}
}
// asTables returns the elements as a []map[string]any when there is at least
// one and every element is a table, the shape the encoder renders as an array
// of tables. The tagged JSON cannot tell an array of tables from a value array
// of inline tables, and both parse back to the same value, so the header form
// is chosen because it is the one the encoder otherwise never exercises. An
// empty array stays a []any, because TOML has no empty array of tables.
func asTables(items []any) any {
if len(items) == 0 {
return items
}
tbls := make([]map[string]any, len(items))
for i, e := range items {
tbl, ok := e.(map[string]any)
if !ok {
return items
}
tbls[i] = tbl
}
return tbls
}
// taggedValue reports whether m is a toml-test value object: a JSON object of
// exactly the two string keys "type" and "value", carrying a type this adapter
// knows. Any other object is a table.
func taggedValue(m map[string]any) (typ, val string, ok bool) {
if len(m) != 2 {
return "", "", false
}
ts, ok := m["type"].(string)
if !ok || !knownType(ts) {
return "", "", false
}
vs, ok := m["value"].(string)
if !ok {
return "", "", false
}
return ts, vs, true
}
func knownType(typ string) bool {
switch typ {
case "string", "integer", "float", "bool",
"datetime", "datetime-local", "date-local", "time-local":
return true
}
return false
}
// decodeTagged returns the Go value for one tagged JSON value. Every type but
// string is parsed by the library itself, so the adapter and the library agree
// on what an integer, a float or a date-time is.
func decodeTagged(typ, val string) (any, error) {
if typ == "string" {
return val, nil
}
v, err := parseAtom(val)
if err != nil {
return nil, fmt.Errorf("%s %q: %w", typ, val, err)
}
// A float with no fractional part and no exponent is described by a bare
// integer literal, so here the tag decides and not the literal.
if n, ok := v.(int64); ok && typ == "float" {
return float64(n), nil
}
if !typeMatches(typ, v) {
return nil, fmt.Errorf("%s %q parsed as %T", typ, val, v)
}
return v, nil
}
// parseAtom parses one bare TOML value, by handing `v = <val>` to the library's
// parser and requiring the result to hold exactly that one statement, so a
// value carrying a newline or a comment cannot smuggle a second one in.
func parseAtom(val string) (any, error) {
tree, err := interpres.ParseMap([]byte("v = " + val + "\n"))
if err != nil {
return nil, err
}
if len(tree) != 1 {
return nil, errors.New("not a single bare value")
}
return tree["v"], nil
}
// typeMatches reports whether v is the Go value the tagged type names.
func typeMatches(typ string, v any) bool {
switch typ {
case "integer":
_, ok := v.(int64)
return ok
case "float":
_, ok := v.(float64)
return ok
case "bool":
_, ok := v.(bool)
return ok
case "datetime":
switch v.(type) {
case time.Time, interpres.OffsetDateTime:
return true
}
return false
case "datetime-local":
_, ok := v.(interpres.LocalDateTime)
return ok
case "date-local":
_, ok := v.(interpres.LocalDate)
return ok
case "time-local":
_, ok := v.(interpres.LocalTime)
return ok
}
return false
}
// tag converts an interpres value into its toml-test tagged-JSON form. Tables
// become JSON objects and arrays become JSON arrays; scalars are wrapped in a
// {"type", "value"} object. An error is returned for value types the encoder
@@ -155,6 +349,8 @@ func tag(v any) (any, error) {
return tagged("float", formatFloat(x)), nil
case time.Time:
return tagged("datetime", x.Format(time.RFC3339Nano)), nil
case interpres.OffsetDateTime:
return tagged("datetime", x.Format(time.RFC3339Nano)), nil
case interpres.LocalDateTime:
return tagged("datetime-local", x.Format("2006-01-02T15:04:05.999999999")), nil
case interpres.LocalDate:
+164 -1
View File
@@ -8,11 +8,12 @@ import (
"encoding/json"
"errors"
"os"
"reflect"
"strings"
"testing"
"time"
"sourcedock.dev/petrbalvin/interpres"
"sourcedock.dev/petrbalvin/interpres/v2"
)
func TestRunParsesValidTOML(t *testing.T) {
@@ -283,3 +284,165 @@ func TestUnknownFlagReturnsTwo(t *testing.T) {
t.Fatalf("Run returned %d, want 2; stderr = %q", code, stderr.String())
}
}
// --- encoder mode ----------------------------------------------------------
func TestRunEncoderScalars(t *testing.T) {
in := `{
"s": {"type": "string", "value": "quote \" and backslash \\"},
"nl": {"type": "string", "value": "line1\nline2"},
"i": {"type": "integer", "value": "-9223372036854775808"},
"g": {"type": "float", "value": "1.5"},
"f": {"type": "float", "value": "inf"},
"b": {"type": "bool", "value": "false"},
"dt": {"type": "datetime", "value": "1979-05-27T07:32:00-07:00"},
"ldt": {"type": "datetime-local", "value": "1979-05-27T07:32:00"},
"ld": {"type": "date-local", "value": "1979-05-27"},
"lt": {"type": "time-local", "value": "07:32:00.999"}
}
`
var stdout, stderr bytes.Buffer
code := Run([]string{"-encode"}, strings.NewReader(in), &stdout, &stderr)
if code != 0 {
t.Fatalf("Run returned %d, stderr = %q", code, stderr.String())
}
want := "b = false\n" +
"dt = 1979-05-27T07:32-07:00\n" +
"f = inf\n" +
"g = 1.5\n" +
"i = -9223372036854775808\n" +
"ld = 1979-05-27\n" +
"ldt = 1979-05-27T07:32\n" +
"lt = 07:32:00.999\n" +
"nl = \"line1\\nline2\"\n" +
"s = \"quote \\\" and backslash \\\\\"\n"
if stdout.String() != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", stdout.String(), want)
}
}
func TestRunEncoderNested(t *testing.T) {
in := `{
"tbl": {"x": {"type": "bool", "value": "true"},
"sub": {"y": {"type": "integer", "value": "1"}}},
"items": [{"n": {"type": "string", "value": "a"}},
{"n": {"type": "string", "value": "b"}}],
"list": [{"type": "integer", "value": "1"}, {"type": "string", "value": "two"}],
"emptyTbl": {},
"emptyArr": []
}
`
var stdout, stderr bytes.Buffer
code := Run([]string{"-encode"}, strings.NewReader(in), &stdout, &stderr)
if code != 0 {
t.Fatalf("Run returned %d, stderr = %q", code, stderr.String())
}
want := "emptyArr = []\n" +
"list = [1, \"two\"]\n" +
"\n[emptyTbl]\n" +
"\n[tbl]\nx = true\n" +
"\n[tbl.sub]\ny = 1\n" +
"\n[[items]]\nn = \"a\"\n" +
"\n[[items]]\nn = \"b\"\n"
if stdout.String() != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", stdout.String(), want)
}
}
func TestRunEncoderFloatTagDecides(t *testing.T) {
// A float with no fraction is described by a bare integer literal, so the
// tag decides the type; the output must stay a float.
var stdout, stderr bytes.Buffer
in := `{"whole": {"type": "float", "value": "1"}, "exp": {"type": "float", "value": "5e+22"}}`
code := Run([]string{"-encode"}, strings.NewReader(in), &stdout, &stderr)
if code != 0 {
t.Fatalf("Run returned %d, stderr = %q", code, stderr.String())
}
if want := "exp = 5e+22\nwhole = 1.0\n"; stdout.String() != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", stdout.String(), want)
}
}
func TestRunEncoderRejectsBadInput(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{"not-json", "not json", "decode JSON"},
{"top-level-array", `[{"type": "integer", "value": "1"}]`, "must be a JSON object"},
{"untagged-scalar", `{"x": 1}`, "unsupported JSON value"},
{"literal-mismatch", `{"x": {"type": "integer", "value": "1.5"}}`, "parsed as float64"},
{"offset-for-local", `{"x": {"type": "datetime-local", "value": "1979-05-27T07:32:00Z"}}`, "parsed as interpres.OffsetDateTime"},
{"bad-literal", `{"x": {"type": "date-local", "value": "nope"}}`, "date-local"},
{"smuggled-statement", `{"x": {"type": "integer", "value": "1\nx = 2"}}`, "not a single bare value"},
}
for _, c := range cases {
var stdout, stderr bytes.Buffer
code := Run([]string{"-encode"}, strings.NewReader(c.in), &stdout, &stderr)
if code != 2 {
t.Errorf("%s: Run returned %d, want 2; stderr = %q", c.name, code, stderr.String())
continue
}
if !strings.Contains(stderr.String(), c.want) {
t.Errorf("%s: stderr = %q, want it to mention %q", c.name, stderr.String(), c.want)
}
if stdout.Len() != 0 {
t.Errorf("%s: stdout should be empty, got %q", c.name, stdout.String())
}
}
}
func TestRunEncoderFlagConflicts(t *testing.T) {
var stdout, stderr bytes.Buffer
if code := Run([]string{"-encode", "-validate"}, strings.NewReader(""), &stdout, &stderr); code != 2 {
t.Errorf("Run returned %d, want 2 for the two modes together", code)
}
if !strings.Contains(stderr.String(), "cannot be combined") {
t.Errorf("stderr = %q, want it to explain the conflict", stderr.String())
}
stdout.Reset()
stderr.Reset()
if code := Run([]string{"-encode", "file.json"}, strings.NewReader(""), &stdout, &stderr); code != 2 {
t.Errorf("Run returned %d, want 2 for an argument", code)
}
}
func TestEncodeAfterDecodeRoundTrip(t *testing.T) {
doc := `title = "x"
flt = 1.5
whole = 7.0
big = 9223372036854775807
when = 1979-05-27T07:32:00-07:00
day = 1979-05-27
clock = 07:32:00.999
list = [1, "two"]
multi = "a\nb"
[tbl]
x = true
[[items]]
n = "a"
`
var tagged, stderr bytes.Buffer
if code := Run(nil, strings.NewReader(doc), &tagged, &stderr); code != 0 {
t.Fatalf("decode returned %d, stderr = %q", code, stderr.String())
}
var out bytes.Buffer
if code := Run([]string{"-encode"}, bytes.NewReader(tagged.Bytes()), &out, &stderr); code != 0 {
t.Fatalf("encode returned %d, stderr = %q", code, stderr.String())
}
want, err := interpres.ParseMap([]byte(doc))
if err != nil {
t.Fatalf("parse of the original: %v", err)
}
got, err := interpres.ParseMap(out.Bytes())
if err != nil {
t.Fatalf("parse of the encoder output (%q): %v", out.String(), err)
}
if !reflect.DeepEqual(want, got) {
t.Errorf("round trip changed the document:\noriginal: %#v\nencoded: %#v\noutput: %q", want, got, out.String())
}
}
+44 -19
View File
@@ -11,9 +11,15 @@ import (
"time"
)
// TOML distinguishes four date-time kinds. interpres decodes an offset
// date-time to a plain time.Time (it carries a zone), and uses the wrapper
// types below for the local variants so callers can tell them apart.
// TOML distinguishes four date-time kinds, and each has its own Go type:
// OffsetDateTime for the offset kind, and the local wrappers below for the
// three that carry no offset. A plain time.Time is accepted wherever an
// offset date-time is, on both the encoding and the decoding side, so a
// timestamp field does not have to name the wrapper.
// OffsetDateTime is a TOML offset date-time, e.g. 1979-05-27T07:32:00-07:00.
// The embedded time.Time is the instant, with the offset the document wrote.
type OffsetDateTime struct{ time.Time }
// LocalDateTime is a TOML local date-time with no offset, e.g.
// 1979-05-27T07:32:00. The embedded time.Time is in UTC.
@@ -27,30 +33,49 @@ type LocalDate struct{ time.Time }
// The embedded time.Time uses the zero date.
type LocalTime struct{ time.Time }
// String returns the TOML-canonical rendering of the offset date-time, e.g.
// "1979-05-27T07:32Z" or "1979-05-27T07:32:00-07:00". The seconds appear only
// when the value carries them, a fractional second drops its trailing zeros,
// and an offset of zero is written "Z".
func (odt OffsetDateTime) String() string { return offsetString(odt.Time) }
// String returns the TOML-canonical rendering of the local date-time, e.g.
// "1979-05-27T07:32:00" or "...:00.000000123" when the time has a fractional
// second. The fractional component is zero-padded to nanosecond precision.
// "1979-05-27T07:32" or "1979-05-27T07:32:00.5" when the time carries a
// fractional second. TOML 1.1 makes the seconds optional, so they appear only
// when they are non-zero, and a fraction drops its trailing zeros.
func (ldt LocalDateTime) String() string {
base := ldt.Format("2006-01-02T15:04:05")
if ns := ldt.Nanosecond(); ns > 0 {
return base + "." + fmt.Sprintf("%09d", ns)
}
return base
return ldt.Format("2006-01-02T") + clockString(ldt.Time)
}
// String returns the TOML-canonical rendering of the local date, e.g.
// "1979-05-27".
func (ld LocalDate) String() string { return ld.Format("2006-01-02") }
// String returns the TOML-canonical rendering of the local time, e.g.
// "07:32:00" or "...:00.000000123" when the time has a fractional second.
// The fractional component is zero-padded to nanosecond precision.
func (lt LocalTime) String() string {
base := lt.Format("15:04:05")
if ns := lt.Nanosecond(); ns > 0 {
return base + "." + fmt.Sprintf("%09d", ns)
// String returns the TOML-canonical rendering of the local time, e.g. "07:32"
// or "07:32:00.5" when the time carries a fractional second.
func (lt LocalTime) String() string { return clockString(lt.Time) }
// clockString renders a time of day the way TOML writes it: the seconds appear
// only when the value carries them, and a fractional second drops its trailing
// zeros, so half a second is "00.5" and not "00.500000000". Both are the same
// value either way; the shorter form is the one TOML 1.1 allows.
func clockString(t time.Time) string {
out := t.Format("15:04")
ns := t.Nanosecond()
if t.Second() != 0 || ns != 0 {
out += t.Format(":05")
}
return base
if ns > 0 {
out += "." + strings.TrimRight(fmt.Sprintf("%09d", ns), "0")
}
return out
}
// offsetString renders an offset date-time, the fourth TOML kind, in the same
// shape: no zero seconds, no trailing zeros in the fraction, and the offset
// written as "Z" when it is zero.
func offsetString(t time.Time) string {
return t.Format("2006-01-02T") + clockString(t) + t.Format("Z07:00")
}
var (
@@ -116,7 +141,7 @@ func parseDateTime(tok string) (any, bool) {
norm := strings.ToUpper(tok)
for _, layout := range offsetDateTimeLayouts {
if t, err := time.Parse(layout, norm); err == nil {
return t, true
return OffsetDateTime{t}, true
}
}
for _, layout := range localDateTimeLayouts {
+87 -6
View File
@@ -4,6 +4,7 @@
package interpres
import (
"encoding"
"fmt"
"reflect"
"slices"
@@ -63,6 +64,19 @@ func (d *decoder) assign(data any, dst reflect.Value) error {
}
}
// A TOML string fills a destination that implements
// encoding.TextUnmarshaler, the rule encoding/json follows. Every other
// value kind keeps its own rule, so an integer still reaches a numeric
// destination.
if s, isString := data.(string); isString {
if tu, ok := textUnmarshalerOf(dst); ok {
if err := tu.UnmarshalText([]byte(s)); err != nil {
return fmt.Errorf("unmarshal text: %w", err)
}
return nil
}
}
switch v := data.(type) {
case map[string]any:
return d.assignTable(v, dst)
@@ -71,6 +85,9 @@ func (d *decoder) assign(data any, dst reflect.Value) error {
case []any:
return d.assignSlice(v, dst)
case string:
if dst.Type() == durationType {
return setDuration(dst, v)
}
return setBasic(dst, reflect.ValueOf(v), "string")
case bool:
return setBasic(dst, reflect.ValueOf(v), "bool")
@@ -78,12 +95,10 @@ func (d *decoder) assign(data any, dst reflect.Value) error {
return setInt(dst, v)
case float64:
return setFloat(dst, v)
case OffsetDateTime:
return setOffsetDateTime(v, dst)
case time.Time:
if dst.Type() != timeType {
return fmt.Errorf("interpres: cannot assign datetime to %s", dst.Type())
}
dst.Set(reflect.ValueOf(v))
return nil
return setDateTime(v, dst)
default:
rv := reflect.ValueOf(data)
if rv.IsValid() && dst.Type() == rv.Type() {
@@ -94,6 +109,26 @@ func (d *decoder) assign(data any, dst reflect.Value) error {
}
}
// textUnmarshalerOf finds the encoding.TextUnmarshaler for dst: on the value
// itself, or on its address, so a pointer-receiver UnmarshalText is invoked on
// an addressable struct field. The TOML date-time types are excluded, because
// they carry time.Time's UnmarshalText through an embedded field while their
// only accepted form is a bare timestamp.
func textUnmarshalerOf(dst reflect.Value) (encoding.TextUnmarshaler, bool) {
if !dst.CanInterface() || isDateTimeType(dst.Type()) {
return nil, false
}
if u, ok := dst.Interface().(encoding.TextUnmarshaler); ok {
return u, true
}
if dst.CanAddr() {
if u, ok := dst.Addr().Interface().(encoding.TextUnmarshaler); ok {
return u, true
}
}
return nil, false
}
func (d *decoder) assignTable(tbl map[string]any, dst reflect.Value) error {
switch dst.Kind() {
case reflect.Struct:
@@ -198,11 +233,57 @@ func (d *decoder) assignTableSlice(items []map[string]any, dst reflect.Value) er
// --- low-level setters -----------------------------------------------------
// setOffsetDateTime stores an offset date-time: in a wrapper destination as it
// is, and in a plain time.Time, which takes the instant with the offset the
// document wrote, so a timestamp field does not have to name the wrapper.
func setOffsetDateTime(v OffsetDateTime, dst reflect.Value) error {
switch dst.Type() {
case offsetDateTimeType:
dst.Set(reflect.ValueOf(v))
case timeType:
dst.Set(reflect.ValueOf(v.Time))
default:
return fmt.Errorf("interpres: cannot assign datetime to %s", dst.Type())
}
return nil
}
// setDateTime stores a time.Time that reached the tree directly, which is the
// shape a tree built by hand carries. Dates the parser produced arrive as
// OffsetDateTime instead.
func setDateTime(v time.Time, dst reflect.Value) error {
switch dst.Type() {
case timeType:
dst.Set(reflect.ValueOf(v))
case offsetDateTimeType:
dst.Set(reflect.ValueOf(OffsetDateTime{v}))
default:
return fmt.Errorf("interpres: cannot assign datetime to %s", dst.Type())
}
return nil
}
func setBasic(dst, val reflect.Value, kind string) error {
if dst.Kind() != val.Kind() {
return fmt.Errorf("interpres: cannot assign %s to %s", kind, dst.Type())
}
dst.Set(val)
// Convert rather than assign: a value of the predeclared type is not
// assignable to a defined type of the same kind, so a plain Set panics on
// a destination such as `type Name string`.
dst.Set(val.Convert(dst.Type()))
return nil
}
// setDuration reads a duration literal into a time.Duration destination. TOML
// has no duration type, so the encoder writes the canonical Go form and the
// decoder reads that back; a bare integer stays the nanosecond count it has
// always been, and reaches the destination through setInt.
func setDuration(dst reflect.Value, s string) error {
d, err := time.ParseDuration(s)
if err != nil {
return fmt.Errorf("interpres: invalid duration %q", s)
}
dst.SetInt(int64(d))
return nil
}
+356 -3
View File
@@ -8,9 +8,11 @@ import (
"errors"
"fmt"
"math"
"net"
"slices"
"strings"
"testing"
"time"
)
func TestSyntaxErrorMessage(t *testing.T) {
@@ -22,7 +24,7 @@ func TestSyntaxErrorMessage(t *testing.T) {
}
func TestParseRejectsInvalidUTF8(t *testing.T) {
_, err := Parse([]byte("v = \"\xff\"\n"))
_, err := ParseMap([]byte("v = \"\xff\"\n"))
if err == nil {
t.Fatal("expected a UTF-8 validation error")
}
@@ -128,7 +130,7 @@ func TestUnmarshalIntoNilAny(t *testing.T) {
func TestParseContextHonoursCancellation(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
if _, err := ParseContext(ctx, []byte("a = 1\n")); !errors.Is(err, context.Canceled) {
if _, err := ParseMapContext(ctx, []byte("a = 1\n")); !errors.Is(err, context.Canceled) {
t.Fatalf("ParseContext returned %v, want context.Canceled", err)
}
}
@@ -158,7 +160,7 @@ func TestContextRoundTrip(t *testing.T) {
in := []byte(`title = "x"
count = 3
`)
if _, err := ParseContext(context.Background(), in); err != nil {
if _, err := ParseMapContext(context.Background(), in); err != nil {
t.Fatalf("ParseContext: %v", err)
}
var out struct {
@@ -725,3 +727,354 @@ func TestDecodeErrorOnMapDestination(t *testing.T) {
t.Fatalf("Path = %v", de.Path)
}
}
func TestUnmarshalIntoDefinedScalarTypes(t *testing.T) {
// A defined type whose underlying kind is string or bool takes the value.
// A bare reflect Set panics on such a type, because a string is not
// assignable to a defined string type without a conversion.
type Name string
type Flag bool
type Cfg struct {
N Name `toml:"n"`
F Flag `toml:"f"`
}
var cfg Cfg
if err := Unmarshal([]byte("n = \"x\"\nf = true\n"), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.N != "x" {
t.Errorf("N = %q, want \"x\"", cfg.N)
}
if !cfg.F {
t.Error("F = false, want true")
}
}
// --- encoding.TextUnmarshaler and time.Duration ----------------------------
// textReceiver implements encoding.TextUnmarshaler on the pointer receiver.
type textReceiver struct{ Text string }
func (t *textReceiver) UnmarshalText(text []byte) error {
t.Text = "got:" + string(text)
return nil
}
// upperText is a defined string type whose UnmarshalText transforms the
// content, so a plain string assignment would leave the wrong value behind.
type upperText string
func (u *upperText) UnmarshalText(text []byte) error {
*u = upperText(strings.ToUpper(string(text)))
return nil
}
// failingTextUnmarshaler fails the decode from UnmarshalText.
type failingTextUnmarshaler struct{}
func (f *failingTextUnmarshaler) UnmarshalText(_ []byte) error { return errors.New("text boom") }
// textAndTOMLReceiver implements both decode interfaces; the TOML method wins.
type textAndTOMLReceiver struct{ From string }
func (t *textAndTOMLReceiver) UnmarshalTOML(any) error { t.From = "toml"; return nil }
func (t *textAndTOMLReceiver) UnmarshalText([]byte) error { t.From = "text"; return nil }
func TestTextUnmarshalerByPointer(t *testing.T) {
type Cfg struct {
R textReceiver `toml:"r"`
}
var cfg Cfg
if err := Unmarshal([]byte(`r = "hello"`), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.R.Text != "got:hello" {
t.Errorf("Text = %q, want \"got:hello\"", cfg.R.Text)
}
}
func TestTextUnmarshalerWinsOverKindAssignment(t *testing.T) {
type Cfg struct {
U upperText `toml:"u"`
}
var cfg Cfg
if err := Unmarshal([]byte(`u = "abc"`), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.U != "ABC" {
t.Errorf("U = %q, want \"ABC\"", cfg.U)
}
}
func TestTextUnmarshalerForNetIP(t *testing.T) {
type Cfg struct {
V4 net.IP `toml:"v4"`
V6 net.IP `toml:"v6"`
IPs []net.IP `toml:"ips"`
}
in := "v4 = \"192.0.2.1\"\nv6 = \"2001:db8::68\"\nips = [\"198.51.100.7\", \"203.0.113.9\"]\n"
var cfg Cfg
if err := Unmarshal([]byte(in), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if got := cfg.V4.String(); got != "192.0.2.1" {
t.Errorf("V4 = %q, want \"192.0.2.1\"", got)
}
if got := cfg.V6.String(); got != "2001:db8::68" {
t.Errorf("V6 = %q, want \"2001:db8::68\"", got)
}
if len(cfg.IPs) != 2 || cfg.IPs[0].String() != "198.51.100.7" || cfg.IPs[1].String() != "203.0.113.9" {
t.Errorf("IPs = %v, want two addresses", cfg.IPs)
}
}
func TestTextUnmarshalerSeesStringsOnly(t *testing.T) {
// An integer keeps its own rule: the text method is not consulted, and the
// value does not reach the receiver.
type Cfg struct {
R textReceiver `toml:"r"`
}
var cfg Cfg
err := Unmarshal([]byte("r = 1\n"), &cfg)
if err == nil {
t.Fatal("expected an integer to be rejected for a text receiver")
}
if cfg.R.Text != "" {
t.Errorf("Text = %q, want it untouched", cfg.R.Text)
}
}
func TestUnmarshalTOMLWinsOverTextUnmarshaler(t *testing.T) {
type Cfg struct {
B textAndTOMLReceiver `toml:"b"`
}
var cfg Cfg
if err := Unmarshal([]byte(`b = "x"`), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.B.From != "toml" {
t.Errorf("From = %q, want \"toml\"", cfg.B.From)
}
}
func TestTextUnmarshalerErrorCarriesPath(t *testing.T) {
type Inner struct {
F failingTextUnmarshaler `toml:"f"`
}
type Cfg struct {
Inner Inner `toml:"inner"`
}
var cfg Cfg
err := Unmarshal([]byte("[inner]\nf = \"x\"\n"), &cfg)
if err == nil {
t.Fatal("expected an error from UnmarshalText")
}
if !strings.Contains(err.Error(), "unmarshal text: text boom") {
t.Errorf("err = %v, want the text error wrapped", err)
}
de, ok := errors.AsType[*DecodeError](err)
if !ok {
t.Fatalf("expected a *DecodeError, got %T: %v", err, err)
}
if !slices.Equal(de.Path, []string{"inner", "f"}) {
t.Fatalf("Path = %v, want [inner f]", de.Path)
}
}
func TestTextUnmarshalerReportsBadText(t *testing.T) {
var cfg struct {
IP net.IP `toml:"ip"`
}
err := Unmarshal([]byte(`ip = "not-an-ip"`), &cfg)
if err == nil {
t.Fatal("expected an error for a malformed address")
}
if !strings.Contains(err.Error(), "unmarshal text:") {
t.Errorf("err = %v, want it wrapped as a text error", err)
}
}
func TestUnmarshalDurations(t *testing.T) {
type Cfg struct {
FromText time.Duration `toml:"from_text"`
FromInt time.Duration `toml:"from_int"`
Fraction time.Duration `toml:"fraction"`
}
in := "from_text = \"1h30m\"\nfrom_int = 5400000000000\nfraction = \"1.5s\"\n"
var cfg Cfg
if err := Unmarshal([]byte(in), &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.FromText != 90*time.Minute {
t.Errorf("FromText = %v, want %v", cfg.FromText, 90*time.Minute)
}
if cfg.FromInt != 90*time.Minute {
t.Errorf("FromInt = %v, want %v", cfg.FromInt, 90*time.Minute)
}
if cfg.Fraction != 1500*time.Millisecond {
t.Errorf("Fraction = %v, want %v", cfg.Fraction, 1500*time.Millisecond)
}
}
func TestUnmarshalDurationRejectsMalformedText(t *testing.T) {
var cfg struct {
D time.Duration `toml:"d"`
}
err := Unmarshal([]byte("d = \"90\"\n"), &cfg)
if err == nil {
t.Fatal("expected an error for a duration without a unit")
}
if !strings.Contains(err.Error(), "invalid duration") {
t.Errorf("err = %v, want an invalid-duration message", err)
}
}
func TestQuotedStringNeverBecomesDateTime(t *testing.T) {
// The date-time types take a bare timestamp only, so the text path is
// excluded for them and a quoted string stays a string.
var stamp struct {
S time.Time `toml:"s"`
}
err := Unmarshal([]byte("s = \"2026-06-26T10:00:00Z\"\n"), &stamp)
if err == nil {
t.Fatal("expected a quoted string to be rejected for time.Time")
}
if !strings.Contains(err.Error(), "cannot assign string") {
t.Errorf("err = %v, want a cannot-assign message", err)
}
var day struct {
D LocalDate `toml:"d"`
}
if err := Unmarshal([]byte("d = \"1979-05-27\"\n"), &day); err == nil {
t.Fatal("expected a quoted string to be rejected for LocalDate")
}
}
func TestDecoderMaxDepth(t *testing.T) {
deep := func(n int) []byte {
return []byte("v = " + strings.Repeat("[", n) + strings.Repeat("]", n) + "\n")
}
var cfg struct {
V any `toml:"v"`
}
if err := NewDecoder().MaxDepth(4).Decode(deep(4), &cfg); err != nil {
t.Fatalf("at the limit: %v", err)
}
err := NewDecoder().MaxDepth(4).Decode(deep(5), &cfg)
if err == nil {
t.Fatal("expected a nesting error")
}
if !strings.Contains(err.Error(), "limit of 4") {
t.Errorf("err = %v, want it to name the limit", err)
}
}
func TestDecoderMaxInputSize(t *testing.T) {
doc := []byte("v = \"ab\"\n")
var cfg struct {
V string `toml:"v"`
}
if err := NewDecoder().MaxInputSize(len(doc)).Decode(doc, &cfg); err != nil {
t.Fatalf("at the limit: %v", err)
}
err := NewDecoder().MaxInputSize(len(doc)-1).Decode(doc, &cfg)
if err == nil {
t.Fatal("expected a size error")
}
if !strings.Contains(err.Error(), "over the limit of 8") {
t.Errorf("err = %v, want it to name the limit", err)
}
// Parse carries the nesting default and no size limit.
if _, err := ParseMap(doc); err != nil {
t.Fatalf("parse: %v", err)
}
}
// --- OffsetDateTime --------------------------------------------------------
func TestOffsetDateTimeIsTheParsedType(t *testing.T) {
// A document's offset date-time arrives as the wrapper, and a plain
// time.Time destination still takes it, so a timestamp field needs no
// change to keep working.
in := []byte("stamp = 2026-06-26T10:00:00-07:00\n")
want := time.Date(2026, 6, 26, 10, 0, 0, 0, time.FixedZone("", -7*3600))
var plain struct {
Stamp time.Time `toml:"stamp"`
}
if err := Unmarshal(in, &plain); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if !plain.Stamp.Equal(want) {
t.Errorf("time.Time destination = %v, want %v", plain.Stamp, want)
}
var wrapped struct {
Stamp OffsetDateTime `toml:"stamp"`
}
if err := Unmarshal(in, &wrapped); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if !wrapped.Stamp.Time.Equal(want) {
t.Errorf("OffsetDateTime destination = %v, want %v", wrapped.Stamp.Time, want)
}
if got := wrapped.Stamp.String(); got != "2026-06-26T10:00-07:00" {
t.Errorf("String() = %q, want 2026-06-26T10:00-07:00", got)
}
}
func TestOffsetDateTimeFromHandBuiltTree(t *testing.T) {
// A tree built by hand may carry a plain time.Time, which is the other
// source of the offset kind; both date-time destinations take it.
tree := map[string]any{"stamp": time.Date(2026, 6, 26, 10, 0, 0, 0, time.UTC)}
want := time.Date(2026, 6, 26, 10, 0, 0, 0, time.UTC)
var plain struct {
Stamp time.Time `toml:"stamp"`
}
if err := newDecoder().decode(tree, &plain); err != nil {
t.Fatalf("decode: %v", err)
}
if !plain.Stamp.Equal(want) {
t.Errorf("time.Time destination = %v, want %v", plain.Stamp, want)
}
var wrapped struct {
Stamp OffsetDateTime `toml:"stamp"`
}
if err := newDecoder().decode(tree, &wrapped); err != nil {
t.Fatalf("decode: %v", err)
}
if !wrapped.Stamp.Time.Equal(want) {
t.Errorf("OffsetDateTime destination = %v, want %v", wrapped.Stamp.Time, want)
}
}
// dateKindReceiver records the Go type UnmarshalTOML was handed.
type dateKindReceiver struct{ Kind string }
func (r *dateKindReceiver) UnmarshalTOML(data any) error {
r.Kind = fmt.Sprintf("%T", data)
return nil
}
func TestUnmarshalerReceivesOffsetDateTime(t *testing.T) {
// The interface sees the wrapper, which names the date-time kind on its
// own; the local kinds keep their own wrappers.
var cfg struct {
O dateKindReceiver `toml:"o"`
L dateKindReceiver `toml:"l"`
}
in := []byte("o = 2026-06-26T10:00:00Z\nl = 2026-06-26T10:00:00\n")
if err := Unmarshal(in, &cfg); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if cfg.O.Kind != "interpres.OffsetDateTime" {
t.Errorf("offset kind = %q, want interpres.OffsetDateTime", cfg.O.Kind)
}
if cfg.L.Kind != "interpres.LocalDateTime" {
t.Errorf("local kind = %q, want interpres.LocalDateTime", cfg.L.Kind)
}
}
+249 -41
View File
@@ -1,36 +1,107 @@
# API
The library exports the surface below from the `sourcedock.dev/petrbalvin/interpres`
The library exports the surface below from the `sourcedock.dev/petrbalvin/interpres/v2`
package. The snippets assume:
```go
import "sourcedock.dev/petrbalvin/interpres"
import "sourcedock.dev/petrbalvin/interpres/v2"
```
The parser accepts TOML 1.0 documents plus the TOML 1.1 extensions: date-times
and times without seconds, the `\e` and `\xHH` escape sequences, and
multi-line inline tables with comments and trailing commas. The encoder emits
TOML 1.0, which is valid under both versions.
The parser implements TOML 1.1: date-times and times without seconds, the
`\e` and `\xHH` escape sequences, and multi-line inline tables with comments
and trailing commas. The encoder emits TOML 1.1.
## Functions
### `func Parse(data []byte) (map[string]any, error)`
### `func Parse(data []byte) (*Document, error)`
Decodes a TOML document into an untyped tree, using the value mapping in the
[Decoding](#decoding) section below. Returns `*SyntaxError` on a malformed
document. Input that is not valid UTF-8 is rejected before the parser runs.
Equivalent to `ParseContext(context.Background(), data)`.
Decodes a TOML document into a [Document](#documents): the values, the order the
keys were written in, whether a table was written inline, and the comments.
The values follow the mapping in the [Decoding](#decoding) section below.
Returns `*SyntaxError` on a malformed document. Input that is not valid UTF-8
is rejected before the parser runs. Equivalent to
`ParseContext(context.Background(), data)`.
```go
tree, err := interpres.Parse([]byte("title = \"x\"\nport = 8080\n"))
doc, err := interpres.Parse([]byte("title = \"x\"\nport = 8080\n"))
tree := doc.Map()
```
### `func ParseContext(ctx context.Context, data []byte) (map[string]any, error)`
### `func ParseContext(ctx context.Context, data []byte) (*Document, error)`
The cancellable variant of `Parse`. An already-cancelled context returns
`ctx.Err()` before any work. During parsing the context is checked every 64
top-level statements, so a long document aborts without running to completion.
### `func ParseMap(data []byte) (map[string]any, error)`
Decodes a TOML document into an untyped tree, the shape this package parsed
into before [Document](#documents) existed: the order of the keys and the
comments are not part of a map, so they are dropped. Use it when only the
values matter, or when the extra bookkeeping of a document is not wanted.
Equivalent to `ParseMapContext(context.Background(), data)`.
```go
tree, err := interpres.ParseMap([]byte("title = \"x\"\nport = 8080\n"))
```
### `func ParseMapContext(ctx context.Context, data []byte) (map[string]any, error)`
The cancellable variant of `ParseMap`.
## Documents
`Parse` returns a `Document`: the value tree together with what a map cannot
carry, which is the order the keys were written in, whether a table was written
as an inline table or under a header, and the comments. `ParseMap` gives the
plain tree when none of that is wanted.
```go
doc, err := interpres.Parse(data)
if err != nil {
return err
}
root := doc.Root()
for _, key := range root.Keys() { // written order, not sorted
entry, _ := root.Get(key)
fmt.Println(key, entry.Value())
}
```
The values are shared with the tree `ParseMap` returns, so a value read from a
document and from `doc.Map()` is the same value.
| Type | Meaning |
|---|---|
| `Document` | the parsed document: `Root()` for the top-level table, `Map()` for the value tree, `Footer()` for a comment block at the end |
| `Table` | one TOML table: `Keys()` and `Entries()` in written order, `Get(key)`, `Values()` for its part of the value tree, `Inline()` |
| `Entry` | one key: `Value()`, `Inline()`, `Table()` when the value is a table, `Elements()` for the tables of an array value |
`Elements()` holds one node per element of an array value: the tables of an
array of tables, and the inline tables inside a value array, with `nil` for the
elements that are not tables.
### Comments
A comment belongs to the line it precedes or follows, and to the node that line
introduced:
| Written | Carried by |
|---|---|
| lines above a key | that key's `Entry`, through `Comments()` |
| a comment beside a key | that key's `Entry`, through `Trailing()` |
| lines above a `[header]` or `[[header]]` | that `Table`, through `Comments()` |
| a comment beside a header | that `Table`, through `Trailing()` |
| a comment block after the last statement | the `Document`, through `Footer()` |
`SetComments` and `SetTrailing` replace them. A line carries no leading `#`
and no surrounding space, so `# note` is stored as `note` and a bare `#` as
`""`.
A `Document` is not a value to marshal: `Marshal` writes values, so it refuses
one and points at `doc.Map()`. Writing a document back, with its order and its
comments, belongs with the editing API.
### `func Unmarshal(data []byte, v any) error`
Parses `data` and stores the result in the value pointed to by `v`, typically a
@@ -75,7 +146,7 @@ and every 64 fields during the reflection walk.
| integer | `int64` |
| float | `float64` |
| boolean | `bool` |
| offset date-time | `time.Time` |
| offset date-time | `OffsetDateTime` |
| local date-time | `LocalDateTime` |
| local date | `LocalDate` |
| local time | `LocalTime` |
@@ -134,20 +205,26 @@ The decoder converts to the destination type with explicit overflow checks:
| `uint`, `uint8`, `uint16`, `uint32`, `uint64` | the value must be non-negative and must not overflow the destination's own width, `uint` on a 32-bit platform included; `uint64` accepts any non-negative `int64` |
| `float32`, `float64` | copied verbatim, except that a finite value beyond the `float32` range is an overflow error rather than a silent infinity; an integer also coerces, so TOML `5` decodes into `5.0` |
| `bool`, `string` | exact kind match only, no coercion across kinds |
| `time.Time` | offset date-times only; no implicit conversion to or from the local variants |
| `time.Time`, `OffsetDateTime` | offset date-times only; no implicit conversion to or from the local variants |
A conversion that the rules do not allow produces an error wrapped with the
offending key or index, for example `p: interpres: integer 300 overflows uint8`.
### Date-time values
Offset date-times decode into `time.Time` and keep their offset. The local
variants decode into `LocalDateTime`, `LocalDate` and `LocalTime`, whose
embedded `time.Time` is normalised to UTC (midnight UTC for a local date, the
zero date for a local time). Every kind may omit the seconds as of TOML 1.1
(`07:32`, `1979-05-27T07:32`); such a value carries a zero second, and the
canonical rendering writes full seconds. There is no implicit conversion
between the offset and local kinds; assigning one to the other is an error.
Offset date-times decode into `OffsetDateTime`, whose embedded `time.Time` is the
instant with the offset the document wrote; a destination of the plain
`time.Time` takes the same value, so a timestamp field does not have to name the
wrapper. The local variants decode into `LocalDateTime`, `LocalDate` and
`LocalTime`, whose embedded `time.Time` is normalised to UTC (midnight UTC for a
local date, the zero date for a local time). Every kind may omit the seconds as
of TOML 1.1 (`07:32`, `1979-05-27T07:32`); such a value carries a zero second,
and the encoder writes the seconds only when the value carries them, so a
document written without seconds comes back without them. There is no implicit
conversion between the offset and local kinds; assigning one to the other is an
error. The date-time types take a bare timestamp and never a quoted string, so a
document that writes a date-time with quotes does not decode into them, and
neither `encoding.TextUnmarshaler` nor the embedded `time.Time` changes that.
### Arrays of tables
@@ -167,7 +244,7 @@ type Unmarshaler interface {
```
`data` is whatever the parser produced for that key: `string`, `bool`, `int64`,
`float64`, `time.Time`, `LocalDateTime`, `LocalDate`, `LocalTime`, `[]any`, or
`float64`, `OffsetDateTime`, `LocalDateTime`, `LocalDate`, `LocalTime`, `[]any`, or
`map[string]any`. The method inspects the value and mutates its own receiver;
the decoder keeps whatever state the receiver stored.
@@ -178,6 +255,38 @@ automatically, and a nil pointer destination is allocated first. An error
returned from `UnmarshalTOML` halts the decode and propagates wrapped with the
key path, for example `addr: unmarshal: not a string`.
### Custom decoding: `encoding.TextUnmarshaler`
A destination type that implements `encoding.TextUnmarshaler` receives a TOML
string as its text content, the rule `encoding/json` follows:
```go
func (ip *IP) UnmarshalText(text []byte) error
```
The decoder looks for the method on the destination and on its address, so a
pointer-receiver `UnmarshalText` is invoked on an addressable struct field, and
the elements of a slice destination are reached the same way. The text path
applies to TOML strings only: every other value kind keeps its own rule, so
`r = 1` does not reach a receiver that expects text. An error from
`UnmarshalText` halts the decode and propagates with the key path and the
prefix `unmarshal text:`, for example `addr: unmarshal text: not an address`.
[`UnmarshalTOML`](#custom-decoding-unmarshaler) wins over `UnmarshalText` when
a type implements both, and the four [date-time
types](#date-time-values) are excluded: a quoted string stays a string and
never becomes an `OffsetDateTime` or one of the local wrappers.
### Durations
TOML has no duration type, so `time.Duration` has a rule of its own. The
encoder writes the canonical Go form in a TOML string, `1h30m0s`, and the
decoder reads that string back with `time.ParseDuration`. A bare integer is
still the nanosecond count it has always been, so `from_int = 5400000000000`
and `from_text = "1h30m"` decode to the same duration. Text that
`time.ParseDuration` rejects, `d = "90"` among it, fails with
`interpres: invalid duration "90"`.
### Strict decoding
By default unknown keys are dropped silently. A `Decoder` built with
@@ -281,7 +390,7 @@ them.
By default every table is emitted with its entries grouped by kind:
1. scalars (`string`, `int64`, `float64`, `bool`, `time.Time`,
`LocalDateTime`, `LocalDate`, `LocalTime`)
`OffsetDateTime`, `LocalDateTime`, `LocalDate`, `LocalTime`)
2. sub-tables (structs and `map[string]V` values)
3. arrays of tables (`[]struct` and `[]map[string]V`)
@@ -316,11 +425,20 @@ type Marshaler interface {
The returned value is encoded as if it had been passed in place of the
receiver, so it may be a scalar, a slice, an array of tables, or another
struct or map, including the `Marshaler` result of another type; the encoder
recurses. An error returned from `MarshalTOML` fails the marshal wrapped with
the key path, for example `interpres: server.port: bad timestamp`. A result
of `nil` with a nil error fails the same way with
`MarshalTOML returned a nil value`: nil has no TOML representation, so
dropping the field silently is not an option.
recurses, and the result is normalised like any other value, so a method may
return a plain `int` or a `time.Duration`.
An error returned from `MarshalTOML` fails the marshal wrapped with the key
path, for example `interpres: server.port: bad timestamp`. A result of `nil` with
a nil error fails the same way with `MarshalTOML returned a nil value`: nil has
no TOML representation, so dropping the field silently is not an option.
The method is reached for every value the walk meets, array elements included:
an element that renders itself as a table keeps the `[[header]]` form, one that
renders itself as a scalar turns the array into a value array, and the method
runs once per element. It is looked up on the value and on its address, so a
pointer-receiver method is called for a field or an element, exactly as
`MarshalText` is.
```go
type Port int
@@ -330,6 +448,27 @@ func (p Port) MarshalTOML() (any, error) {
}
```
### Custom encoding: `encoding.TextMarshaler`
A type that implements `encoding.TextMarshaler` is encoded as a TOML string
holding the text the method returns, which is the rule `encoding/json` follows:
```go
func (ip IP) MarshalText() ([]byte, error)
```
The encoder looks for the method on the value and on its address, so a
pointer-receiver `MarshalText` is found on a struct field of an addressable
value (pass a pointer to `Marshal`) and always on a slice element. `net.IP`,
`netip.Addr` and user types follow this rule, and a struct that implements the
interface becomes a string rather than a table. `MarshalTOML` wins when a type
implements both, the four [date-time types](#date-time-values) keep their bare
timestamp form, and text that is not valid UTF-8 is an error rather than a
replacement character.
A duration carries no text method of its own; see [Durations](#durations) for
its rule.
### Arrays
An array whose every element is a table (`[]struct`, `[]map[string]V`, after
@@ -357,10 +496,9 @@ omitted, because TOML forbids an empty `[[a]]`. Other empty arrays emit as
### Long strings
By default every string is emitted as a basic `"..."` string with the escapes
TOML requires, and a string containing a newline is emitted as an escaped
multi-line basic string. `UseLiteralMultiline(threshold)` switches strings that
contain a newline and are at least `threshold` bytes long to the literal
`'''...'''` form, which carries the newlines verbatim:
TOML requires, a newline among them as `\n`. `UseLiteralMultiline(threshold)`
switches strings that contain a newline and are at least `threshold` bytes long
to the literal `'''...'''` form, which carries the newlines verbatim:
```go
out, err := interpres.NewEncoder().UseLiteralMultiline(80).Marshal(cfg)
@@ -372,6 +510,48 @@ carry verbatim (an embedded run of three single quotes, a control character
other than tab or newline, or a carriage return outside a CRLF pair) also keeps
the basic form, so the output always re-parses to the same value.
### Inline tables
A table element of a value array, and a sub-table inlined by
[`InlineTables`](#compact-documents), is written as one `{a = 1, b = 2}` line
while it fits. An inline table that would pass the hundredth column carries
newlines and a trailing comma instead, which TOML 1.1 allows:
```toml
arr = [1, {
n = 1,
name = "a value long enough to push this line well past the one hundred column limit",
}]
```
The closing brace and the entries are indented one tab per nesting level, a
nested table is measured on its own line, and the output re-parses to the same
value either way.
### Compact documents
`InlineTables(threshold)` writes a sub-table as an inline table when its
single-line rendering is at most `threshold` bytes, and as a table header
section when it is longer. A document of small tables therefore grows shorter:
```go
out, err := interpres.NewEncoder().InlineTables(60).Marshal(cfg)
```
With `60` and a table of three short entries, the same value is written
```toml
server = {host = "127.0.0.1", port = 9090, tls = {on = false}}
```
instead of three lines under a `[server]` header and a `[server.tls]` section.
A nested sub-table takes part in the same way, and the whole option is off at
`0` or less. Two limits are deliberate. An array of tables keeps the `[[a]]`
header form, because its inline form re-parses as a value array and would change
the value's Go type. And because an inlined table is a value line, every one of
them precedes the first header of its document, so a table inlined next to a
header is not read back as part of that header's section.
### Cancellation
`MarshalContext` and `(*Encoder).MarshalContext` accept a `context.Context`. The
@@ -386,6 +566,8 @@ The output is not byte-identical to any document that produced the value:
- map keys are emitted in sorted order
- the choice between `[table]` headers and inline tables is not preserved
- strings use the basic quoted form unless the literal option above applies
- a date-time drops its zero seconds and the trailing zeros of its fraction, so
`07:32:00` is written `07:32`; both are the same value
- floats always carry a `.` or an exponent, so a float `1` is emitted as `1.0`
and stays distinguishable from the integer `1` across a round-trip; negative
zero is normalised to `0.0`
@@ -413,9 +595,11 @@ sequenceDiagram
### `type SyntaxError struct{ Line int; Msg string }`
Describes a malformed TOML document; `Line` is 1-based and `Error()` renders as
`interpres: line N: msg`. Read the structured fields with a type assertion or
`errors.AsType`:
Describes a document the parser rejected, with the 1-based `Line` at which it
gave up and `Error()` rendering as `interpres: line N: msg`. A malformed
document is the usual cause; the nesting limit and an input that is not valid
UTF-8 report through the same type. Read the structured fields with a type
assertion or `errors.AsType`:
```go
if se, ok := errors.AsType[*interpres.SyntaxError](err); ok {
@@ -451,6 +635,19 @@ with `DisallowUnknownFields`, then call `Decode` or `DecodeContext` any number
of times. A configured `Decoder` holds no per-call state and is safe for
concurrent use.
| Method | Default | Effect |
|---|---|---|
| `DisallowUnknownFields()` | off | a key with no matching struct field is an error |
| `MaxDepth(depth int)` | `10000` | bound how deeply arrays and inline tables may nest |
| `MaxInputSize(size int)` | no limit | bound the size of the document, in bytes |
The nesting limit protects the stack, because the parser is a recursive
descent: a deeper document is rejected with a `SyntaxError` naming the limit
rather than running the stack out. `Parse` and `ParseContext` carry that same
default but take no options. The size limit is off by default, because the
caller already holds the bytes and the size is therefore a policy, not a
protection the library can impose on its own.
### `type Encoder`
Configurable emission policy, constructed with `NewEncoder`. The option state
@@ -462,12 +659,14 @@ encoder:
| `GroupByKind(v bool)` | `true` | group entries as scalars, then sub-tables, then arrays of tables; `false` preserves declaration order |
| `OmitEmptyArrays()` | off | skip `key = []` for empty scalar arrays |
| `UseLiteralMultiline(threshold int)` | `0` | emit multi-line strings of at least `threshold` bytes as literal `'''...'''` |
| `InlineTables(threshold int)` | `0` | write a sub-table inline when its single-line form is at most `threshold` bytes |
```go
out, err := interpres.NewEncoder().
GroupByKind(false).
OmitEmptyArrays().
UseLiteralMultiline(80).
InlineTables(60).
MarshalContext(ctx, cfg)
```
@@ -475,6 +674,11 @@ A configured `Encoder` holds no per-call state; each `Marshal` or
`MarshalContext` call copies the options and is safe for concurrent use, as
long as no setter races with a call.
### `type Document`, `type Table`, `type Entry`
See [Documents](#documents). A `Document` is what `Parse` returns, and it is
not a value `Marshal` accepts.
### `type Marshaler interface{ MarshalTOML() (any, error) }`
See [Custom encoding](#custom-encoding-marshaler).
@@ -486,24 +690,28 @@ See [Custom decoding](#custom-decoding-unmarshaler).
### Date-time wrappers
```go
type OffsetDateTime struct{ time.Time } // 1979-05-27T07:32:00Z
type LocalDateTime struct{ time.Time } // 1979-05-27T07:32:00
type LocalDate struct{ time.Time } // 1979-05-27
type LocalTime struct{ time.Time } // 07:32:00.999999
```
Each carries a `String()` method returning the TOML-canonical rendering, with
the fractional second zero-padded to nanosecond precision when present. The
types are produced by `Parse` and accepted by `Marshal`.
Each carries a `String()` method returning the TOML-canonical rendering: the
seconds appear only when the value carries them, and a fractional second drops
its trailing zeros, so `07:32:00` renders as `07:32` and a half second as
`00.5`. The types are produced by `Parse` and accepted by `Marshal`, which
writes them through `String()`.
## Errors
The entry points return:
- `*SyntaxError` for a malformed document, with the 1-based line
- `*SyntaxError` for a malformed document, with the 1-based line; the nesting
limit reports through it as well
- `*DecodeError` for a decoding failure, with the key path in `Path`
- `*EncodeError` for an encoding failure, with the key path in `Path`
- a plain error for the rest: a non-pointer decode target, a cancelled
context, a key that is not valid UTF-8
context, a key that is not valid UTF-8, an input over the size limit
Decode and encode failures carry the key path or element index in the typed
wrappers above, so `errors.Is` and `errors.AsType` see through them and the
+7 -6
View File
@@ -6,10 +6,10 @@ source tree; nothing is aspirational.
## Overview
interpres is one public library package, one command, and one example. The
library implements the whole of TOML 1.0 and 1.1, decoding and encoding, in the
standard library alone; the command wraps the parser for the toml-test
compliance harness, against which it stands at 214 valid and 467 invalid cases
with zero failures; the example demonstrates the API.
library implements the whole of TOML 1.1, decoding and encoding, in the
standard library alone; the command wraps the parser and the encoder for the
toml-test compliance harness, against which it stands at 214 valid, 467 invalid
and 214 encoder cases with zero failures; the example demonstrates the API.
```mermaid
flowchart TD
@@ -36,14 +36,15 @@ strict validation.
| Path | Responsibility |
|---|---|
| `.` (package `interpres`) | The whole library. `interpres.go` declares the exported surface (`Parse`, `Unmarshal`, `Marshal`, the `*Context` variants, `Decoder`, `Encoder`, `Marshaler`, `Unmarshaler`, `SyntaxError`, the local date-time types); everything below it is unexported. |
| `cmd/interpres-decode` | The toml-test adapter. Reads TOML on stdin, writes tagged JSON on stdout. Owns no parsing logic. |
| `cmd/interpres-decode` | The toml-test adapter, both directions. Reads TOML on stdin, writes tagged JSON on stdout; with `-encode` it reads tagged JSON and writes TOML. Owns no parsing logic and no emission logic. |
| `examples/basic` | A runnable tour of the API. Documentation in executable form, not part of the library. |
Inside the library package, one file owns one concern:
| File | Responsibility |
|---|---|
| `parser.go` | The recursive-descent parser. Produces the `map[string]any` tree and enforces the structural rules of TOML 1.0 and 1.1 (table redefinitions, dotted keys, arrays of tables, multi-line inline tables). Reports a 1-based line on failure. |
| `parser.go` | The recursive-descent parser. Produces the `map[string]any` tree, records the nodes a [Document](API.md#documents) is built from, and enforces the structural rules of TOML 1.1 (table redefinitions, dotted keys, arrays of tables, multi-line inline tables). Reports a 1-based line on failure. |
| `document.go` | The parsed-document types: `Document`, `Table` and `Entry`, which carry the key order, whether a table was written inline, and the comments. The values they expose are the parser's own tree, not a copy. |
| `number.go` | Strict numeric tokens: integers in the four radixes with `_` separators, and floats including `inf` and `nan`. Rejects leading zeros, misplaced underscores and malformed fractions. |
| `datetime.go` | The three local date-time wrapper types and `parseDateTime`, which classifies a token into the four date-time kinds under the strict TOML grammar. |
| `decode.go` | Maps the parsed tree onto Go values by reflection: struct fields, maps, slices, scalar conversion with overflow checks, `Unmarshaler` dispatch. |
+41 -13
View File
@@ -1,25 +1,32 @@
# Command line
The reference below is taken from the program itself. `interpres-decode` is
the toml-test harness adapter, and it also validates documents. Install it
with Go itself, no release assets involved:
the toml-test harness adapter in both directions, decoding TOML into tagged
JSON and encoding tagged JSON back into TOML, and it also validates documents.
Install it with Go itself, no release assets involved:
```sh
go install sourcedock.dev/petrbalvin/interpres/cmd/interpres-decode@latest
go install sourcedock.dev/petrbalvin/interpres/v2/cmd/interpres-decode@latest
```
## Synopsis
```sh
interpres-decode [flags]
interpres-decode -encode
interpres-decode -validate [file ...]
```
Without `-validate` the program is the toml-test adapter: it takes no
arguments, reads one TOML document from stdin, and writes the toml-test
tagged-JSON form to stdout. Build it locally with `just build`, which
compiles it into `bin/interpres-decode`, or run it straight from the module
directory with `just run`.
Without `-validate` or `-encode` the program is the decoding half of the
toml-test adapter: it takes no arguments, reads one TOML document from stdin,
and writes the toml-test tagged-JSON form to stdout. Build it locally with
`just build`, which compiles it into `bin/interpres-decode`, or run it
straight from the module directory with `just run`.
With `-encode` the direction is reversed: the program reads a tagged-JSON
description from stdin and writes the TOML document it describes to stdout,
which is the shape toml-test expects of an encoder command. It takes no
arguments either, and `-validate` and `-encode` cannot be combined.
With `-validate` the program parses each named file instead, or stdin when no
file is named, and prints one line per invalid document to stderr. It is
@@ -31,15 +38,16 @@ means stdin.
| Flag | Effect |
|---|---|
| `-validate` | validate the documents instead of emitting tagged JSON |
| `-encode` | read tagged JSON from stdin and write TOML instead |
| `-h` | print the usage |
## Exit codes
| Code | Meaning |
|---|---|
| `0` | adapter: the document parsed and the tagged JSON was written; validate: every document parsed |
| `0` | adapter: the document parsed and the tagged JSON was written; encode: the TOML was written; validate: every document parsed |
| `1` | adapter: parse error; validate: at least one document is invalid |
| `2` | a usage error, a read failure, or a value with no tagged representation |
| `2` | a usage error, a read failure, malformed tagged JSON, or a value with no TOML representation |
## Wire format
@@ -66,6 +74,14 @@ wrapped in an object with a `type` and a `value`:
| local date | `date-local` | `1979-05-27` |
| local time | `time-local` | `07:32:00.999999` |
The `-encode` mode reads exactly this form back. Two properties of it are
worth knowing. A float whose value has no fraction and no exponent is written
as a bare integer string, `{"type": "float", "value": "1"}`, so there the tag
decides the type and not the literal. And the form cannot tell an array of
tables from a value array of inline tables, so the adapter writes the header
form, `[[a]]`, for an array whose every element is a JSON object; a mixed
array keeps the value form.
## Examples
Echo a small document through the adapter:
@@ -78,8 +94,18 @@ port = 9090
' | ./bin/interpres-decode
```
The output is the equivalent value tree as one JSON object. Validate the
TOML files of another repository in CI:
The output is the equivalent value tree as one JSON object. Turn a description
back into TOML with `-encode`:
```sh
echo '{"title": {"type": "string", "value": "hello"}}' | ./bin/interpres-decode -encode
```
```toml
title = "hello"
```
Validate the TOML files of another repository in CI:
```sh
interpres-decode -validate config.toml deploy/example.toml
@@ -101,5 +127,7 @@ just toml-test
```
That recipe needs the `toml-test` binary on `PATH`, installed with
`go install github.com/toml-lang/toml-test/v2/cmd/toml-test@v2.2.0`. The full
`go install github.com/toml-lang/toml-test/v2/cmd/toml-test@v2.2.0`. It runs
the suite in both directions: the decoder against the valid and invalid
corpora, and the encoder against the tagged JSON of the valid one. The full
reference for the library itself is [API.md](API.md).
+1 -1
View File
@@ -41,7 +41,7 @@ prints the same list.
| `just run` | `go run ./cmd/interpres-decode`, reads TOML from stdin |
| `just dev` | the same run, for iterating |
| `just example` | `go run ./examples/basic`, the usage tour |
| `just toml-test` | builds the adapter and runs the official toml-test compliance suite against it |
| `just toml-test` | builds the adapter and runs the official toml-test compliance suite against it, decoder and encoder |
| `just coverage-html` | `just test`, then `go tool cover -html` into `coverage.html` |
| `just install` | builds, then copies the binary into `~/.local/bin` (`BINDIR` overrides) |
| `just uninstall` | removes the installed binary |
+196
View File
@@ -0,0 +1,196 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: MIT
package interpres
// A Document is a parsed TOML document: the values, plus what the map shape
// cannot carry, which is the order the keys were written in, whether a table
// was written inline or under a header, and the comments.
//
// The values are the tree ParseMap returns, shared rather than copied, so a
// value read from a Document and from that map is the same value. A comment
// belongs to the statement it precedes: the lines above a key belong to the
// key, the lines above a header belong to the header's table, and a comment
// block after the last statement belongs to the Document.
//
// Comments inside array and inline-table values are not carried yet; the
// parser skips them as it always has.
type Document struct {
root *Table
footer []string
}
// Root returns the document's root table.
func (d *Document) Root() *Table { return d.root }
// Map returns the value tree, the shape ParseMap gives. It is the tree the
// document was parsed into, not a copy.
func (d *Document) Map() map[string]any { return d.root.values }
// Footer returns the comment lines that follow the last statement, and every
// line of a document that holds no statement at all.
func (d *Document) Footer() []string { return d.footer }
// SetFooter replaces those lines.
func (d *Document) SetFooter(lines []string) { d.footer = lines }
// A Table is one TOML table: its values, its keys in written order, and the
// comments around the header or the key that introduced it.
type Table struct {
values map[string]any
entries []*Entry
index map[string]*Entry
// inline records that the table was written as an inline table, `{…}`,
// rather than under a header or as a dotted key.
inline bool
// comments are the lines above the table's header, trailing is the comment
// on the header's own line. Both are empty for a table a dotted key
// introduced, which has no line of its own.
comments []string
trailing string
}
func newTable(values map[string]any) *Table {
return &Table{values: values, index: map[string]*Entry{}}
}
// Keys returns the table's keys in the order they were written.
func (t *Table) Keys() []string {
keys := make([]string, len(t.entries))
for i, e := range t.entries {
keys[i] = e.key
}
return keys
}
// Values returns the table's values, which is the map the value tree holds for
// it.
func (t *Table) Values() map[string]any { return t.values }
// Entries returns the table's entries in written order.
func (t *Table) Entries() []*Entry { return t.entries }
// Get returns the entry for key, and whether the table has one.
func (t *Table) Get(key string) (*Entry, bool) {
e, ok := t.index[key]
return e, ok
}
// Inline reports whether the table was written as an inline table, `{…}`,
// rather than under a header or introduced by a dotted key.
func (t *Table) Inline() bool { return t.inline }
// Comments returns the comment lines above the table's header, or above the
// key that introduced it. Lines carry no leading '#' and no surrounding space.
func (t *Table) Comments() []string { return t.comments }
// SetComments replaces those lines. Each line is written back with a "# " in
// front of it, so a line should not carry one.
func (t *Table) SetComments(lines []string) { t.comments = lines }
// Trailing returns the comment on the header's own line, without the '#'.
func (t *Table) Trailing() string { return t.trailing }
// SetTrailing replaces that comment.
func (t *Table) SetTrailing(line string) { t.trailing = line }
// addValue records a key of the table, in written order.
func (t *Table) addValue(key string, val any, inline bool) *Entry {
e := &Entry{table: t, key: key, inline: inline}
t.entries = append(t.entries, e)
t.index[key] = e
if node, ok := val.(map[string]any); ok {
e.child = newTable(node)
}
return e
}
// addTable records a key whose value is a table introduced by a header or a
// dotted key, and returns the table's node.
func (t *Table) addTable(key string, values map[string]any) *Table {
if e, ok := t.index[key]; ok {
// The key was seen before, as the leaf of an earlier dotted key.
if e.child == nil {
e.child = newTable(values)
}
return e.child
}
e := t.addValue(key, values, false)
e.child = newTable(values)
return e.child
}
// addElement records one element of an array of tables, and returns its node.
func (t *Table) addElement(key string, values map[string]any) *Table {
e, ok := t.index[key]
if !ok {
e = t.addValue(key, nil, false)
e.elements = []*Table{}
}
el := newTable(values)
e.elements = append(e.elements, el)
return el
}
// child returns the node of a table-valued key, or nil.
func (t *Table) child(key string) *Table {
if e, ok := t.index[key]; ok {
return e.child
}
return nil
}
// lastElement returns the node of the newest element of an array of tables.
func (t *Table) lastElement(key string) *Table {
if e, ok := t.index[key]; ok && len(e.elements) > 0 {
return e.elements[len(e.elements)-1]
}
return nil
}
// An Entry is one key of a table: the value and the comments around the key.
type Entry struct {
table *Table
key string
inline bool
comments []string
trailing string
// child is the table the value is, and elements are the tables of an array
// of tables; one of them is set only when the value has that shape.
child *Table
elements []*Table
}
// Key returns the key as it was written.
func (e *Entry) Key() string { return e.key }
// Value returns the value the key holds. It is read from the table's map, so
// it stays current if that map is changed.
func (e *Entry) Value() any { return e.table.values[e.key] }
// Inline reports whether the value was written as an inline table, `{…}`.
func (e *Entry) Inline() bool { return e.inline }
// Table returns the table the value is, or nil when it is not a table.
func (e *Entry) Table() *Table { return e.child }
// Elements returns the tables of an array of tables, or nil when the value is
// not one.
func (e *Entry) Elements() []*Table { return e.elements }
// Comments returns the comment lines above the key. Lines carry no leading '#'
// and no surrounding space.
func (e *Entry) Comments() []string { return e.comments }
// SetComments replaces those lines. Each line is written back with a "# " in
// front of it, so a line should not carry one.
func (e *Entry) SetComments(lines []string) { e.comments = lines }
// Trailing returns the comment on the key's own line, without the '#'.
func (e *Entry) Trailing() string { return e.trailing }
// SetTrailing replaces that comment.
func (e *Entry) SetTrailing(line string) { e.trailing = line }
+314
View File
@@ -0,0 +1,314 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: MIT
package interpres
import (
"slices"
"strings"
"testing"
)
// mustEntry returns the entry a table must have, and fails the test when it
// does not.
func mustEntry(t *testing.T, tbl *Table, key string) *Entry {
t.Helper()
e, ok := tbl.Get(key)
if !ok {
t.Fatalf("%q is missing from the table", key)
}
return e
}
func TestDocumentKeepsKeyOrder(t *testing.T) {
doc, err := Parse([]byte(`
b = 1
a = 2
inline = {y = 1, x = 2}
[table]
z = 3
m = 4
`))
if err != nil {
t.Fatalf("parse: %v", err)
}
// The root's keys come back in written order, not sorted.
if got := doc.Root().Keys(); !slices.Equal(got, []string{"b", "a", "inline", "table"}) {
t.Errorf("root keys = %v, want [b a inline table]", got)
}
// So do the keys of an inline table, which the map shape loses.
inline, ok := doc.Root().Get("inline")
if !ok {
t.Fatal("inline is missing from the root")
}
if !inline.Inline() {
t.Error("inline is not marked inline")
}
if got := inline.Table().Keys(); !slices.Equal(got, []string{"y", "x"}) {
t.Errorf("inline keys = %v, want [y x]", got)
}
// And the keys of a table written under a header, which is not inline.
tbl, ok := doc.Root().Get("table")
if !ok {
t.Fatal("table is missing from the root")
}
if tbl.Inline() {
t.Error("table is marked inline")
}
if got := tbl.Table().Keys(); !slices.Equal(got, []string{"z", "m"}) {
t.Errorf("table keys = %v, want [z m]", got)
}
}
func TestDocumentValuesAreTheTree(t *testing.T) {
doc, err := Parse([]byte("n = 7\ns = \"x\"\n\n[t]\nk = true\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if got := mustEntry(t, doc.Root(), "n").Value(); got != int64(7) {
t.Errorf("n = %#v, want int64(7)", got)
}
tbl := mustEntry(t, doc.Root(), "t").Table()
if got := mustEntry(t, tbl, "k").Value(); got != true {
t.Errorf("t.k = %#v, want true", got)
}
// Map gives the tree ParseMap would have returned, the same values.
tree := doc.Map()
if tree["n"] != int64(7) || tree["s"] != "x" {
t.Errorf("Map = %#v", tree)
}
if tree["t"].(map[string]any)["k"] != true {
t.Errorf("Map[t] = %#v", tree["t"])
}
if tbl.Values()["k"] != true {
t.Errorf("t.Values() = %#v", tbl.Values())
}
}
func TestDocumentComments(t *testing.T) {
doc, err := Parse([]byte(`# above b
b = 1 # trailing b
# above the table
[table] # trailing table
# above m
m = 2
# footer
`))
if err != nil {
t.Fatalf("parse: %v", err)
}
b, ok := doc.Root().Get("b")
if !ok {
t.Fatal("b is missing")
}
if got := b.Comments(); !slices.Equal(got, []string{"above b"}) {
t.Errorf("b comments = %q, want [above b]", got)
}
if got := b.Trailing(); got != "trailing b" {
t.Errorf("b trailing = %q, want \"trailing b\"", got)
}
tbl, ok := doc.Root().Get("table")
if !ok {
t.Fatal("table is missing")
}
// A [header] line introduces the table, so the comments around it belong
// to the table node; the entry that names it stays bare.
if got := tbl.Table().Comments(); !slices.Equal(got, []string{"above the table"}) {
t.Errorf("table comments = %q, want [above the table]", got)
}
if got := tbl.Table().Trailing(); got != "trailing table" {
t.Errorf("table trailing = %q, want \"trailing table\"", got)
}
if got := tbl.Comments(); got != nil {
t.Errorf("entry comments = %q, want none", got)
}
m, ok := tbl.Table().Get("m")
if !ok {
t.Fatal("table.m is missing")
}
if got := m.Comments(); !slices.Equal(got, []string{"above m"}) {
t.Errorf("m comments = %q, want [above m]", got)
}
if got := doc.Footer(); !slices.Equal(got, []string{"footer"}) {
t.Errorf("footer = %q, want [footer]", got)
}
}
func TestDocumentCommentsAreWritable(t *testing.T) {
doc, err := Parse([]byte("# above\nk = 1\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
entry, ok := doc.Root().Get("k")
if !ok {
t.Fatal("k is missing")
}
entry.SetComments([]string{"first", "second"})
entry.SetTrailing("beside")
if got := entry.Comments(); !slices.Equal(got, []string{"first", "second"}) {
t.Errorf("comments = %q", got)
}
if got := entry.Trailing(); got != "beside" {
t.Errorf("trailing = %q", got)
}
tbl := doc.Root()
tbl.SetComments([]string{"above the root"})
if got := tbl.Comments(); !slices.Equal(got, []string{"above the root"}) {
t.Errorf("root comments = %q", got)
}
doc.SetFooter([]string{"end"})
if got := doc.Footer(); !slices.Equal(got, []string{"end"}) {
t.Errorf("footer = %q", got)
}
}
func TestDocumentArrayOfTables(t *testing.T) {
doc, err := Parse([]byte(`# first element
[[item]]
a = 1
[[item]]
b = 2 # beside b
`))
if err != nil {
t.Fatalf("parse: %v", err)
}
entry, ok := doc.Root().Get("item")
if !ok {
t.Fatal("item is missing")
}
elems := entry.Elements()
if len(elems) != 2 {
t.Fatalf("elements = %d, want 2", len(elems))
}
if got := elems[0].Keys(); !slices.Equal(got, []string{"a"}) {
t.Errorf("first element keys = %v, want [a]", got)
}
if got := elems[0].Comments(); !slices.Equal(got, []string{"first element"}) {
t.Errorf("first element comments = %q", got)
}
if got := elems[1].Keys(); !slices.Equal(got, []string{"b"}) {
t.Errorf("second element keys = %v, want [b]", got)
}
if got := mustEntry(t, elems[1], "b").Trailing(); got != "beside b" {
t.Errorf("b trailing = %q, want \"beside b\"", got)
}
// The value keeps the map shape the decoder reads.
if _, ok := entry.Value().([]map[string]any); !ok {
t.Errorf("item value = %#v, want []map[string]any", entry.Value())
}
}
func TestDocumentDottedKeysAndValueArrays(t *testing.T) {
doc, err := Parse([]byte("a.b.c = 1\narr = [1, {x = 1}]\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
// A dotted key builds tables, and they are not inline ones.
a, ok := doc.Root().Get("a")
if !ok {
t.Fatal("a is missing")
}
if a.Inline() {
t.Error("a is marked inline")
}
b, ok := a.Table().Get("b")
if !ok {
t.Fatal("a.b is missing")
}
if b.Inline() {
t.Error("a.b is marked inline")
}
if got := b.Table().Keys(); !slices.Equal(got, []string{"c"}) {
t.Errorf("a.b keys = %v, want [c]", got)
}
// An inline table inside a value array keeps its node in the elements
// slice; the scalar before it has none.
arr, ok := doc.Root().Get("arr")
if !ok {
t.Fatal("arr is missing")
}
elems := arr.Elements()
if len(elems) != 2 {
t.Fatalf("elements = %d, want 2", len(elems))
}
if elems[0] != nil {
t.Errorf("elements[0] = %#v, want nil for a scalar", elems[0])
}
if got := elems[1].Keys(); !slices.Equal(got, []string{"x"}) {
t.Errorf("elements[1] keys = %v, want [x]", got)
}
}
func TestDocumentWithoutStatements(t *testing.T) {
doc, err := Parse([]byte("# only a comment\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if got := doc.Root().Keys(); len(got) != 0 {
t.Errorf("keys = %v, want none", got)
}
if got := doc.Footer(); !slices.Equal(got, []string{"only a comment"}) {
t.Errorf("footer = %q, want [only a comment]", got)
}
empty, err := Parse(nil)
if err != nil {
t.Fatalf("parse of nothing: %v", err)
}
if len(empty.Root().Keys()) != 0 || len(empty.Footer()) != 0 {
t.Errorf("empty document = %v / %q", empty.Root().Keys(), empty.Footer())
}
}
func TestParseMapIsTheValueTree(t *testing.T) {
// ParseMap is the path that does not build a document, and it gives the
// tree the decoder reads.
tree, err := ParseMap([]byte("a = 1\n\n[t]\nb = \"x\"\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if tree["a"] != int64(1) {
t.Errorf("a = %#v", tree["a"])
}
if tree["t"].(map[string]any)["b"] != "x" {
t.Errorf("t = %#v", tree["t"])
}
}
func TestMarshalRejectsDocument(t *testing.T) {
// A Document is not a value to marshal: its order and comments would be
// dropped, and a struct walk would silently write nothing at all.
doc, err := Parse([]byte("a = 1\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if _, err := Marshal(doc); err == nil {
t.Fatal("expected an error for a Document")
} else if !strings.Contains(err.Error(), "Map()") {
t.Errorf("err = %v, want it to point at Map()", err)
}
if _, err := Marshal(*doc); err == nil {
t.Fatal("expected an error for a Document value")
}
// The tree marshals, which is the way through.
out, err := Marshal(doc.Map())
if err != nil {
t.Fatalf("marshal of the tree: %v", err)
}
if want := "a = 1\n"; string(out) != want {
t.Errorf("output = %q, want %q", out, want)
}
}
+385 -41
View File
@@ -6,6 +6,7 @@ package interpres
import (
"bytes"
"context"
"encoding"
"errors"
"fmt"
"maps"
@@ -22,21 +23,56 @@ var (
localDateTimeType = reflect.TypeFor[LocalDateTime]()
localDateType = reflect.TypeFor[LocalDate]()
localTimeType = reflect.TypeFor[LocalTime]()
offsetDateTimeType = reflect.TypeFor[OffsetDateTime]()
timeGoType = reflect.TypeFor[time.Time]()
durationType = reflect.TypeFor[time.Duration]()
textMarshalerType = reflect.TypeFor[encoding.TextMarshaler]()
)
// inlineLimit is the column past which an inline table is written across
// lines. TOML 1.1 lets an inline table carry newlines and a trailing comma, so
// a long one stays readable instead of running off the line.
const inlineLimit = 100
// noInlineBreak is the limit a measuring encoder carries, high enough that the
// form it renders is always the single-line one.
const noInlineBreak = 1 << 30
// encoder produces a TOML document from a Go value via a small intermediate
// representation that preserves the order in which fields were declared.
type encoder struct {
buf bytes.Buffer
ctx context.Context
opts Encoder
// inlineDepth is the nesting level inside inline tables, which decides
// their indentation.
inlineDepth int
// limit is the column at which an inline table is broken; only a
// measuring encoder raises it.
limit int
}
func newEncoder() *encoder { return &encoder{} }
func newEncoder() *encoder { return &encoder{limit: inlineLimit} }
// flat returns an encoder that measures a value by rendering it on one line,
// so a caller can decide which form to write before writing it.
func (e *encoder) flat() *encoder {
return &encoder{ctx: e.ctx, opts: e.opts, limit: noInlineBreak}
}
func (e *encoder) bytes() []byte { return e.buf.Bytes() }
// column reports how many bytes the current line already holds, so a form can
// be measured against the limit before it is written.
func (e *encoder) column() int {
if i := bytes.LastIndexByte(e.buf.Bytes(), '\n'); i >= 0 {
return e.buf.Len() - i - 1
}
return e.buf.Len()
}
func (e *encoder) checkCtx() error {
if e.ctx == nil {
return nil
@@ -50,6 +86,15 @@ func (e *encoder) encode(v any) error {
if err := e.checkCtx(); err != nil {
return err
}
switch x := v.(type) {
case *Document:
if x == nil {
return fmt.Errorf("interpres: cannot marshal nil value")
}
return fmt.Errorf("interpres: cannot marshal a Document; marshal its Map() to write the values")
case Document:
return fmt.Errorf("interpres: cannot marshal a Document; marshal its Map() to write the values")
}
rv := reflect.ValueOf(v)
if !rv.IsValid() {
return fmt.Errorf("interpres: cannot marshal nil value")
@@ -307,8 +352,7 @@ func buildMapDoc(v reflect.Value, doc *tomlDoc, ctx string) error {
var errNilMarshalTOML = errors.New("MarshalTOML returned a nil value")
func addField(doc *tomlDoc, name string, v reflect.Value, ctx string) error {
if v.CanInterface() {
if m, ok := v.Interface().(Marshaler); ok {
if m, ok := marshalerOf(v); ok {
mv, err := m.MarshalTOML()
if err != nil {
return &EncodeError{Path: joinKey(ctx, name), Err: err}
@@ -318,6 +362,14 @@ func addField(doc *tomlDoc, name string, v reflect.Value, ctx string) error {
}
v = reflect.ValueOf(mv)
}
// A type that renders itself as text becomes a TOML string, whether it is
// a scalar kind or a struct.
s, isText, err := textValue(v)
if err != nil {
return &EncodeError{Path: joinKey(ctx, name), Err: err}
}
if isText {
return doc.appendScalar(name, s, ctx)
}
v = followPtr(v)
if !v.IsValid() {
@@ -387,12 +439,29 @@ func addArrayValue(doc *tomlDoc, name string, v reflect.Value, ctx string) error
return doc.appendScalar(name, []any{}, ctx)
}
// Every element is resolved through MarshalTOML first, so an element that
// renders itself as a scalar, a table or a value array is classified by
// what it produces rather than by its Go kind, and its method runs once.
elems := make([]reflect.Value, n)
for i := range n {
if i%ctxCheckInterval == 0 {
if err := doc.checkCtx(); err != nil {
return err
}
}
ev, err := resolveElement(v.Index(i), fmt.Sprintf("%s[%d]", joinKey(ctx, name), i))
if err != nil {
return err
}
elems[i] = ev
}
// An array keeps the [[header]] form only when every element is a table.
// TOML lets one array mix tables with scalars, and that mix renders as a
// value array with the table elements written inline.
allTables := true
for i := range n {
if !isTableElementValue(v.Index(i)) {
for _, ev := range elems {
if !ev.IsValid() || !isTableElementValue(ev) {
allTables = false
break
}
@@ -406,16 +475,12 @@ func addArrayValue(doc *tomlDoc, name string, v reflect.Value, ctx string) error
}
if allTables {
subs := make([]*tomlDoc, n)
for i := range n {
for i, ev := range elems {
if i%ctxCheckInterval == 0 {
if err := doc.checkCtx(); err != nil {
return err
}
}
ev := followPtr(v.Index(i))
if !ev.IsValid() {
return &EncodeError{Path: fmt.Sprintf("%s[%d]", joinKey(ctx, name), i), Err: errors.New("nil element")}
}
sub := &tomlDoc{ctx: doc.ctx, opts: doc.opts}
switch ev.Kind() {
case reflect.Struct:
@@ -441,29 +506,12 @@ func addArrayValue(doc *tomlDoc, name string, v reflect.Value, ctx string) error
// Value array. Table elements normalise to map[string]any and the emitter
// writes them as inline tables.
items := make([]any, n)
for i := range n {
for i, ev := range elems {
if i%ctxCheckInterval == 0 {
if err := doc.checkCtx(); err != nil {
return err
}
}
ev := followPtr(v.Index(i))
if !ev.IsValid() {
return &EncodeError{Path: fmt.Sprintf("%s[%d]", joinKey(ctx, name), i), Err: errors.New("nil element")}
}
if ev.CanInterface() {
if m, ok := ev.Interface().(Marshaler); ok {
mv, err := m.MarshalTOML()
if err != nil {
return &EncodeError{Path: fmt.Sprintf("%s[%d]", joinKey(ctx, name), i), Err: err}
}
if mv == nil {
return &EncodeError{Path: fmt.Sprintf("%s[%d]", joinKey(ctx, name), i), Err: errNilMarshalTOML}
}
ev = reflect.ValueOf(mv)
ev = followPtr(ev)
}
}
val, err := normaliseValue(ev)
if err != nil {
return &EncodeError{Path: fmt.Sprintf("%s[%d]", joinKey(ctx, name), i), Err: err}
@@ -473,6 +521,50 @@ func addArrayValue(doc *tomlDoc, name string, v reflect.Value, ctx string) error
return doc.appendScalar(name, items, ctx)
}
// marshalerOf finds the Marshaler a value carries: on the value itself, or on
// its address, so a pointer-receiver MarshalTOML is found on an addressable
// struct field or slice element, exactly as textMarshalerOf finds MarshalText.
func marshalerOf(v reflect.Value) (Marshaler, bool) {
if !v.CanInterface() {
return nil, false
}
if m, ok := v.Interface().(Marshaler); ok {
return m, true
}
if v.CanAddr() {
if m, ok := v.Addr().Interface().(Marshaler); ok {
return m, true
}
}
return nil, false
}
// resolveElement looks through pointers and runs MarshalTOML, so an array
// element is classified by what its method produces. path names the element,
// for the errors the method can raise.
func resolveElement(v reflect.Value, path string) (reflect.Value, error) {
ev := followPtr(v)
if !ev.IsValid() {
return ev, &EncodeError{Path: path, Err: errors.New("nil element")}
}
m, ok := marshalerOf(ev)
if !ok {
return ev, nil
}
mv, err := m.MarshalTOML()
if err != nil {
return reflect.Value{}, &EncodeError{Path: path, Err: err}
}
if mv == nil {
return reflect.Value{}, &EncodeError{Path: path, Err: errNilMarshalTOML}
}
ev = followPtr(reflect.ValueOf(mv))
if !ev.IsValid() {
return ev, &EncodeError{Path: path, Err: errors.New("nil element")}
}
return ev, nil
}
// normaliseValue converts a reflect.Value into one of the canonical scalar or
// nested-array representations the emitter understands. Slices and arrays are
// recursively normalised so that nested arrays (e.g. [][]int) work.
@@ -484,8 +576,7 @@ func normaliseValue(v reflect.Value) (any, error) {
if v.Kind() == reflect.Interface {
return nil, fmt.Errorf("cannot encode nil value")
}
if v.CanInterface() {
if m, ok := v.Interface().(Marshaler); ok {
if m, ok := marshalerOf(v); ok {
mv, err := m.MarshalTOML()
if err != nil {
return nil, err
@@ -493,13 +584,33 @@ func normaliseValue(v reflect.Value) (any, error) {
if mv == nil {
return nil, errNilMarshalTOML
}
// The result is normalised like any other value, so a method may return
// a duration, a defined type or another Marshaler. A result of the
// receiver's own type is written as it is, because recursing into it
// would never end.
if rv := reflect.ValueOf(mv); rv.Type() != v.Type() {
return normaliseValue(rv)
}
return mv, nil
}
}
// The datetime structs are TOML scalars; the emitter renders each of them.
if t := v.Type(); t == timeGoType || isLocalDateType(t) {
if isScalarStruct(v.Type()) {
return v.Interface(), nil
}
// TOML has no duration type, so a duration goes out in its canonical Go
// form, the shape it comes back in.
if v.Type() == durationType {
return time.Duration(v.Int()).String(), nil
}
// A type that renders itself as text becomes a TOML string, scalar kinds
// and structs alike.
s, isText, err := textValue(v)
if err != nil {
return nil, err
}
if isText {
return s, nil
}
switch v.Kind() {
case reflect.String:
return v.String(), nil
@@ -564,19 +675,84 @@ func followPtr(v reflect.Value) reflect.Value {
}
// isScalarStruct reports whether t is a struct type that the encoder treats
// as a TOML scalar (time.Time, LocalDateTime, LocalDate, LocalTime).
// as a TOML scalar: time.Time and the four date-time wrappers.
func isScalarStruct(t reflect.Type) bool {
return t == timeGoType || isLocalDateType(t)
return t == timeGoType || t == offsetDateTimeType || isLocalDateType(t)
}
func isLocalDateType(t reflect.Type) bool {
return t == localDateTimeType || t == localDateType || t == localTimeType
}
// isDateTimeType reports whether t is one of the date-time types, which the
// encoder emits as bare atoms. Pointers are looked through. The types carry
// time.Time's text methods through an embedded field, and the atom form takes
// precedence over them.
func isDateTimeType(t reflect.Type) bool {
for t.Kind() == reflect.Pointer {
t = t.Elem()
}
return isScalarStruct(t)
}
// isTextMarshalerType reports whether t or *t implements
// encoding.TextMarshaler. An array of such values stays a value array, because
// each element's TOML form is a string.
func isTextMarshalerType(t reflect.Type) bool {
if isDateTimeType(t) {
return false
}
return t.Implements(textMarshalerType) || reflect.PointerTo(t).Implements(textMarshalerType)
}
// textValue returns the string a value renders itself as through
// encoding.TextMarshaler. The date-time types are excluded, because their
// embedded time.Time would answer with an RFC 3339 string where the TOML form
// is a bare timestamp. A nil pointer offers no text and is left to the ordinary
// nil handling, which omits the field.
func textValue(v reflect.Value) (string, bool, error) {
for v.Kind() == reflect.Interface && !v.IsNil() {
v = v.Elem()
}
if !v.IsValid() || isDateTimeType(v.Type()) {
return "", false, nil
}
if v.Kind() == reflect.Pointer && v.IsNil() {
return "", false, nil
}
m, ok := textMarshalerOf(v)
if !ok {
return "", false, nil
}
b, err := m.MarshalText()
if err != nil {
return "", true, err
}
return string(b), true, nil
}
// textMarshalerOf finds the encoding.TextMarshaler for v: on the value itself,
// or on its address, so a pointer-receiver MarshalText is found on an
// addressable struct field.
func textMarshalerOf(v reflect.Value) (encoding.TextMarshaler, bool) {
if !v.CanInterface() {
return nil, false
}
if m, ok := v.Interface().(encoding.TextMarshaler); ok {
return m, true
}
if v.CanAddr() {
if m, ok := v.Addr().Interface().(encoding.TextMarshaler); ok {
return m, true
}
}
return nil, false
}
func isTableElementType(t reflect.Type) bool {
switch t.Kind() {
case reflect.Struct:
return !isScalarStruct(t)
return !isScalarStruct(t) && !isTextMarshalerType(t)
case reflect.Map:
return t.Key().Kind() == reflect.String
}
@@ -618,7 +794,20 @@ func (e *encoder) emitDoc(doc *tomlDoc, prefix []string) error {
return err
}
}
// An inlined sub-table is a value line, so it has to precede every
// header of this document: a line written after a [header] would be
// read back as part of that table.
headers := make([]entry, 0, len(tables))
for _, t := range tables {
inlined, err := e.writeInlineSubTableIfSmall(t.key, t.doc)
if err != nil {
return err
}
if !inlined {
headers = append(headers, t)
}
}
for _, t := range headers {
path := append(append([]string{}, prefix...), t.key)
e.writeBlankLine()
e.buf.WriteByte('[')
@@ -659,6 +848,13 @@ func (e *encoder) emitDoc(doc *tomlDoc, prefix []string) error {
return err
}
case entryTable:
inlined, err := e.writeInlineSubTableIfSmall(ent.key, ent.doc)
if err != nil {
return err
}
if inlined {
continue
}
path := append(append([]string{}, prefix...), ent.key)
e.writeBlankLine()
e.buf.WriteByte('[')
@@ -800,7 +996,10 @@ func (e *encoder) writeValue(val any) error {
case float64:
return e.writeFloat(v)
case time.Time:
e.buf.WriteString(v.Format(time.RFC3339Nano))
e.buf.WriteString(offsetString(v))
return nil
case OffsetDateTime:
e.buf.WriteString(v.String())
return nil
case LocalDateTime:
e.buf.WriteString(v.String())
@@ -824,7 +1023,7 @@ func (e *encoder) writeValue(val any) error {
e.buf.WriteByte(']')
return nil
case map[string]any:
return e.writeInlineTable(v)
return e.writeInlineMap(v)
case nil:
return fmt.Errorf("interpres: cannot encode nil value")
default:
@@ -832,10 +1031,24 @@ func (e *encoder) writeValue(val any) error {
}
}
// writeInlineTable renders m as a TOML inline table with sorted keys, the
// order buildMapDoc uses for header tables. It backs the table elements of a
// value array, where the [[header]] form is not available.
func (e *encoder) writeInlineTable(m map[string]any) error {
// writeInlineMap renders m as a TOML inline table, on one line when it fits
// there and across lines when it does not.
func (e *encoder) writeInlineMap(m map[string]any) error {
flat := e.flat()
if err := flat.writeInlineMapFlat(m); err != nil {
return err
}
if e.column()+flat.buf.Len() <= e.limit {
e.buf.Write(flat.buf.Bytes())
return nil
}
return e.writeInlineMapMultiline(m)
}
// writeInlineMapFlat renders m as a single-line inline table with sorted keys,
// the order buildMapDoc uses for header tables. It backs the table elements of
// a value array, where the [[header]] form is not available.
func (e *encoder) writeInlineMapFlat(m map[string]any) error {
keys := slices.Sorted(maps.Keys(m))
e.buf.WriteByte('{')
for i, k := range keys {
@@ -854,6 +1067,137 @@ func (e *encoder) writeInlineTable(m map[string]any) error {
return nil
}
// writeInlineMapMultiline renders m with one entry per line and a trailing
// comma, the form TOML 1.1 allows for an inline table too long for one line.
func (e *encoder) writeInlineMapMultiline(m map[string]any) error {
keys := slices.Sorted(maps.Keys(m))
e.buf.WriteString("{\n")
e.inlineDepth++
for _, k := range keys {
e.writeInlineIndent()
if err := e.writeKey(k); err != nil {
return err
}
e.buf.WriteString(" = ")
if err := e.writeValue(m[k]); err != nil {
return err
}
e.buf.WriteString(",\n")
}
e.inlineDepth--
e.writeInlineIndent()
e.buf.WriteByte('}')
return nil
}
// writeInlineIndent writes one tab per inline-table nesting level.
func (e *encoder) writeInlineIndent() {
for range e.inlineDepth {
e.buf.WriteByte('\t')
}
}
// errInlineArrayOfTables reports an attempt to render an array of tables
// inline, which has no form that keeps the value's type.
var errInlineArrayOfTables = errors.New("interpres: an array of tables has no inline form")
// inlinableDoc reports whether doc can be written as an inline table without
// changing the type of any value: scalars, value arrays and further sub-tables
// are fine, while an array of tables is not, because its inline form would
// re-parse as a value array.
func inlinableDoc(doc *tomlDoc) bool {
for _, ent := range doc.entries {
switch ent.kind {
case entryArray:
return false
case entryTable:
if !inlinableDoc(ent.doc) {
return false
}
}
}
return true
}
// writeInlineDocEntry writes one "key = value" binding of an inline table,
// without the separator that follows it.
func (e *encoder) writeInlineDocEntry(ent entry) error {
if err := e.writeKey(ent.key); err != nil {
return err
}
e.buf.WriteString(" = ")
switch ent.kind {
case entryTable:
return e.writeInlineDoc(ent.doc)
case entryArray:
return errInlineArrayOfTables
default:
return e.writeValue(ent.val)
}
}
// writeInlineDoc renders doc as a single-line inline table in entry order, the
// order the fields were declared in.
func (e *encoder) writeInlineDoc(doc *tomlDoc) error {
e.buf.WriteByte('{')
for i, ent := range doc.entries {
if i > 0 {
e.buf.WriteString(", ")
}
if err := e.writeInlineDocEntry(ent); err != nil {
return err
}
}
e.buf.WriteByte('}')
return nil
}
// writeInlineDocMultiline renders doc with one entry per line and a trailing
// comma, the form TOML 1.1 allows for an inline table too long for one line.
func (e *encoder) writeInlineDocMultiline(doc *tomlDoc) error {
e.buf.WriteString("{\n")
e.inlineDepth++
for _, ent := range doc.entries {
e.writeInlineIndent()
if err := e.writeInlineDocEntry(ent); err != nil {
return err
}
e.buf.WriteString(",\n")
}
e.inlineDepth--
e.writeInlineIndent()
e.buf.WriteByte('}')
return nil
}
// writeInlineSubTableIfSmall writes "key = {…}" for a sub-table whose
// single-line rendering fits the compact threshold, and reports whether it did
// so. An array of tables is never inlined, because its inline form would
// re-parse as a value array and change the value's Go type.
func (e *encoder) writeInlineSubTableIfSmall(name string, doc *tomlDoc) (bool, error) {
if e.opts.inlineTablesAt <= 0 || !inlinableDoc(doc) {
return false, nil
}
flat := e.flat()
if err := flat.writeInlineDoc(doc); err != nil {
return false, err
}
if flat.buf.Len() > e.opts.inlineTablesAt {
return false, nil
}
if err := e.writeKey(name); err != nil {
return false, err
}
e.buf.WriteString(" = ")
if e.column()+flat.buf.Len() <= e.limit {
e.buf.Write(flat.buf.Bytes())
} else if err := e.writeInlineDocMultiline(doc); err != nil {
return false, err
}
e.buf.WriteByte('\n')
return true, nil
}
func (e *encoder) writeStringVal(s string) error {
if e.opts.literalMultilineAt > 0 && strings.ContainsRune(s, '\n') &&
len(s) >= e.opts.literalMultilineAt && canBeLiteralMultiline(s) {
+596 -22
View File
@@ -8,6 +8,7 @@ import (
"context"
"errors"
"math"
"net"
"reflect"
"strings"
"testing"
@@ -288,7 +289,7 @@ func TestEncoderLiteralMultilineFallsBackWhenUnsafe(t *testing.T) {
if !bytes.HasPrefix(out, []byte("s = \"")) {
t.Errorf("%s: expected the basic quoted form, got:\n%s", c.name, out)
}
re, err := Parse(out)
re, err := ParseMap(out)
if err != nil {
t.Errorf("%s: re-parse: %v\ndoc:\n%s", c.name, err, out)
continue
@@ -325,7 +326,7 @@ func TestMarshalerReturningTime(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "m = 2026-06-26T10:00:00Z\n"
want := "m = 2026-06-26T10:00Z\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
@@ -447,7 +448,7 @@ func TestMarshalDuplicateKeyResolvesToOneField(t *testing.T) {
if want := "name = \"outer\"\n"; string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
if _, err := Parse(out); err != nil {
if _, err := ParseMap(out); err != nil {
t.Errorf("re-parse: %v\ndoc:\n%s", err, out)
}
}
@@ -465,7 +466,7 @@ func TestMarshalEmbeddedScalarStruct(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "name = \"x\"\ns = 2026-06-26T00:00:00\n"
want := "name = \"x\"\ns = 2026-06-26T00:00\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
@@ -540,7 +541,7 @@ func TestMarshalDateTime(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "offset = 2026-06-26T10:00:00Z\nlocal = 2026-06-26T07:32:00\nday = 2026-06-26\nclock = 07:32:00\n"
want := "offset = 2026-06-26T10:00Z\nlocal = 2026-06-26T07:32\nday = 2026-06-26\nclock = 07:32\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
@@ -609,7 +610,7 @@ func TestMarshalMixedArrayWithInlineTable(t *testing.T) {
// Parse accepts a mixed array (TOML allows any value kinds in one array),
// so Marshal of the parsed tree must re-emit it. The table element has no
// header form inside a value array and renders inline.
tree, err := Parse([]byte("arr = [1, {a = 2}, \"x\"]\n"))
tree, err := ParseMap([]byte("arr = [1, {a = 2}, \"x\"]\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -621,7 +622,7 @@ func TestMarshalMixedArrayWithInlineTable(t *testing.T) {
if string(out) != want {
t.Fatalf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
re, err := Parse(out)
re, err := ParseMap(out)
if err != nil {
t.Fatalf("re-parse: %v", err)
}
@@ -639,7 +640,7 @@ func TestMarshalValueArrayOfTablesStaysInline(t *testing.T) {
"a = [{x = 1}, {x = 2}]\n",
"b = [{x = 1}, 2, \"three\"]\n",
} {
tree, err := Parse([]byte(doc))
tree, err := ParseMap([]byte(doc))
if err != nil {
t.Fatalf("%s: parse: %v", doc, err)
}
@@ -650,7 +651,7 @@ func TestMarshalValueArrayOfTablesStaysInline(t *testing.T) {
if bytes.HasPrefix(out, []byte("[[")) {
t.Errorf("%s: emitted the [[header]] form for a value array:\n%s", doc, out)
}
re, err := Parse(out)
re, err := ParseMap(out)
if err != nil {
t.Fatalf("%s: re-parse: %v\ndoc:\n%s", doc, err, out)
}
@@ -687,14 +688,14 @@ func TestMarshalInlineTableWithDatetime(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "mix = [1979-05-27T07:32:00Z, {t = 1979-05-27T07:32:00}]\n"
want := "mix = [1979-05-27T07:32Z, {t = 1979-05-27T07:32}]\n"
if string(out) != want {
t.Fatalf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalArrayOfTablesStaysHeaderForm(t *testing.T) {
tree, err := Parse([]byte("[[items]]\nname = \"a\"\n\n[[items]]\nname = \"b\"\n"))
tree, err := ParseMap([]byte("[[items]]\nname = \"a\"\n\n[[items]]\nname = \"b\"\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -721,7 +722,7 @@ func TestMarshalFloatExponentNoLeadingZero(t *testing.T) {
}
// Parse to check the output is valid TOML (for a strict parser that
// rejects leading zeros in exponents).
if _, err := Parse(out); err != nil {
if _, err := ParseMap(out); err != nil {
t.Fatalf("marshalled output is not valid TOML:\n%s\nerror: %v", out, err)
}
if string(out) != "large = 1e+6\nsmall = 1e-5\n" {
@@ -906,7 +907,7 @@ func TestMarshalTagOptionOmitZero(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
want = "name = \"x\"\ncount = 1\nratio = 0.5\nwhen = 2026-09-17T12:00:00Z\nalways = \"kept\"\n\n[server]\nhost = \"h\"\n"
want = "name = \"x\"\ncount = 1\nratio = 0.5\nwhen = 2026-09-17T12:00Z\nalways = \"kept\"\n\n[server]\nhost = \"h\"\n"
if string(out) != want {
t.Fatalf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
@@ -1160,7 +1161,7 @@ func TestMarshalThenParseRoundTrip(t *testing.T) {
if err != nil {
t.Fatalf("marshal: %v", err)
}
tree1, err := Parse(out)
tree1, err := ParseMap(out)
if err != nil {
t.Fatalf("parse of marshalled: %v\noutput:\n%s", err, out)
}
@@ -1195,8 +1196,10 @@ qty = 2
[meta]
created = 2026-06-26T10:00:00Z
mixed = [1, {n = 1, name = "a value long enough to push this line well past the one hundred column limit"}]
`)
tree1, err := Parse(src)
tree1, err := ParseMap(src)
if err != nil {
t.Fatalf("parse src: %v", err)
}
@@ -1204,7 +1207,7 @@ created = 2026-06-26T10:00:00Z
if err != nil {
t.Fatalf("marshal: %v", err)
}
tree2, err := Parse(out)
tree2, err := ParseMap(out)
if err != nil {
t.Fatalf("re-parse marshalled: %v\noutput:\n%s", err, out)
}
@@ -1257,24 +1260,35 @@ func TestLocalDateString(t *testing.T) {
}
func TestLocalDateTimeString(t *testing.T) {
// The rendering drops zero seconds and the trailing zeros of a fraction,
// which TOML 1.1 allows and which keeps a value written without seconds
// written without them.
ldt := LocalDateTime{Time: time.Date(1979, 5, 27, 7, 32, 0, 0, time.UTC)}
if got := ldt.String(); got != "1979-05-27T07:32:00" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32:00", got)
if got := ldt.String(); got != "1979-05-27T07:32" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32", got)
}
ldt2 := LocalDateTime{Time: time.Date(1979, 5, 27, 7, 32, 0, 5, time.UTC)}
if got := ldt2.String(); got != "1979-05-27T07:32:00.000000005" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32:00.000000005", got)
}
ldt3 := LocalDateTime{Time: time.Date(1979, 5, 27, 7, 32, 0, 500, time.UTC)}
if got := ldt3.String(); got != "1979-05-27T07:32:00.000000500" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32:00.000000500", got)
if got := ldt3.String(); got != "1979-05-27T07:32:00.0000005" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32:00.0000005", got)
}
ldt4 := LocalDateTime{Time: time.Date(1979, 5, 27, 7, 32, 30, 500000000, time.UTC)}
if got := ldt4.String(); got != "1979-05-27T07:32:30.5" {
t.Errorf("LocalDateTime.String() = %q, want 1979-05-27T07:32:30.5", got)
}
}
func TestLocalTimeString(t *testing.T) {
lt := LocalTime{Time: time.Date(0, 1, 1, 7, 32, 0, 0, time.UTC)}
if got := lt.String(); got != "07:32:00" {
t.Errorf("LocalTime.String() = %q, want 07:32:00", got)
if got := lt.String(); got != "07:32" {
t.Errorf("LocalTime.String() = %q, want 07:32", got)
}
lt2 := LocalTime{Time: time.Date(0, 1, 1, 7, 32, 15, 250000000, time.UTC)}
if got := lt2.String(); got != "07:32:15.25" {
t.Errorf("LocalTime.String() = %q, want 07:32:15.25", got)
}
}
@@ -1372,3 +1386,563 @@ func TestEncodeErrorHeterogeneousArrayPath(t *testing.T) {
t.Fatalf("Path = %q, want %q", ee.Path, "items[0]")
}
}
// --- encoding.TextMarshaler and time.Duration ------------------------------
// textTag is a value-receiver encoding.TextMarshaler, so the encoder finds the
// method on the value itself.
type textTag string
func (t textTag) MarshalText() ([]byte, error) { return []byte("tag:" + string(t)), nil }
// textPointer carries MarshalText on the pointer receiver only, so the encoder
// has to look at the address of an addressable field.
type textPointer struct{ V string }
func (t *textPointer) MarshalText() ([]byte, error) { return []byte(strings.ToUpper(t.V)), nil }
// textAndTOML implements both encoding interfaces; the TOML method wins.
type textAndTOML struct{}
func (textAndTOML) MarshalTOML() (any, error) { return "toml", nil }
func (textAndTOML) MarshalText() ([]byte, error) { return []byte("text"), nil }
// brokenText fails the marshal from MarshalText.
type brokenText struct{}
func (brokenText) MarshalText() ([]byte, error) { return nil, errors.New("text boom") }
// notUTF8 renders bytes that no TOML string can carry.
type notUTF8 struct{}
func (notUTF8) MarshalText() ([]byte, error) { return []byte{0xff, 0xfe}, nil }
// textTagBoth renders itself with a prefix and strips it again on decode, so
// the round trip through a TOML string is lossless.
type textTagBoth string
func (t textTagBoth) MarshalText() ([]byte, error) { return []byte("tag:" + string(t)), nil }
func (t *textTagBoth) UnmarshalText(text []byte) error {
trimmed, ok := strings.CutPrefix(string(text), "tag:")
if !ok {
return errors.New("textTagBoth: missing the tag prefix")
}
*t = textTagBoth(trimmed)
return nil
}
func TestMarshalTextValues(t *testing.T) {
// The pointer receiver is reachable only through an addressable field, so
// the whole value is marshalled through a pointer here.
type Cfg struct {
IP net.IP `toml:"ip"`
Duration time.Duration `toml:"duration"`
Tag textTag `toml:"tag"`
Pointer textPointer `toml:"pointer"`
Both textAndTOML `toml:"both"`
}
out, err := Marshal(&Cfg{
IP: net.IPv4(192, 0, 2, 1),
Duration: 90 * time.Minute,
Tag: "x",
Pointer: textPointer{V: "abc"},
})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "ip = \"192.0.2.1\"\nduration = \"1h30m0s\"\ntag = \"tag:x\"\npointer = \"ABC\"\nboth = \"toml\"\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalTextValuesInContainers(t *testing.T) {
// Slice elements are addressable, so a pointer-receiver MarshalText is used
// there too, and an array of such values stays a value array: each element's
// TOML form is a string, so the [[header]] form cannot carry it.
type Cfg struct {
Map map[string]net.IP `toml:"map"`
Durs []time.Duration `toml:"durs"`
Ptrs []textPointer `toml:"ptrs"`
Empty []textPointer `toml:"empty"`
}
out, err := Marshal(Cfg{
Map: map[string]net.IP{"a": net.IPv4(10, 0, 0, 1)},
Durs: []time.Duration{0, 250 * time.Millisecond},
Ptrs: []textPointer{{V: "a"}, {V: "b"}},
Empty: []textPointer{},
})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "durs = [\"0s\", \"250ms\"]\nptrs = [\"A\", \"B\"]\nempty = []\n\n[map]\na = \"10.0.0.1\"\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalTextLeavesDateTimesAlone(t *testing.T) {
// The four date-time types carry time.Time's text methods through an
// embedded field; their TOML form is a bare atom, never a quoted string.
stamp := time.Date(2026, 6, 26, 10, 0, 0, 0, time.UTC)
type Cfg struct {
Stamp time.Time `toml:"stamp"`
Ptr *time.Time `toml:"ptr"`
Day LocalDate `toml:"day"`
At LocalDateTime `toml:"at"`
Clock LocalTime `toml:"clock"`
}
out, err := Marshal(&Cfg{
Stamp: stamp,
Ptr: &stamp,
Day: LocalDate{time.Date(1979, 5, 27, 0, 0, 0, 0, time.UTC)},
At: LocalDateTime{time.Date(1979, 5, 27, 7, 32, 0, 0, time.UTC)},
Clock: LocalTime{time.Date(0, 1, 1, 7, 32, 0, 0, time.UTC)},
})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "stamp = 2026-06-26T10:00Z\nptr = 2026-06-26T10:00Z\nday = 1979-05-27\nat = 1979-05-27T07:32\nclock = 07:32\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalTextNilPointerOmitted(t *testing.T) {
type Cfg struct {
P *textPointer `toml:"p"`
K string `toml:"k"`
}
out, err := Marshal(&Cfg{K: "x"})
if err != nil {
t.Fatalf("marshal: %v", err)
}
if want := "k = \"x\"\n"; string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalTextErrorCarriesPath(t *testing.T) {
type Inner struct {
F brokenText `toml:"f"`
}
type Cfg struct {
Inner Inner `toml:"inner"`
}
_, err := Marshal(Cfg{})
if err == nil {
t.Fatal("expected an error from MarshalText")
}
if !strings.Contains(err.Error(), "text boom") {
t.Errorf("err = %v, want substring \"text boom\"", err)
}
ee, ok := errors.AsType[*EncodeError](err)
if !ok {
t.Fatalf("expected an *EncodeError, got %T: %v", err, err)
}
if ee.Path != "inner.f" {
t.Fatalf("Path = %q, want %q", ee.Path, "inner.f")
}
}
func TestMarshalTextRejectsInvalidUTF8(t *testing.T) {
// A TOML string holds UTF-8 only, so text that is not gets an error rather
// than replacement characters.
_, err := Marshal(struct {
V notUTF8 `toml:"v"`
}{})
if err == nil {
t.Fatal("expected an error for text that is not valid UTF-8")
}
if !strings.Contains(err.Error(), "UTF-8") {
t.Errorf("err = %v, want a UTF-8 message", err)
}
}
func TestMarshalTextValuesRoundTrip(t *testing.T) {
type Cfg struct {
Duration time.Duration `toml:"duration"`
IP net.IP `toml:"ip"`
Tag textTagBoth `toml:"tag"`
}
in := Cfg{Duration: 90 * time.Minute, IP: net.IPv4(198, 51, 100, 7), Tag: "y"}
out, err := Marshal(&in)
if err != nil {
t.Fatalf("marshal: %v", err)
}
var back Cfg
if err := Unmarshal(out, &back); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if back.Duration != in.Duration {
t.Errorf("Duration = %v, want %v", back.Duration, in.Duration)
}
if !back.IP.Equal(in.IP) {
t.Errorf("IP = %v, want %v", back.IP, in.IP)
}
if back.Tag != in.Tag {
t.Errorf("Tag = %q, want %q", back.Tag, in.Tag)
}
}
// --- TOML 1.1 output forms -------------------------------------------------
func TestMarshalDateTimeRendering(t *testing.T) {
// The seconds are written only when the value carries them, and a fraction
// drops its trailing zeros. Both are the same value either way; the shorter
// form is the one TOML 1.1 allows.
base := time.Date(2026, 6, 26, 10, 0, 0, 0, time.UTC)
cases := []struct {
name string
val any
want string
}{
{"offset-zero-seconds", base, "v = 2026-06-26T10:00Z\n"},
{"offset-wrapper", OffsetDateTime{Time: base}, "v = 2026-06-26T10:00Z\n"},
{"offset-seconds", base.Add(30 * time.Second), "v = 2026-06-26T10:00:30Z\n"},
{"offset-fraction", base.Add(500 * time.Millisecond), "v = 2026-06-26T10:00:00.5Z\n"},
{"offset-zone", time.Date(2026, 6, 26, 10, 0, 0, 0, time.FixedZone("", -7*3600)), "v = 2026-06-26T10:00-07:00\n"},
{"local-zero-seconds", LocalDateTime{Time: base}, "v = 2026-06-26T10:00\n"},
{"local-fraction", LocalDateTime{Time: base.Add(2500 * time.Millisecond)}, "v = 2026-06-26T10:00:02.5\n"},
{"date", LocalDate{Time: base}, "v = 2026-06-26\n"},
{"time-zero-seconds", LocalTime{Time: base}, "v = 10:00\n"},
{"time-seconds", LocalTime{Time: base.Add(15 * time.Second)}, "v = 10:00:15\n"},
{"time-nanoseconds", LocalTime{Time: base.Add(123456789 * time.Nanosecond)}, "v = 10:00:00.123456789\n"},
}
for _, c := range cases {
out, err := Marshal(map[string]any{"v": c.val})
if err != nil {
t.Errorf("%s: marshal: %v", c.name, err)
continue
}
if string(out) != c.want {
t.Errorf("%s: output mismatch:\ngot: %q\nwant: %q", c.name, out, c.want)
}
}
}
func TestMarshalInlineTableBreaksWhenLong(t *testing.T) {
// A table element of a value array is written inline; a long one carries
// newlines and a trailing comma instead of running past the line limit,
// which TOML 1.1 allows an inline table to do.
const long = "a-very-long-value-that-pushes-the-line-well-past-the-one-hundred-column-limit"
type Cfg struct {
Arr []any `toml:"arr"`
}
out, err := Marshal(Cfg{Arr: []any{int64(1), map[string]any{"n": int64(1), "name": long}}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "arr = [1, {\n\tn = 1,\n\tname = \"" + long + "\",\n}]\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
// The same values without the long string stay on one line.
out, err = Marshal(Cfg{Arr: []any{int64(1), map[string]any{"n": int64(1), "name": "short"}}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
if want := "arr = [1, {n = 1, name = \"short\"}]\n"; string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
// The broken form parses back to the same tree.
tree, err := ParseMap(out)
if err != nil {
t.Fatalf("parse of the encoder output: %v", err)
}
if got := len(tree["arr"].([]any)); got != 2 {
t.Fatalf("arr has %d elements, want 2", got)
}
}
func TestMarshalNestedInlineTableBreaksIndependently(t *testing.T) {
// A nested table breaks on its own measure, so a table whose entries stay
// short keeps the one-line form inside a parent that broke.
const long = "a-very-long-value-that-pushes-the-line-well-past-the-one-hundred-column-limit"
type Cfg struct {
Arr []any `toml:"arr"`
}
out, err := Marshal(Cfg{Arr: []any{int64(1), map[string]any{"n": int64(1), "sub": map[string]any{"name": long}}}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "arr = [1, {\n\tn = 1,\n\tsub = {name = \"" + long + "\"},\n}]\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
type inlineTLS struct {
On bool `toml:"on"`
}
type inlineServer struct {
Host string `toml:"host"`
Port int `toml:"port"`
TLS inlineTLS `toml:"tls"`
}
type inlineBig struct {
A int `toml:"a"`
B int `toml:"b"`
C int `toml:"c"`
}
func TestEncoderInlineTables(t *testing.T) {
type Cfg struct {
Server inlineServer `toml:"server"`
Big inlineBig `toml:"big"`
}
cfg := Cfg{Server: inlineServer{Host: "127.0.0.1", Port: 9090}, Big: inlineBig{A: 1, B: 2, C: 3}}
// The default keeps every sub-table a header section.
headerForm, err := Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "[server]\nhost = \"127.0.0.1\"\nport = 9090\n\n[server.tls]\non = false\n\n[big]\na = 1\nb = 2\nc = 3\n"
if string(headerForm) != want {
t.Errorf("default output mismatch:\ngot: %q\nwant: %q", headerForm, want)
}
// With the option both fit the threshold and become inline tables, nested
// ones included.
out, err := NewEncoder().InlineTables(60).Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
want = "server = {host = \"127.0.0.1\", port = 9090, tls = {on = false}}\nbig = {a = 1, b = 2, c = 3}\n"
if string(out) != want {
t.Errorf("compact output mismatch:\ngot: %q\nwant: %q", out, want)
}
// A threshold below the rendering keeps the header form.
out, err = NewEncoder().InlineTables(10).Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
if string(out) != string(headerForm) {
t.Errorf("small threshold output mismatch:\ngot: %q\nwant: %q", out, headerForm)
}
}
func TestEncoderInlineTablesOrderAndRoundTrip(t *testing.T) {
// An inlined sub-table is a value line, so it precedes every header of the
// document; written after a header it would be read back as part of that
// table. The compact form and the header form parse to the same tree.
type Four struct {
A int `toml:"a"`
B int `toml:"b"`
C int `toml:"c"`
D int `toml:"d"`
}
type Cfg struct {
Small inlineTLS `toml:"small"`
Big Four `toml:"big"`
}
cfg := Cfg{Small: inlineTLS{On: true}, Big: Four{A: 1, B: 2, C: 3, D: 4}}
headerForm, err := Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
compact, err := NewEncoder().InlineTables(20).Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "small = {on = true}\n\n[big]\na = 1\nb = 2\nc = 3\nd = 4\n"
if string(compact) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", compact, want)
}
got, err := ParseMap(compact)
if err != nil {
t.Fatalf("parse of the compact output: %v", err)
}
ref, err := ParseMap(headerForm)
if err != nil {
t.Fatalf("parse of the header output: %v", err)
}
if !reflect.DeepEqual(got, ref) {
t.Errorf("the compact form changed the tree:\ncompact: %#v\nheaders: %#v", got, ref)
}
if _, ok := got["big"].(map[string]any); !ok {
t.Errorf("big = %#v, want a table", got["big"])
}
}
func TestEncoderInlineTablesKeepsArraysOfTables(t *testing.T) {
// An array of tables has no inline form that keeps the value's type, so the
// option leaves it alone and the tree keeps its []map[string]any shape.
type Item struct {
N int `toml:"n"`
}
type Cfg struct {
Items []Item `toml:"items"`
Small inlineTLS `toml:"small"`
}
cfg := Cfg{Items: []Item{{N: 1}}, Small: inlineTLS{On: true}}
out, err := NewEncoder().InlineTables(60).Marshal(cfg)
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "small = {on = true}\n\n[[items]]\nn = 1\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
tree, err := ParseMap(out)
if err != nil {
t.Fatalf("parse: %v", err)
}
if _, ok := tree["items"].([]map[string]any); !ok {
t.Errorf("items = %#v, want []map[string]any", tree["items"])
}
}
// --- Marshaler on array elements -------------------------------------------
// countingMarshaler reports how often its method ran, so a test can check that
// the encoder resolves an element once.
type countingMarshaler struct{ calls *int }
func (c countingMarshaler) MarshalTOML() (any, error) {
*c.calls++
return map[string]any{"n": int64(*c.calls)}, nil
}
// ptrMarshaler carries MarshalTOML on the pointer receiver only.
type ptrMarshaler struct{ V string }
func (p *ptrMarshaler) MarshalTOML() (any, error) { return map[string]any{"v": p.V}, nil }
func TestMarshalerElementOfArrayOfTables(t *testing.T) {
// An element is classified by what MarshalTOML returns, so methods that
// render tables keep the [[header]] form the Go kind would have given them.
type Cfg struct {
Items []marshalerFunc `toml:"items"`
}
out, err := Marshal(Cfg{Items: []marshalerFunc{
func() (any, error) { return map[string]any{"k": "a"}, nil },
func() (any, error) { return map[string]any{"k": "b"}, nil },
}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "[[items]]\nk = \"a\"\n\n[[items]]\nk = \"b\"\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalerElementScalarResultMakesValueArray(t *testing.T) {
// One element rendering itself as a scalar turns the whole array into a
// value array, with the table elements written inline.
type Cfg struct {
Items []marshalerFunc `toml:"items"`
}
out, err := Marshal(Cfg{Items: []marshalerFunc{
func() (any, error) { return map[string]any{"k": "a"}, nil },
func() (any, error) { return "x", nil },
}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "items = [{k = \"a\"}, \"x\"]\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalerElementRunsOnce(t *testing.T) {
// Classification and emission share one result, so the method runs exactly
// once per element even when it decides the array's form.
calls := 0
type Cfg struct {
Items []countingMarshaler `toml:"items"`
}
out, err := Marshal(Cfg{Items: []countingMarshaler{{calls: &calls}, {calls: &calls}, {calls: &calls}}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
if calls != 3 {
t.Errorf("MarshalTOML ran %d times, want 3", calls)
}
want := "[[items]]\nn = 1\n\n[[items]]\nn = 2\n\n[[items]]\nn = 3\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalerPointerReceiverOnElement(t *testing.T) {
// A slice element is addressable, so a pointer-receiver MarshalTOML is
// found there, and the method's table keeps the [[header]] form.
type Cfg struct {
Items []ptrMarshaler `toml:"items"`
}
out, err := Marshal(&Cfg{Items: []ptrMarshaler{{V: "a"}, {V: "b"}}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "[[items]]\nv = \"a\"\n\n[[items]]\nv = \"b\"\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
// The same method is found on an addressable struct field, whose result is
// a table and so keeps a header section.
type Field struct {
F ptrMarshaler `toml:"f"`
}
out, err = Marshal(&Field{F: ptrMarshaler{V: "c"}})
if err != nil {
t.Fatalf("marshal: %v", err)
}
if want := "[f]\nv = \"c\"\n"; string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
func TestMarshalerElementErrorCarriesPath(t *testing.T) {
type Cfg struct {
Items []any `toml:"items"`
}
_, err := Marshal(Cfg{Items: []any{map[string]any{"k": "a"}, failingMarshalerFunc{}}})
if err == nil {
t.Fatal("expected an error from MarshalTOML")
}
if !strings.Contains(err.Error(), "oops") {
t.Errorf("err = %v, want substring \"oops\"", err)
}
ee, ok := errors.AsType[*EncodeError](err)
if !ok {
t.Fatalf("expected an *EncodeError, got %T: %v", err, err)
}
if ee.Path != "items[1]" {
t.Fatalf("Path = %q, want %q", ee.Path, "items[1]")
}
}
func TestMarshalerResultIsNormalised(t *testing.T) {
// A result is normalised like any other value, so a method may return a
// plain int or a duration where the Go kind alone would not encode.
type Cfg struct {
Plain marshalerFunc `toml:"plain"`
Duration marshalerFunc `toml:"duration"`
Elements []any `toml:"elements"`
}
out, err := Marshal(Cfg{
Plain: func() (any, error) { return 7, nil },
Duration: func() (any, error) { return 90 * time.Minute, nil },
Elements: []any{marshalerFunc(func() (any, error) { return 8, nil })},
})
if err != nil {
t.Fatalf("marshal: %v", err)
}
want := "plain = 7\nduration = \"1h30m0s\"\nelements = [8]\n"
if string(out) != want {
t.Errorf("output mismatch:\ngot: %q\nwant: %q", out, want)
}
}
+1 -1
View File
@@ -13,7 +13,7 @@ import (
"os"
"time"
"sourcedock.dev/petrbalvin/interpres"
"sourcedock.dev/petrbalvin/interpres/v2"
)
// document is a small but realistic configuration: it has scalars, a
+2 -2
View File
@@ -41,7 +41,7 @@ func FuzzParse(f *testing.F) {
f.Add([]byte(s))
}
f.Fuzz(func(t *testing.T, data []byte) {
tree, err := Parse(data)
tree, err := ParseMap(data)
if err != nil {
return
}
@@ -49,7 +49,7 @@ func FuzzParse(f *testing.F) {
if err != nil {
t.Fatalf("marshal of a parsed tree failed: %v\ntree: %#v", err, tree)
}
re, err := Parse(out)
re, err := ParseMap(out)
if err != nil {
t.Fatalf("re-parse of the emitted document failed: %v\ndoc:\n%s", err, out)
}
+1 -1
View File
@@ -1,3 +1,3 @@
module sourcedock.dev/petrbalvin/interpres
module sourcedock.dev/petrbalvin/interpres/v2
go 1.27.1
+148 -31
View File
@@ -11,9 +11,10 @@
//
// out, err := interpres.Marshal(cfg)
//
// or, for an untyped tree:
// or, for the document with its key order and comments:
//
// tree, err := interpres.Parse(data)
// doc, err := interpres.Parse(data)
// tree := doc.Map()
//
// A Decoder allows strict decoding that rejects keys without a matching
// struct field, mirroring (*json.Decoder).DisallowUnknownFields.
@@ -89,31 +90,77 @@ func (e *EncodeError) Error() string { return "interpres: " + e.Path + ": " + e.
// Unwrap returns the failure the path points at.
func (e *EncodeError) Unwrap() error { return e.Err }
// Parse decodes a TOML document into a nested map[string]any.
// Parse decodes a TOML document into a Document: the values, the order the
// keys were written in, whether a table was written inline, and the comments.
// ParseMap gives the plain value tree instead.
//
// Values are mapped to Go types as follows: strings to string, integers to
// int64, floats to float64, booleans to bool, date-times to time.Time, arrays
// to []any, and tables (including inline tables) to map[string]any.
// int64, floats to float64, booleans to bool, offset date-times to
// OffsetDateTime, the local date-time kinds to their wrappers, arrays to
// []any, and tables (including inline tables) to map[string]any.
//
// Parse is equivalent to ParseContext with context.Background.
func Parse(data []byte) (map[string]any, error) {
func Parse(data []byte) (*Document, error) {
return ParseContext(context.Background(), data)
}
// ParseContext decodes a TOML document into a nested map[string]any, obeying
// ctx. The context is checked between top-level statements so cancellation is
// honoured before the parser has done substantial work.
func ParseContext(ctx context.Context, data []byte) (map[string]any, error) {
// ParseContext decodes a TOML document into a Document, obeying ctx. The
// context is checked between top-level statements so cancellation is honoured
// before the parser has done substantial work.
func ParseContext(ctx context.Context, data []byte) (*Document, error) {
_, doc, err := parseWithOptions(ctx, data, parseOptions{}, true)
return doc, err
}
// ParseMap decodes a TOML document into a nested map[string]any, the value
// tree without the order and the comments a Document carries. It is the shape
// this package parsed into before [Document] existed.
//
// ParseMap is equivalent to ParseMapContext with context.Background.
func ParseMap(data []byte) (map[string]any, error) {
return ParseMapContext(context.Background(), data)
}
// ParseMapContext is the cancellable variant of ParseMap.
func ParseMapContext(ctx context.Context, data []byte) (map[string]any, error) {
tree, _, err := parseWithOptions(ctx, data, parseOptions{}, false)
return tree, err
}
// parseOptions bound the work one parse may do. A zero field takes the
// default.
type parseOptions struct {
maxDepth int
maxInputSize int
}
// parseWithOptions parses data, building the node tree of a Document when
// wantDoc asks for it, and returns both the value tree and that document.
func parseWithOptions(ctx context.Context, data []byte, opts parseOptions, wantDoc bool) (map[string]any, *Document, error) {
if err := ctx.Err(); err != nil {
return nil, err
return nil, nil, err
}
if opts.maxInputSize > 0 && len(data) > opts.maxInputSize {
return nil, nil, fmt.Errorf("interpres: input is %d bytes, over the limit of %d", len(data), opts.maxInputSize)
}
if !utf8.Valid(data) {
return nil, &SyntaxError{Line: 1, Msg: "input is not valid UTF-8"}
return nil, nil, &SyntaxError{Line: 1, Msg: "input is not valid UTF-8"}
}
maxDepth := opts.maxDepth
if maxDepth <= 0 {
maxDepth = maxNestingDepth
}
// The parser scans data in place; it only reads the buffer, and every
// string it stores in the tree is copied out of it.
p := &parser{src: data, line: 1, ctx: ctx}
return p.parse()
p := &parser{src: data, line: 1, ctx: ctx, maxDepth: maxDepth, wantDoc: wantDoc}
tree, err := p.parse()
if err != nil {
return nil, nil, err
}
if !wantDoc {
return tree, nil, nil
}
return tree, &Document{root: p.doc, footer: p.footer}, nil
}
// Unmarshal parses a TOML document and stores the result in the value pointed
@@ -123,6 +170,11 @@ func ParseContext(ctx context.Context, data []byte) (map[string]any, error) {
// case-insensitive match on the field name when no tag is present. A tag of
// "-" skips the field.
//
// A destination implementing Unmarshaler receives the parsed value as it is,
// a TOML string fills a destination implementing encoding.TextUnmarshaler, and
// a time.Duration destination takes a duration literal such as `1h30m` or a
// bare integer as its nanosecond count.
//
// Unmarshal is equivalent to UnmarshalContext with context.Background.
func Unmarshal(data []byte, v any) error {
return UnmarshalContext(context.Background(), data, v)
@@ -130,7 +182,7 @@ func Unmarshal(data []byte, v any) error {
// UnmarshalContext is the cancellable variant of Unmarshal.
func UnmarshalContext(ctx context.Context, data []byte, v any) error {
tree, err := ParseContext(ctx, data)
tree, err := ParseMapContext(ctx, data)
if err != nil {
return err
}
@@ -138,9 +190,11 @@ func UnmarshalContext(ctx context.Context, data []byte, v any) error {
}
// A Decoder decodes a TOML document into a Go value with configurable
// strictness.
// strictness and configurable limits on the parse it performs.
type Decoder struct {
disallowUnknown bool
maxDepth int
maxInputSize int
}
// NewDecoder returns a Decoder.
@@ -153,6 +207,28 @@ func (d *Decoder) DisallowUnknownFields() *Decoder {
return d
}
// MaxDepth bounds how deeply arrays and inline tables may nest in a document
// this decoder accepts. The parser is a recursive descent, so a document that
// nests without bound would exhaust the stack; one that nests deeper than the
// limit is rejected with a SyntaxError naming it instead. Use 0 or any
// negative value for the default of 10000, which no hand-written document
// approaches.
func (d *Decoder) MaxDepth(depth int) *Decoder {
d.maxDepth = depth
return d
}
// MaxInputSize bounds the size of a document this decoder accepts, in bytes; a
// larger one is rejected before parsing starts. Use 0 or any negative value for
// no limit, which is the default: the caller already holds the bytes, so the
// size is a policy the caller sets rather than a protection the library
// imposes on its own. Parse and ParseContext take no limit beyond the nesting
// default.
func (d *Decoder) MaxInputSize(size int) *Decoder {
d.maxInputSize = size
return d
}
// Decode parses data and stores the result in the value pointed to by v,
// honouring the decoder's strictness settings.
//
@@ -163,7 +239,10 @@ func (d *Decoder) Decode(data []byte, v any) error {
// DecodeContext is the cancellable variant of Decode.
func (d *Decoder) DecodeContext(ctx context.Context, data []byte, v any) error {
tree, err := ParseContext(ctx, data)
tree, _, err := parseWithOptions(ctx, data, parseOptions{
maxDepth: d.maxDepth,
maxInputSize: d.maxInputSize,
}, false)
if err != nil {
return err
}
@@ -177,6 +256,10 @@ func (d *Decoder) DecodeContext(ctx context.Context, data []byte, v any) error {
// then encodes as if the returned value had been passed in its place, which
// is useful for emitting a Go type as a different TOML shape (for example, a
// struct as an inline table or a primitive alias as a richer value).
//
// MarshalTOML wins over encoding.TextMarshaler when a type implements both.
// A type that implements only encoding.TextMarshaler is encoded as a TOML
// string holding its text, and needs no method here.
type Marshaler interface {
MarshalTOML() (any, error)
}
@@ -184,22 +267,28 @@ type Marshaler interface {
// Unmarshaler is the inverse of Marshaler: a type that wants control over
// how it is decoded from a TOML value may implement UnmarshalTOML. The data
// argument is whatever the parser produced for that key: one of string,
// bool, int64, float64, time.Time, LocalDateTime, LocalDate, LocalTime,
// []any, or map[string]any. UnmarshalTOML may parse, inspect, or transform
// the value however it likes, then store the result by mutating its
// receiver through the standard pointer-indirection rules of the reflect
// package (i.e. via reflect.Value.Set or by reassigning fields through a
// pointer the receiver holds).
// bool, int64, float64, OffsetDateTime, LocalDateTime, LocalDate, LocalTime,
// []any, or map[string]any. A tree built by hand may carry a plain time.Time
// where the parser would put an OffsetDateTime.
//
// UnmarshalTOML may parse, inspect, or transform the value however it likes,
// then store the result by mutating its receiver through the standard
// pointer-indirection rules of the reflect package (i.e. via
// reflect.Value.Set or by reassigning fields through a pointer the receiver
// holds).
//
// UnmarshalTOML is invoked from (*Decoder).Decode / Unmarshal when the
// destination type implements the interface. The decoder does not need to
// consult the concrete return value; whatever the receiver stores is kept.
//
// UnmarshalTOML wins over encoding.TextUnmarshaler when a type implements
// both. A type that implements only encoding.TextUnmarshaler is filled from a
// TOML string holding its text, and needs no method here.
type Unmarshaler interface {
UnmarshalTOML(data any) error
}
// Marshal returns the TOML encoding of v. The output stays within TOML 1.0,
// so it is valid under both TOML 1.0 and 1.1.
// Marshal returns the TOML encoding of v. The output is valid TOML 1.1.
//
// Marshal traverses v using reflection and applies the following rules:
//
@@ -219,10 +308,17 @@ type Unmarshaler interface {
// value array (for example an inline table in a mixed array) emits as an
// inline table.
// - Scalars encode as TOML scalars: bool, int64, float64, string, time.Time
// (offset date-time), and LocalDateTime/LocalDate/LocalTime (local
// variants).
// and OffsetDateTime (offset date-time), and LocalDateTime/LocalDate/
// LocalTime (local variants). A date-time writes its seconds only when the value carries
// them, and drops the trailing zeros of a fractional second.
// - A table element of a value array, and a sub-table inlined by
// Encoder.InlineTables, is written as an inline table, across lines when it
// does not fit one.
// - Values implementing Marshaler are encoded by calling MarshalTOML and
// using its result.
// - Values implementing encoding.TextMarshaler, and not one of the
// date-time types, encode as a TOML string holding the text the method
// returns. time.Duration is written in its canonical Go form, `1h30m0s`.
// - nil pointer fields are omitted.
//
// Marshal cannot encode cyclic data structures; passing one will loop until
@@ -246,14 +342,16 @@ func MarshalContext(ctx context.Context, v any) ([]byte, error) {
// An Encoder encodes Go values into TOML.
//
// All options default to behaviour that preserves byte-for-byte compatibility
// with previous releases and passes the toml-test compliance suite:
// All options default to the behaviour that passes the toml-test compliance
// suite in both directions:
//
// GroupByKind: true (scalars first, then tables, then arrays of tables)
// OmitEmptyArrays: false (a nil/empty []string slice emits [] as a value;
// a nil/empty []Item struct slice is still skipped)
// LiteralMultilineAt: 0 (always emit basic multi-line strings with
// escape sequences, never literal ones)
// LiteralMultilineAt: 0 (always emit the escaped basic form, never a
// literal one)
// InlineTablesAt: 0 (always emit a table header, never an inline
// table)
//
// Use the chainable option methods to opt out. The option state is private;
// callers that need the underlying knobs reach for the methods rather than
@@ -262,6 +360,7 @@ type Encoder struct {
groupByKind bool // default true; set via (*Encoder).GroupByKind
omitEmptyArrays bool // default false; set via (*Encoder).OmitEmptyArrays
literalMultilineAt int // default 0; set via (*Encoder).UseLiteralMultiline
inlineTablesAt int // default 0; set via (*Encoder).InlineTables
}
// NewEncoder returns an Encoder with default options.
@@ -294,6 +393,24 @@ func (e *Encoder) UseLiteralMultiline(threshold int) *Encoder {
return e
}
// InlineTables sets the size limit, in bytes of the single-line rendering, at
// which a sub-table is written as an inline table instead of a table header,
// which makes a document of small tables shorter. Use 0 or any negative value
// to disable (always emit a header).
//
// A sub-table is inlined only when doing so keeps every value's type: an array
// of tables keeps its header form, because its inline form would re-parse as a
// value array. An inlined table that does not fit the line is written across
// lines, which TOML 1.1 allows.
//
// With GroupByKind(false) the layout is already for presentation only, and an
// inlined table follows the same rule as any other value line: it lands in the
// section of the header that precedes it.
func (e *Encoder) InlineTables(threshold int) *Encoder {
e.inlineTablesAt = threshold
return e
}
// Marshal encodes v to TOML bytes. It is equivalent to calling Marshal with v.
//
// Marshal is equivalent to MarshalContext with context.Background.
+73 -48
View File
@@ -4,6 +4,7 @@
package interpres
import (
"errors"
"math"
"strings"
"testing"
@@ -11,7 +12,7 @@ import (
)
func TestParseScalars(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
title = "interpres"
count = 42
ratio = 3.14
@@ -49,7 +50,7 @@ expv = 1e3
}
func TestParseInfNan(t *testing.T) {
tree, err := Parse([]byte("pos = inf\nneg = -inf\nbad = nan\n"))
tree, err := ParseMap([]byte("pos = inf\nneg = -inf\nbad = nan\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -65,7 +66,7 @@ func TestParseInfNan(t *testing.T) {
}
func TestParseStrings(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
basic = "a\tb\nc"
literal = 'C:\path\no\escape'
quote = "say \"hi\""
@@ -89,7 +90,7 @@ unicode = "\u00e9"
}
func TestParseMultilineString(t *testing.T) {
tree, err := Parse([]byte("text = \"\"\"\nfirst\nsecond\"\"\"\n"))
tree, err := ParseMap([]byte("text = \"\"\"\nfirst\nsecond\"\"\"\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -99,7 +100,7 @@ func TestParseMultilineString(t *testing.T) {
}
func TestParseMultilineLineEndingBackslash(t *testing.T) {
tree, err := Parse([]byte("text = \"\"\"\\\n one \\\n two\"\"\"\n"))
tree, err := ParseMap([]byte("text = \"\"\"\\\n one \\\n two\"\"\"\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -109,7 +110,7 @@ func TestParseMultilineLineEndingBackslash(t *testing.T) {
}
func TestParseTablesAndDottedKeys(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
owner.name = "Petr"
[server]
@@ -137,7 +138,7 @@ enabled = true
}
func TestParseArrayOfTables(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
[[forms]]
name = "contact"
@@ -157,7 +158,7 @@ name = "feedback"
}
func TestParseArraysAndInlineTables(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
ports = [80, 443]
mixed = [
"a",
@@ -183,7 +184,7 @@ point = { x = 1, y = 2 }
}
func TestParseDateTime(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
offset = 1979-05-27T07:32:00Z
local = 1979-05-27T07:32:00
day = 1979-05-27
@@ -192,7 +193,7 @@ clock = 07:32:00
if err != nil {
t.Fatalf("parse: %v", err)
}
if off, ok := tree["offset"].(time.Time); !ok || off.Year() != 1979 || off.Hour() != 7 {
if off, ok := tree["offset"].(OffsetDateTime); !ok || off.Year() != 1979 || off.Hour() != 7 {
t.Errorf("offset = %#v (%T)", tree["offset"], tree["offset"])
}
if ldt, ok := tree["local"].(LocalDateTime); !ok || ldt.Year() != 1979 || ldt.Hour() != 7 {
@@ -207,15 +208,15 @@ clock = 07:32:00
}
func TestDateTimeFormats(t *testing.T) {
tree, err := Parse([]byte("a = 1987-07-05 17:45:00Z\nb = 1987-07-05t17:45:00z\nc = 1977-12-21T10:32:00.555\n"))
tree, err := ParseMap([]byte("a = 1987-07-05 17:45:00Z\nb = 1987-07-05t17:45:00z\nc = 1977-12-21T10:32:00.555\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if _, ok := tree["a"].(time.Time); !ok {
t.Errorf("a is %T, want time.Time", tree["a"])
if _, ok := tree["a"].(OffsetDateTime); !ok {
t.Errorf("a is %T, want OffsetDateTime", tree["a"])
}
if _, ok := tree["b"].(time.Time); !ok {
t.Errorf("b is %T, want time.Time", tree["b"])
if _, ok := tree["b"].(OffsetDateTime); !ok {
t.Errorf("b is %T, want OffsetDateTime", tree["b"])
}
if _, ok := tree["c"].(LocalDateTime); !ok {
t.Errorf("c is %T, want LocalDateTime", tree["c"])
@@ -352,7 +353,7 @@ func TestSkippedFieldTag(t *testing.T) {
}
func TestSyntaxErrorReportsLine(t *testing.T) {
_, err := Parse([]byte("a = 1\nb = \nc = 3\n"))
_, err := ParseMap([]byte("a = 1\nb = \nc = 3\n"))
if err == nil {
t.Fatal("expected a syntax error")
}
@@ -366,7 +367,7 @@ func TestSyntaxErrorReportsLine(t *testing.T) {
}
func TestComments(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
# a leading comment
key = "value" # trailing comment
# another
@@ -380,7 +381,7 @@ key = "value" # trailing comment
}
func TestDuplicateKeyRejected(t *testing.T) {
_, err := Parse([]byte("a = 1\na = 2\n"))
_, err := ParseMap([]byte("a = 1\na = 2\n"))
if err == nil {
t.Fatal("expected duplicate key error")
}
@@ -395,7 +396,7 @@ func TestRejectsInvalidNumbers(t *testing.T) {
"0x", "0o", "0b", "0b2", "0o8", "0xG",
"+0x1",
} {
if _, err := Parse([]byte("v = " + tok + "\n")); err == nil {
if _, err := ParseMap([]byte("v = " + tok + "\n")); err == nil {
t.Errorf("%q: expected an error, got none", tok)
}
}
@@ -408,22 +409,22 @@ func TestParseRejectsOffsetOutOfRange(t *testing.T) {
"1979-05-27T07:32:00+24:00",
"1979-05-27T07:32:00+99:99",
} {
if _, err := Parse([]byte("v = " + tok + "\n")); err == nil {
if _, err := ParseMap([]byte("v = " + tok + "\n")); err == nil {
t.Errorf("%q: expected an error, got none", tok)
}
}
}
func TestParseAcceptsOffsetBounds(t *testing.T) {
tree, err := Parse([]byte("a = 1979-05-27T07:32:00+23:59\nb = 1979-05-27T07:32:00-23:59\n"))
tree, err := ParseMap([]byte("a = 1979-05-27T07:32:00+23:59\nb = 1979-05-27T07:32:00-23:59\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
a := tree["a"].(time.Time)
a := tree["a"].(OffsetDateTime)
if _, offset := a.Zone(); offset != 23*3600+59*60 {
t.Fatalf("a offset = %d, want %d", offset, 23*3600+59*60)
}
b := tree["b"].(time.Time)
b := tree["b"].(OffsetDateTime)
if _, offset := b.Zone(); offset != -(23*3600 + 59*60) {
t.Fatalf("b offset = %d", offset)
}
@@ -449,7 +450,7 @@ func TestAcceptsNumberEdgeCases(t *testing.T) {
"-2.5E-3": -2.5e-3,
}
for tok, want := range cases {
tree, err := Parse([]byte("v = " + tok + "\n"))
tree, err := ParseMap([]byte("v = " + tok + "\n"))
if err != nil {
t.Errorf("%q: %v", tok, err)
continue
@@ -461,14 +462,14 @@ func TestAcceptsNumberEdgeCases(t *testing.T) {
}
func TestRejectsTableRedefinition(t *testing.T) {
_, err := Parse([]byte("[a]\nx = 1\n\n[a]\ny = 2\n"))
_, err := ParseMap([]byte("[a]\nx = 1\n\n[a]\ny = 2\n"))
if err == nil {
t.Fatal("expected a table-redefinition error")
}
}
func TestAllowsImplicitThenExplicitTable(t *testing.T) {
tree, err := Parse([]byte("[a.b]\nx = 1\n\n[a]\ny = 2\n"))
tree, err := ParseMap([]byte("[a.b]\nx = 1\n\n[a]\ny = 2\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -482,13 +483,13 @@ func TestAllowsImplicitThenExplicitTable(t *testing.T) {
}
func TestRejectsControlCharInString(t *testing.T) {
if _, err := Parse([]byte("v = \"a\x01b\"\n")); err == nil {
if _, err := ParseMap([]byte("v = \"a\x01b\"\n")); err == nil {
t.Fatal("expected a control-character error")
}
}
func TestAllowsEscapedControlChar(t *testing.T) {
tree, err := Parse([]byte(`v = "\u0000"`))
tree, err := ParseMap([]byte(`v = "\u0000"`))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -498,7 +499,7 @@ func TestAllowsEscapedControlChar(t *testing.T) {
}
func TestMultilineQuotesAtDelimiter(t *testing.T) {
tree, err := Parse([]byte("a = '''''two quotes'''''\n"))
tree, err := ParseMap([]byte("a = '''''two quotes'''''\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -517,7 +518,7 @@ func TestRejectsInlineTableExtension(t *testing.T) {
"by nested array header": "a = { b = {} }\n[[a.b.c]]\nx = 2\n",
}
for name, doc := range cases {
if _, err := Parse([]byte(doc)); err == nil {
if _, err := ParseMap([]byte(doc)); err == nil {
t.Errorf("%s: expected an inline-table extension error", name)
}
}
@@ -532,7 +533,7 @@ func TestArrayOfTablesFreshScopePerElement(t *testing.T) {
"dotted key": "[[a]]\nb.c = 1\n[[a]]\n[a.b]\nd = 2\n",
}
for name, doc := range cases {
tree, err := Parse([]byte(doc))
tree, err := ParseMap([]byte(doc))
if err != nil {
t.Errorf("%s: %v", name, err)
continue
@@ -547,7 +548,7 @@ func TestArrayOfTablesFreshScopePerElement(t *testing.T) {
"header over dotted in one element": "[[a]]\nb.c = 1\n[a.b]\nd = 2\n",
"table over nested array": "[[a]]\n[[a.b]]\n[a.b]\nx = 1\n",
} {
if _, err := Parse([]byte(doc)); err == nil {
if _, err := ParseMap([]byte(doc)); err == nil {
t.Errorf("%s: expected an error, got none", name)
}
}
@@ -565,14 +566,14 @@ func TestRejectsSpecInvalid(t *testing.T) {
// makes the seconds optional.
}
for name, doc := range cases {
if _, err := Parse([]byte(doc)); err == nil {
if _, err := ParseMap([]byte(doc)); err == nil {
t.Errorf("%s: expected an error", name)
}
}
}
func TestArrayOfTablesPerElementSubtable(t *testing.T) {
tree, err := Parse([]byte(`
tree, err := ParseMap([]byte(`
[[forms]]
name = "a"
@@ -603,7 +604,7 @@ host = "h2"
// --- TOML 1.1 --------------------------------------------------------------
func TestParseAcceptsNoSecondsDatetimes(t *testing.T) {
tree, err := Parse([]byte(`t = 13:37
tree, err := ParseMap([]byte(`t = 13:37
dt = 1979-05-27T07:32
odt1 = 1979-05-27 07:32Z
odt2 = 1979-05-27 07:32-07:00
@@ -611,26 +612,28 @@ odt2 = 1979-05-27 07:32-07:00
if err != nil {
t.Fatalf("parse: %v", err)
}
if got := tree["t"].(LocalTime).String(); got != "13:37:00" {
t.Errorf("t = %q, want %q", got, "13:37:00")
// A value written without seconds comes back without them: the seconds are
// only written when the value carries them.
if got := tree["t"].(LocalTime).String(); got != "13:37" {
t.Errorf("t = %q, want %q", got, "13:37")
}
if got := tree["dt"].(LocalDateTime).String(); got != "1979-05-27T07:32:00" {
t.Errorf("dt = %q, want %q", got, "1979-05-27T07:32:00")
if got := tree["dt"].(LocalDateTime).String(); got != "1979-05-27T07:32" {
t.Errorf("dt = %q, want %q", got, "1979-05-27T07:32")
}
if got := tree["odt1"].(time.Time).Format(time.RFC3339Nano); got != "1979-05-27T07:32:00Z" {
if got := tree["odt1"].(OffsetDateTime).Format(time.RFC3339Nano); got != "1979-05-27T07:32:00Z" {
t.Errorf("odt1 = %q", got)
}
if got := tree["odt2"].(time.Time).Format(time.RFC3339Nano); got != "1979-05-27T07:32:00-07:00" {
if got := tree["odt2"].(OffsetDateTime).Format(time.RFC3339Nano); got != "1979-05-27T07:32:00-07:00" {
t.Errorf("odt2 = %q", got)
}
// The fraction still requires the seconds it belongs to.
if _, err := Parse([]byte("a = 07:32.5\n")); err == nil {
if _, err := ParseMap([]byte("a = 07:32.5\n")); err == nil {
t.Error("07:32.5: expected an error, got none")
}
}
func TestParseAcceptsEscapeAndHexEscapes(t *testing.T) {
tree, err := Parse([]byte(`esc = "\e"
tree, err := ParseMap([]byte(`esc = "\e"
hex = "\x20\x7f\xf8"
nul = "\x00"
multi = """\x68\x65"""
@@ -657,14 +660,14 @@ lit = '\x20'
}
// Two digits exactly; a short or non-hex escape is an error.
for _, doc := range []string{`a = "\x4"`, `a = "\x"`, `a = "\xgg"`} {
if _, err := Parse([]byte(doc)); err == nil {
if _, err := ParseMap([]byte(doc)); err == nil {
t.Errorf("%s: expected an error, got none", doc)
}
}
}
func TestParseAcceptsMultilineInlineTables(t *testing.T) {
tree, err := Parse([]byte("tbl = {\n\thello = \"world\",\n\tarr = [1,\n\t\t2,\n\t],\n\tsub = {\n\t\tk = 1,\n\t},\n\tbare = 2}\n"))
tree, err := ParseMap([]byte("tbl = {\n\thello = \"world\",\n\tarr = [1,\n\t\t2,\n\t],\n\tsub = {\n\t\tk = 1,\n\t},\n\tbare = 2}\n"))
if err != nil {
t.Fatalf("parse: %v", err)
}
@@ -679,7 +682,7 @@ func TestParseAcceptsMultilineInlineTables(t *testing.T) {
t.Errorf("sub = %#v", tbl["sub"])
}
// Comments inside the table, and a trailing comma at both depths.
tree, err = Parse([]byte("m = { # one\n\t# two\n\ta = 1, # three\n\t# four\n}\n"))
tree, err = ParseMap([]byte("m = { # one\n\t# two\n\ta = 1, # three\n\t# four\n}\n"))
if err != nil {
t.Fatalf("parse with comments: %v", err)
}
@@ -687,7 +690,7 @@ func TestParseAcceptsMultilineInlineTables(t *testing.T) {
t.Errorf("m = %#v", m)
}
// The old single-line shapes keep working, with and without the comma.
if _, err := Parse([]byte("a = { b = 1, c = 2 }\n")); err != nil {
if _, err := ParseMap([]byte("a = { b = 1, c = 2 }\n")); err != nil {
t.Errorf("single line: %v", err)
}
// Still rejected: two commas, a missing value, and an unclosed table.
@@ -696,8 +699,30 @@ func TestParseAcceptsMultilineInlineTables(t *testing.T) {
"missing value": "a = {\n\tb =\n}\n",
"unterminated": "a = { b = 1,\n",
} {
if _, err := Parse([]byte(doc)); err == nil {
if _, err := ParseMap([]byte(doc)); err == nil {
t.Errorf("%s: expected an error, got none", name)
}
}
}
func TestParseNestingLimit(t *testing.T) {
// The parser is a recursive descent, so a document that nests without bound
// is rejected instead of exhausting the stack.
deep := func(n int) []byte {
return []byte("v = " + strings.Repeat("[", n) + strings.Repeat("]", n) + "\n")
}
if _, err := ParseMap(deep(100)); err != nil {
t.Fatalf("a document well inside the limit: %v", err)
}
_, err := ParseMap(deep(maxNestingDepth + 1))
if err == nil {
t.Fatal("expected a nesting error")
}
var se *SyntaxError
if !errors.As(err, &se) {
t.Fatalf("expected a *SyntaxError, got %T: %v", err, err)
}
if !strings.Contains(se.Msg, "nesting") {
t.Errorf("Msg = %q, want it to name the nesting limit", se.Msg)
}
}
+2 -2
View File
@@ -92,9 +92,9 @@ run:
dev:
go run -buildvcs=true {{package}}
# Runs the official toml-test compliance suite against the built adapter; toml-test must be on PATH (go install github.com/toml-lang/toml-test/v2/cmd/toml-test@v2.2.0); not standard because no canonical recipe covers a domain compliance suite.
# Runs the official toml-test compliance suite in both directions, decoder and encoder, against the built adapter; toml-test must be on PATH (go install github.com/toml-lang/toml-test/v2/cmd/toml-test@v2.2.0); not standard because no canonical recipe covers a domain compliance suite.
toml-test: build
toml-test test -decoder=bin/interpres-decode -toml=1.1
toml-test test -decoder=bin/interpres-decode -encoder='bin/interpres-decode -encode' -toml=1.1
# Coverage report as an HTML map from the gate's profile; not standard because the gate needs only the numeric floor, and a browser artefact is exploration, not a gate.
coverage-html: test
+238 -23
View File
@@ -30,6 +30,12 @@ type parser struct {
line int
ctx context.Context
// maxDepth and depth bound the nesting the recursive descent may follow:
// arrays and inline tables nest through parseValue, and without a limit a
// hostile document would exhaust the stack.
maxDepth int
depth int
root map[string]any
current map[string]any
headers map[string]bool
@@ -38,8 +44,46 @@ type parser struct {
arrays map[string]bool
currentPath []string
// wantDoc asks for the node tree the Document is built from; doc is that
// tree, and it stays nil when only the value tree is wanted. currentNode
// is the node of p.current; pending collects the comment lines since the
// last statement and trailing the comment on the statement's own line;
// lastEntry and lastTable name the statement those comments belong to;
// lastInline and lastArrayElems carry the nodes of the value parseValue has
// just produced.
wantDoc bool
doc *Table
currentNode *Table
pending []string
trailing string
lastEntry *Entry
lastTable *Table
lastInline *Table
lastArrayElems []*Table
// footer holds the comment lines that follow the last statement, which
// belong to the document rather than to any table or key.
footer []string
}
// maxNestingDepth bounds how deeply arrays and inline tables may nest when no
// limit is set. It matches the default encoding/json uses for the same reason,
// and sits far above any document a person writes.
const maxNestingDepth = 10000
// enterNesting counts one level of array or inline-table nesting and reports a
// document that nests deeper than the limit allows.
func (p *parser) enterNesting() error {
p.depth++
if p.depth > p.maxDepth {
return p.errf("nesting exceeds the limit of %d", p.maxDepth)
}
return nil
}
func (p *parser) leaveNesting() { p.depth-- }
func (p *parser) parse() (map[string]any, error) {
p.root = map[string]any{}
p.current = p.root
@@ -48,6 +92,10 @@ func (p *parser) parse() (map[string]any, error) {
p.dotted = map[string]bool{}
p.arrays = map[string]bool{}
p.currentPath = nil
if p.wantDoc {
p.doc = newTable(p.root)
p.currentNode = p.doc
}
for i := 0; ; i++ {
if i%ctxCheckInterval == 0 {
@@ -61,6 +109,7 @@ func (p *parser) parse() (map[string]any, error) {
if p.eof() {
break
}
p.lastEntry, p.lastTable = nil, nil
c := p.peek()
switch {
case c == '[':
@@ -75,10 +124,38 @@ func (p *parser) parse() (map[string]any, error) {
if err := p.expectLineEnd(); err != nil {
return nil, err
}
p.attachComments()
}
p.attachFooter()
return p.root, nil
}
// attachComments hands the collected comments to the statement just parsed:
// the lines above it to its entry or table, the comment on its own line as the
// trailing one.
func (p *parser) attachComments() {
if p.doc != nil {
switch {
case p.lastEntry != nil:
p.lastEntry.comments = p.pending
p.lastEntry.trailing = p.trailing
case p.lastTable != nil:
p.lastTable.comments = p.pending
p.lastTable.trailing = p.trailing
}
}
p.pending, p.trailing = nil, ""
}
// attachFooter hands the comment lines that follow the last statement to the
// document, which is where a comment block at the end of a file belongs.
func (p *parser) attachFooter() {
if p.doc != nil && len(p.pending) > 0 {
p.footer = p.pending
}
p.pending = nil
}
// checkCtx returns ctx.Err() when the context has been cancelled, nil
// otherwise. The call is a no-op when ctx is nil or the zero Background
// context, both of which never cancel.
@@ -117,7 +194,7 @@ func (p *parser) parseTableHeader() error {
}
if array {
tbl, err := p.appendArrayTable(key)
tbl, elem, err := p.appendArrayTable(key)
if err != nil {
return err
}
@@ -127,6 +204,8 @@ func (p *parser) parseTableHeader() error {
p.arrays[pathKey(key)] = true
p.current = tbl
p.currentPath = key
p.currentNode = elem
p.lastTable = elem
return nil
}
@@ -136,69 +215,91 @@ func (p *parser) parseTableHeader() error {
}
p.headers[pk] = true
tbl, err := p.tableAt(key)
tbl, node, err := p.tableAt(key)
if err != nil {
return err
}
p.current = tbl
p.currentPath = key
p.currentNode = node
p.lastTable = node
return nil
}
// tableAt walks (creating intermediate tables) to the table named by key,
// relative to the document root, rejecting any step into a frozen inline table.
func (p *parser) tableAt(key []string) (map[string]any, error) {
func (p *parser) tableAt(key []string) (map[string]any, *Table, error) {
cur := p.root
node := p.doc
path := make([]string, 0, len(key))
for _, k := range key {
path = append(path, k)
if p.frozen[pathKey(path)] {
return nil, p.errf("cannot extend inline table %q", strings.Join(path, "."))
return nil, nil, p.errf("cannot extend inline table %q", strings.Join(path, "."))
}
existing, ok := cur[k]
if !ok {
next := map[string]any{}
cur[k] = next
cur = next
if node != nil {
node = node.addTable(k, next)
}
continue
}
switch v := existing.(type) {
case map[string]any:
cur = v
if node != nil {
node = node.addTable(k, v)
}
case []map[string]any:
if len(v) == 0 {
return nil, p.errf("key %q is an empty array of tables", k)
return nil, nil, p.errf("key %q is an empty array of tables", k)
}
cur = v[len(v)-1]
if node != nil {
node = node.lastElement(k)
}
default:
return nil, p.errf("key %q is not a table", k)
return nil, nil, p.errf("key %q is not a table", k)
}
}
return cur, nil
return cur, node, nil
}
func (p *parser) appendArrayTable(key []string) (map[string]any, error) {
func (p *parser) appendArrayTable(key []string) (map[string]any, *Table, error) {
parent := p.root
node := p.doc
path := make([]string, 0, len(key))
for _, k := range key[:len(key)-1] {
path = append(path, k)
if p.frozen[pathKey(path)] {
return nil, p.errf("cannot extend inline table %q", strings.Join(path, "."))
return nil, nil, p.errf("cannot extend inline table %q", strings.Join(path, "."))
}
existing, ok := parent[k]
if !ok {
next := map[string]any{}
parent[k] = next
parent = next
if node != nil {
node = node.addTable(k, next)
}
continue
}
switch v := existing.(type) {
case map[string]any:
parent = v
if node != nil {
node = node.addTable(k, v)
}
case []map[string]any:
parent = v[len(v)-1]
if node != nil {
node = node.lastElement(k)
}
default:
return nil, p.errf("key %q is not a table", k)
return nil, nil, p.errf("key %q is not a table", k)
}
}
@@ -210,9 +311,13 @@ func (p *parser) appendArrayTable(key []string) (map[string]any, error) {
case []map[string]any:
parent[leaf] = append(existing, tbl)
default:
return nil, p.errf("key %q is not an array of tables", leaf)
return nil, nil, p.errf("key %q is not an array of tables", leaf)
}
return tbl, nil
var elem *Table
if node != nil {
elem = node.addElement(leaf, tbl)
}
return tbl, elem, nil
}
// --- key/value -------------------------------------------------------------
@@ -239,6 +344,9 @@ func (p *parser) parseKeyValue() error {
// top-level statement reuses it for the leaf.
abs := make([]string, 0, len(p.currentPath)+len(key))
abs = append(abs, p.currentPath...)
// dests collects the map each dotted key descended into, which the node
// tree needs to build the matching tables around the value.
var dests []map[string]any
for _, k := range key[:len(key)-1] {
abs = append(abs, k)
if p.frozen[pathKey(abs)] {
@@ -253,6 +361,7 @@ func (p *parser) parseKeyValue() error {
next := map[string]any{}
dest[k] = next
dest = next
dests = append(dests, next)
continue
}
m, ok := existing.(map[string]any)
@@ -260,6 +369,7 @@ func (p *parser) parseKeyValue() error {
return p.errf("key %q is not a table", k)
}
dest = m
dests = append(dests, m)
}
leaf := key[len(key)-1]
abs = append(abs, leaf)
@@ -267,10 +377,47 @@ func (p *parser) parseKeyValue() error {
return p.errf("duplicate key %q", leaf)
}
dest[leaf] = val
if p.doc != nil {
node := p.currentNode
for i, k := range key[:len(key)-1] {
node = node.addTable(k, dests[i])
}
_, inline := val.(map[string]any)
entry := node.addValue(leaf, val, inline)
if inline {
entry.child = p.takeInline(val)
}
if nodes := p.takeArrayElems(val); nodes != nil {
entry.elements = nodes
}
p.lastEntry = entry
}
p.freezeInline(abs, val)
return nil
}
// takeInline returns the node of the inline table just parsed, when v is that
// table's value, and clears it so a later value cannot pick it up.
func (p *parser) takeInline(v any) *Table {
node := p.lastInline
p.lastInline = nil
if _, ok := v.(map[string]any); !ok {
return nil
}
return node
}
// takeArrayElems returns the element nodes of the array just parsed, when v is
// that array's value, and clears them.
func (p *parser) takeArrayElems(v any) []*Table {
nodes := p.lastArrayElems
p.lastArrayElems = nil
if _, ok := v.([]any); !ok {
return nil
}
return nodes
}
// freezeInline marks the path of an inline table (and any nested inline tables)
// as immutable, so a later header or dotted key cannot extend it.
func (p *parser) freezeInline(path []string, val any) {
@@ -360,6 +507,9 @@ func (p *parser) parseValue() (any, error) {
if p.eof() {
return nil, p.errf("expected a value")
}
// A container value leaves its node behind for the caller to pick up; a
// value that follows must not find the previous one.
p.lastInline, p.lastArrayElems = nil, nil
switch c := p.peek(); {
case c == '"':
return p.parseBasicString()
@@ -675,9 +825,23 @@ func (p *parser) readUnicode(n int) (rune, error) {
// --- arrays and inline tables ---------------------------------------------
func (p *parser) parseArray() (any, error) {
func (p *parser) parseArray() (val any, err error) {
if err := p.enterNesting(); err != nil {
return nil, err
}
defer p.leaveNesting()
p.pos++ // '['
arr := []any{}
// elems carries the node of each element that is an inline table, so the
// caller can keep its key order; the entries are nil for other values.
var elems []*Table
if p.doc != nil {
defer func() {
if err == nil {
p.lastArrayElems = elems
}
}()
}
for {
if err := p.skipNestedSpace(); err != nil {
return nil, err
@@ -693,6 +857,9 @@ func (p *parser) parseArray() (any, error) {
if err != nil {
return nil, err
}
if p.doc != nil {
elems = append(elems, p.takeInline(v))
}
arr = append(arr, v)
if err := p.skipNestedSpace(); err != nil {
return nil, err
@@ -712,10 +879,26 @@ func (p *parser) parseArray() (any, error) {
}
}
func (p *parser) parseInlineTable() (any, error) {
func (p *parser) parseInlineTable() (val any, err error) {
if err := p.enterNesting(); err != nil {
return nil, err
}
defer p.leaveNesting()
p.pos++ // '{'
tbl := map[string]any{}
assigned := map[string]bool{}
// The inline table is a node of its own, so the keys keep their order; the
// caller picks the node up when the table parses.
var node *Table
if p.doc != nil {
node = newTable(tbl)
node.inline = true
defer func() {
if err == nil {
p.lastInline = node
}
}()
}
// TOML 1.1 lets an inline table span lines: interior whitespace includes
// newlines and comments, and a trailing comma is allowed before the
// closing brace.
@@ -747,6 +930,7 @@ func (p *parser) parseInlineTable() (any, error) {
dest := tbl
path := make([]string, 0, len(key))
var dests []map[string]any
for _, k := range key[:len(key)-1] {
path = append(path, k)
if assigned[pathKey(path)] {
@@ -757,6 +941,7 @@ func (p *parser) parseInlineTable() (any, error) {
m := map[string]any{}
dest[k] = m
dest = m
dests = append(dests, m)
continue
}
m, isMap := existing.(map[string]any)
@@ -764,6 +949,7 @@ func (p *parser) parseInlineTable() (any, error) {
return nil, p.errf("key %q is already defined", k)
}
dest = m
dests = append(dests, m)
}
leaf := key[len(key)-1]
path = append(path, leaf)
@@ -772,6 +958,20 @@ func (p *parser) parseInlineTable() (any, error) {
}
dest[leaf] = val
assigned[pathKey(path)] = true
if node != nil {
child := node
for i, k := range key[:len(key)-1] {
child = child.addTable(k, dests[i])
}
_, inline := val.(map[string]any)
entry := child.addValue(leaf, val, inline)
if inline {
entry.child = p.takeInline(val)
}
if nodes := p.takeArrayElems(val); nodes != nil {
entry.elements = nodes
}
}
if err := p.skipNestedSpace(); err != nil {
return nil, err
@@ -866,7 +1066,9 @@ func (p *parser) skipNestedSpace() error {
p.line++
p.pos++
case '#':
if err := p.skipComment(); err != nil {
// A comment between values inside an array or an inline table is
// skipped; the Document does not carry those yet.
if _, err := p.skipComment(); err != nil {
return err
}
default:
@@ -890,9 +1092,11 @@ func (p *parser) skipBlank() error {
p.line++
p.pos++
case '#':
if err := p.skipComment(); err != nil {
line, err := p.skipComment()
if err != nil {
return err
}
p.pending = append(p.pending, line)
default:
return nil
}
@@ -900,27 +1104,34 @@ func (p *parser) skipBlank() error {
return nil
}
func (p *parser) skipComment() error {
func (p *parser) skipComment() (string, error) {
p.pos++ // consume '#'
start := p.pos
for !p.eof() {
c := p.peek()
switch {
case c == '\n':
return nil
return commentText(string(p.src[start:p.pos])), nil
case c == '\r':
if p.pos+1 < len(p.src) && p.src[p.pos+1] == '\n' {
return nil
return commentText(string(p.src[start:p.pos])), nil
}
return p.errf("bare carriage return is not allowed")
return "", p.errf("bare carriage return is not allowed")
case c == '\t':
p.pos++
case c < 0x20 || c == 0x7f:
return p.errf("control character U+%04X is not allowed in a comment", c)
return "", p.errf("control character U+%04X is not allowed in a comment", c)
default:
p.pos++
}
}
return nil
return commentText(string(p.src[start:p.pos])), nil
}
// commentText drops the one space that usually follows the '#', so a line
// stored in a Document reads as the comment itself.
func commentText(s string) string {
return strings.TrimPrefix(s, " ")
}
// expectCRLF consumes a carriage return that must be immediately followed by a
@@ -941,9 +1152,13 @@ func (p *parser) expectLineEnd() error {
return nil
}
if p.peek() == '#' {
if err := p.skipComment(); err != nil {
line, err := p.skipComment()
if err != nil {
return err
}
if p.doc != nil {
p.trailing = line
}
}
if p.eof() {
return nil