perf(parse): cut allocations and validate UTF-8 in the scan

Per-statement work no longer allocates where it does not need to: parseKeyPath returns the first key component directly, absolute paths and key buffers sit on the parser, the definition maps are created on first write, arrays presize, and a string without escapes is copied out in one piece instead of built byte by byte. Repeated keys share one string across array-of-tables elements through a parser-local intern table whose lookup works on the bytes, so a repeated key costs no allocation.

UTF-8 validity is no longer a whole-input pass before the parser runs: the scan validates where it meets a multi-byte sequence, and a SyntaxError now names the line the invalid byte sits on. TestParseRejectsInvalidUTF8 covers nine positions; docs/API.md describes the new reporting.

Long document: 67 664 to 31 765 allocations per parse, about 51 to about 106 MB/s; representative document: 157 to 104 allocations.
This commit is contained in:
2026-09-20 22:15:10 +02:00
parent f37c05f9a2
commit d00cbb983a
6 changed files with 427 additions and 117 deletions
+3 -4
View File
@@ -24,7 +24,6 @@ import (
"context"
"errors"
"fmt"
"unicode/utf8"
)
// A SyntaxError describes a malformed TOML document, including the 1-based
@@ -143,9 +142,9 @@ func parseWithOptions(ctx context.Context, data []byte, opts parseOptions, wantD
if opts.maxInputSize > 0 && len(data) > opts.maxInputSize {
return nil, nil, fmt.Errorf("interpres: input is %d bytes, over the limit of %d", len(data), opts.maxInputSize)
}
if !utf8.Valid(data) {
return nil, nil, &SyntaxError{Line: 1, Msg: "input is not valid UTF-8"}
}
// UTF-8 validity is not checked in a pass of its own: the scanner
// validates the multi-byte sequences where it meets them, so an invalid
// byte is reported on its own line instead of always on line 1.
maxDepth := opts.maxDepth
if maxDepth <= 0 {
maxDepth = maxNestingDepth