perf(parse): cut allocations and validate UTF-8 in the scan

Per-statement work no longer allocates where it does not need to: parseKeyPath returns the first key component directly, absolute paths and key buffers sit on the parser, the definition maps are created on first write, arrays presize, and a string without escapes is copied out in one piece instead of built byte by byte. Repeated keys share one string across array-of-tables elements through a parser-local intern table whose lookup works on the bytes, so a repeated key costs no allocation.

UTF-8 validity is no longer a whole-input pass before the parser runs: the scan validates where it meets a multi-byte sequence, and a SyntaxError now names the line the invalid byte sits on. TestParseRejectsInvalidUTF8 covers nine positions; docs/API.md describes the new reporting.

Long document: 67 664 to 31 765 allocations per parse, about 51 to about 106 MB/s; representative document: 157 to 104 allocations.
This commit is contained in:
2026-09-20 22:15:10 +02:00
parent f37c05f9a2
commit d00cbb983a
6 changed files with 427 additions and 117 deletions
+2 -1
View File
@@ -19,7 +19,8 @@ Decodes a TOML document into a [Document](#documents): the values, the order the
keys were written in, whether a table was written inline, and the comments.
The values follow the mapping in the [Decoding](#decoding) section below.
Returns `*SyntaxError` on a malformed document. Input that is not valid UTF-8
is rejected before the parser runs. Equivalent to
is rejected with a `SyntaxError` naming the line where the invalid byte
appears, because validity is checked during the scan. Equivalent to
`ParseContext(context.Background(), data)`.
```go