perf(parse): cut allocations and validate UTF-8 in the scan

Per-statement work no longer allocates where it does not need to: parseKeyPath returns the first key component directly, absolute paths and key buffers sit on the parser, the definition maps are created on first write, arrays presize, and a string without escapes is copied out in one piece instead of built byte by byte. Repeated keys share one string across array-of-tables elements through a parser-local intern table whose lookup works on the bytes, so a repeated key costs no allocation.

UTF-8 validity is no longer a whole-input pass before the parser runs: the scan validates where it meets a multi-byte sequence, and a SyntaxError now names the line the invalid byte sits on. TestParseRejectsInvalidUTF8 covers nine positions; docs/API.md describes the new reporting.

Long document: 67 664 to 31 765 allocations per parse, about 51 to about 106 MB/s; representative document: 157 to 104 allocations.
This commit is contained in:
2026-09-20 22:15:10 +02:00
parent f37c05f9a2
commit d00cbb983a
6 changed files with 427 additions and 117 deletions
+14
View File
@@ -74,6 +74,20 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- The module path carries the /v2 suffix the Go toolchain requires of
every major version 2 module: imports change to
`sourcedock.dev/petrbalvin/interpres/v2`.
- Input that is not valid UTF-8 is now rejected where the parser's scan
meets the invalid byte, with a `SyntaxError` naming that line, instead of
a whole-input check that always reported line 1. Invalid input is still
rejected; the reported location is now the byte's own.
**Performance**
- Parsing is faster than in 1.1.0 while carrying the new document layer:
the suite's representative document decodes at about 79 MB/s with 104
allocations per call, and the long array-of-tables document at about
106 MB/s against 56 MB/s in 1.1.0, with allocations on that document
halved from 67 664 to 31 765. Date-time tokens are validated by a byte
scan instead of regular expressions, repeated keys share one string
across array-of-tables elements, and per-statement buffers are reused.
### Fixed