Grouped emission walks the entries slice in three passes instead of copying them into per-kind slices, error paths render their key path only when an error names it, inline tables are measured once because a measuring encoder no longer nests another measurement, clean string runs are written in one write, and date-times render through the single-buffer path. The output buffer comes from a sync.Pool and returns to it within a 1 MiB retention cap, the output handed to the caller as a copy; repeated marshals keep the live heap flat, verified over 500 runs. The encoder resolves Marshaler and the text interfaces through the same per-type flag cache and hint the decoder uses, and docs/ARCHITECTURE.md now lists the caches and the pool as the library's shared state.
Representative document: 177 to 141 allocations, 12 230 to 6170 bytes per operation; long document: 97 175 to 63 660 allocations, 6.97 to 2.79 MB per operation, about 7.8 to about 3.4 ms.
The decoder asked every value whether it implements Unmarshaler or encoding.TextUnmarshaler by boxing it into an interface and asserting, which allocated on every scalar field. A per-type flag cache answers first and an interface value is built only where the assertion can succeed; interface destinations are still asked dynamically. A monomorphic hint in front of the cache keeps the hot walk off the sync.Map probe, and it re-points at published cache entries so a miss allocates nothing.
Representative document: 233 to 167 allocations; long document typed decode: 107 674 to 63 772 allocations, about 6.1 to about 3.8 ms.
Per-statement work no longer allocates where it does not need to: parseKeyPath returns the first key component directly, absolute paths and key buffers sit on the parser, the definition maps are created on first write, arrays presize, and a string without escapes is copied out in one piece instead of built byte by byte. Repeated keys share one string across array-of-tables elements through a parser-local intern table whose lookup works on the bytes, so a repeated key costs no allocation.
UTF-8 validity is no longer a whole-input pass before the parser runs: the scan validates where it meets a multi-byte sequence, and a SyntaxError now names the line the invalid byte sits on. TestParseRejectsInvalidUTF8 covers nine positions; docs/API.md describes the new reporting.
Long document: 67 664 to 31 765 allocations per parse, about 51 to about 106 MB/s; representative document: 157 to 104 allocations.
Date-time tokens were validated by two regular expressions and then tried against up to sixteen time.Parse layouts; the profile named the regexp backtracker among the hottest nodes, and the failed attempts allocated ParseErrors by the million. scanDateTimeShape walks the strict TOML grammar as bytes and dispatches one layout per shape, which time.Parse accepts because parsing takes a fractional second whether the layout signs it or not. The String methods build their output in a single buffer instead of concatenating Format results.
The suite gains BenchmarkStrictDecodeLong and BenchmarkMarshalLong over the same 2000-entry document ParseLong uses, so long-input work is visible on every path, not the parse alone. docs/BENCHMARKING.md lists them.