perf(parse): cut allocations and validate UTF-8 in the scan
Per-statement work no longer allocates where it does not need to: parseKeyPath returns the first key component directly, absolute paths and key buffers sit on the parser, the definition maps are created on first write, arrays presize, and a string without escapes is copied out in one piece instead of built byte by byte. Repeated keys share one string across array-of-tables elements through a parser-local intern table whose lookup works on the bytes, so a repeated key costs no allocation. UTF-8 validity is no longer a whole-input pass before the parser runs: the scan validates where it meets a multi-byte sequence, and a SyntaxError now names the line the invalid byte sits on. TestParseRejectsInvalidUTF8 covers nine positions; docs/API.md describes the new reporting. Long document: 67 664 to 31 765 allocations per parse, about 51 to about 106 MB/s; representative document: 157 to 104 allocations.
This commit is contained in:
+3
-4
@@ -24,7 +24,6 @@ import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"unicode/utf8"
|
||||
)
|
||||
|
||||
// A SyntaxError describes a malformed TOML document, including the 1-based
|
||||
@@ -143,9 +142,9 @@ func parseWithOptions(ctx context.Context, data []byte, opts parseOptions, wantD
|
||||
if opts.maxInputSize > 0 && len(data) > opts.maxInputSize {
|
||||
return nil, nil, fmt.Errorf("interpres: input is %d bytes, over the limit of %d", len(data), opts.maxInputSize)
|
||||
}
|
||||
if !utf8.Valid(data) {
|
||||
return nil, nil, &SyntaxError{Line: 1, Msg: "input is not valid UTF-8"}
|
||||
}
|
||||
// UTF-8 validity is not checked in a pass of its own: the scanner
|
||||
// validates the multi-byte sequences where it meets them, so an invalid
|
||||
// byte is reported on its own line instead of always on line 1.
|
||||
maxDepth := opts.maxDepth
|
||||
if maxDepth <= 0 {
|
||||
maxDepth = maxNestingDepth
|
||||
|
||||
Reference in New Issue
Block a user