mirror of
https://github.com/deepseek-ai/deepseek-harness.git
synced 2026-09-11 04:00:38 +00:00
Three separate paths in the CPython child allocated state proportional to a value's width or a string's length, so a legitimate input the byte budgets admit could die as the program's own MemoryError. `_lossless_json_violation` enqueued one traversal tuple per member while running, in `dispatch`, over MODEL-CONSTRUCTED binding arguments that no child-side byte budget bounds first. It now uses the same (kind, container, iterator) cursor the other two walks already had, checking dict keys as the cursor pulls each entry. Measured over `[0] * 6_000_000` (~17 MB of JSON): 459.1 MiB of traversal tuples before, 0.0 MiB after. `_decode_json_plain` matched JSON strings with a `(?:[^"\\]|\\.)*` repetition, which makes CPython's engine retain backtracking state proportional to the string's width: 146 MiB for a 1 MiB string, 557.8 MiB for 4 MiB. A legitimate multi-megabyte binding reply raised MemoryError inside `_pump_replies`, and because that pump is the only settler of the call's future, the run stranded until the wall clock reported `timeout`. Strings now scan chunk-to-chunk over a character class, which the engine matches without backtracking state; the same 4 MiB decode peaks at the 4.0 MiB result. `_check_done_value` charged strings and dict keys what `_dump_string(...).encode()` returned, building the escaped copy plus its encode to MEASURE it -- ~6x the original each for control-heavy text, so metering a value the budget then rejects could itself breach RLIMIT_AS and report `exception` where the seam promises `output-limit`. The new `_json_str_cost` counts instead, reusing `_json_string_cost`'s C-level passes and reproducing `_dump_string`'s exact surrogate rules (fold spelled-out pairs, charge six ASCII bytes per lone surrogate). Identical values, 228.9 MiB -> 19.1 MiB of peak on a 20M-NUL string. Each fix ships a regression test. The two RLIMIT_AS repros are Linux-only: Darwin does not apply the limit, so the peaks above are measured directly and recorded in the test comments.