Files
deepseek-harness/packages/code-runtime/code-runtime-python
Chinesezjc 295e020ea4 fix(code-runtime-python): reject only an oversized FIRST frame before the join, not a multi-frame buffer
The pre-join check charged the whole unframed buffer, which legitimately holds
several frames each within FRAME_PARSE_CAP_BYTES: a first frame of exactly the
cap followed by a second frame crossed the counter and was misreported as a
worker-exit. The pre-join rejection now fires only while the held bytes are a
single unframed line (this chunk carries no newline); once a newline arrives,
a FIRST-FRAME check measures the bytes up to the first newline across the held
chunks (including sealed blocks) and rejects only that frame before the join —
keeping the peak at one copy of its wire bytes — while later frames in the
same buffer are handled by the restored per-line check. Regression cases: a
72 MiB newline-free buffer is rejected pre-join (fail-before: joining would
have doubled it); two within-cap frames whose combined buffer crosses the cap
both survive (fail-before: the unconditional counter check turns it red).
2026-08-31 14:53:33 +08:00
..

description, kind
description kind
CPython subprocess implementation of the DeepSeek Harness code-execution seam, with fd-3 bindings, resource limits, log capture, and process-group teardown. package-reference

@deepseek-ai/dsh-code-runtime-python

English | 中文

CPython-subprocess implementation of the @deepseek-ai/dsh-code-runtime seam. Companion to @deepseek-ai/dsh-code-runtime-worker-thread; trades the Node worker thread for a fresh python3 subprocess so model code is Python instead of TypeScript.

The package owns the wire protocol for that seam: the host-side frame codec and the Python-side mirror of the same message vocabulary. On top of that protocol it ships PythonCodeRuntime (the plugin's default export), which registers as codeRuntime with language: 'python' and isolation: 'process'. Each run() spawns a fresh python3 -I process, sends a boot frame and the program over fd 3, and resolves a CodeRunResult for every program outcome — run() rejects only for seam misuse, such as a malformed binding namespace or a call on a runtime whose fiber was already disposed. Configuration is rejected earlier, when the plugin loads: a non-Unix platform, a non-positive or non-integer budget, a maxLogBytes below the truncation-marker floor (64), a timer value setTimeout would clamp, a budget larger than one fd-3 frame can carry, and an addressSpaceMb/output-budget pair whose worst-case peak would breach RLIMIT_AS all throw from the constructor, so a misconfiguration fails at assembly rather than on a later run. The child runs the program as the body of an async function, so top-level await and return both work; binding calls travel back over fd 3 as JSON-lines. Containment (not a security boundary — model code has bash-equivalent trust) comes from an empty environment, RLIMIT_CPU/RLIMIT_AS, a wall-clock ceiling, and a SIGTERM→grace→SIGKILL teardown on the child's process group.

Wire protocol

The host and the CPython subprocess exchange a versionless, JSON-lines protocol on the child's fd 3 — one JSON object per line, leaving stdout/stderr free for the program's own output. src/protocol.ts is the host side; py/protocol.py mirrors its message shapes and the shared truncation-marker text on the Python side.

  • fd 3, not stdout — Node pins the channel positionally with stdio: ['pipe','pipe','pipe','pipe']; the Python bootstrap reads the same PROTOCOL_FD constant. JSON-lines framing.
  • Host treats every inbound frame as hostile — model code has full access to fd 3 and can post anything through it, so validateChildFrame shape-validates and REBUILDS each frame before the host reads it: forged extra fields never ride along, a non-number call id can never be echoed into a reply, and junk drops to undefined rather than throwing in the host's message handler. The Python side trusts host replies (the host is not model-controlled).
  • Lossless-JSON crossing — completion values and binding arguments cross as exact JSON. encodeJsonPlain serializes a JSON.parse-produced value without recursion, so a deep value below the byte budget crosses intact instead of dying on JSON.stringify's stack limit; checkDoneValue meters a forged completion value's byte length AND number losslessness in one bounded traversal that rejects an over-budget payload before enqueuing its children; hasUnsafeIntegerToken reads the raw frame text to catch an integer token that JSON.parse would silently round; hasNonLosslessNumber rejects a non-finite or negative-zero number in unbounded call.args. Beyond-safe-range integral doubles serialize through BigInt digits so the exact integer crosses, not the rounded String() form.
  • Shared truncation markerlogTruncationMarker(maxBytes) produces byte-identical text on both sides, so a truncated log run reads the same however the cap was hit. The log frame's truncated flag distinguishes the child ledger's own marker from program output.

Configuration

Every cap is a validated Config field with a default, changeable from cordis.yml (no hardcoded tunables). cpuSeconds (default 60) is the RLIMIT_CPU whole-second budget; the child sets the soft limit to cpuSeconds and the hard limit to cpuSeconds + 1, so the kernel's SIGXCPU at the soft limit classifies as a timeout while the +1s hard limit is a SIGKILL backstop. maxWallMs (default 600000) is the wall-clock ceiling that backstops CPU time for a program awaiting a promise nobody resolves. addressSpaceMb (default 512) is the RLIMIT_AS cap, not applied on Darwin (the dyld shared cache mapped into every process exceeds any practical cap there; cpuSeconds and maxWallMs still bound the run). maxLogBytes (default 65536) is the shared captured-log byte budget; maxValueBytes (default 32768) caps the completion value; graceMs (default 3000) is the SIGTERMSIGKILL grace window; pythonBin (default python3) is the interpreter, resolved against PATH before the child spawns with an empty environment.

Model Experience

Indirectly, through Code Mode in dsh-tools, which renders this backend's exact completion value when it fits (or an explicit invalid-output / output-limit failure), plus the exact [dsh-code-runtime-python] log capture truncated at <maxLogBytes> bytes log marker, into a retained run_code result.

KV Cache effect

No direct invalidation; the named consumer owns any request-prefix changes.

Known Limitations and Deferred Work

  • The cross-language guard covers executed values and frame field sets, not field typestests/protocol-mirror.e2e.ts compares PROTOCOL_FD, the log truncation marker, and each TypedDict's required and optional fields against a real python3. Comparing field types across TypeScript and Python has no mechanical equivalent here, so review plus the backend's real-subprocess suite owns type-level drift.

  • RLIMIT_AS is not enforced on macOS — the dyld shared cache mapped into every process at exec exceeds any practical address-space cap, and the kernel rejects the setrlimit call, so addressSpaceMb is skipped there. cpuSeconds and maxWallMs still bound every run.

  • PID-reuse protection is inert on macOSreadProcessStart reads /proc/<pid>/stat, which Darwin does not provide, so the identity re-check that guards killGroup against signalling a recycled pgid always passes there; killGroup signals the pgid without the identity re-check on macOS rather than paying a ps fork on a teardown path. The process-group teardown and the closeDeadline bound still contain the run.

  • C-ext stdio buffers are not drained at settlement. The child runs with -u (unbuffered), so sys.__stdout__/sys.__stderr__ and os.write bytes are visible to the host's stray capture immediately; but a C extension's private C-stdio (FILE*) buffering is outside the interpreter, and its unwritten bytes are lost when the host SIGTERMs the child after the done frame. Model code should flush C-level stdio explicitly before returning if it must survive.

  • An fd-3 frame whose raw length exceeds 64 MiB is dropped before decoding. The receive path caps raw frames at FRAME_PARSE_CAP_BYTES before toString/JSON.parse (a compact wide frame near the 256 MiB wire ceiling could decode to far more host memory than the wire admitted). maxLogBytes/maxValueBytes are load-bounded to that parser cap so an honest child's frames always fit; a model-constructed binding ARGUMENT above 64 MiB (a value with no seam-level budget) is likewise dropped, stranding that call to the wall clock — an accepted residual of the same OOM guard.

  • A truncated log's serialized array runs to maxLogBytes plus the marker. The truncation marker is envelope, not payload — it rides uncharged so it can always be emitted — and the outer-array envelope is reserved one byte in the ledger. A truncated run with admitted entries therefore serializes its logs array to at most maxLogBytes + marker + 1; the marker alone fits any admissible budget (the 64-byte floor guarantees it).

  • A descendant that calls setsid() / start_new_session=True escapes teardown. Termination signals the child's process group with kill(-pid); a descendant that moves itself into a fresh session is no longer in that group and no signal reaches it. If it also releases the inherited stdout/stderr/fd-3 pipes, the leader's close still settles the run, and after the closeDeadline bound the fiber goes quiescent while that orphan keeps running. This is the containment boundary, not a security one — model code has bash-equivalent trust, and a bash tool can setsid away just the same. Reaching such an orphan would require tracking every descendant pid (as the bash-local backend's process-inspector does) and is deferred; the process-group teardown reaps everything that stays in the group.

  • A combined log-and-value peak is not modelled by the load gate. Each budget is checked against addressSpaceMb on its own. A model daemon thread that keeps writing while the completion value is metered and framed can refill the log pending toward maxLogBytes during that window, so the two peaks add in a way no gate admits or rejects. A gate over (maxLogBytes + maxValueBytes) was considered and deferred: its discriminating case cannot be scheduled deterministically under RLIMIT_AS, so the gate would only prove its own arithmetic. When the combined peak is reached the run dies as worker-exit -- containment holds and only the failure classification is degraded.

  • A 1-second dual-limit ulimit -t 1 CPU overrun is reported as worker-exit, not a timeout. When the host starts under a hard CPU limit equal to the soft (ulimit -t N sets both) and that limit is 1, _clamped cannot lower the soft to 0, so the kernel SIGKILLs the busy loop in the same tick and SIGXCPU is never delivered. The host classifies a CPU overrun only on signal === 'SIGXCPU', so the overrun is reported as worker-exit. For a dual limit of 2 or more the soft is lowered by one unit, SIGXCPU fires, and the run is a timeout. Containment holds in both cases; only the classification is degraded.

  • A program that traps SIGXCPU can exceed the soft CPU limit during settlement encoding and still report success. The settlement CPU recheck (die_if_cpu_exhausted) runs unconditionally after the program returns and before the log flush and completion encode; a program that exceeded the soft limit before returning is caught there and dies on the re-delivered SIGXCPU, classified as a timeout. The only false-success window is a program that PASSES the recheck and then, with SIGXCPU trapped, exceeds the soft limit during the settlement flush/encode window. A post-encode recheck is not done because it would charge the settlement encode's own CPU to the program, misclassifying a legitimate near-limit program. Containment holds — the hard limit (soft + 1s) and the wall clock still bound it — and only the classification is degraded.

  • The encoder's direct dependencies resolve at call time. _encode_json_plain reaches _dump_scalar/_dump_string/json via module-global lookup, so a program running as __main__ that rebinds one of those names (e.g. __main__._dump_scalar = boom) after returning a legitimate value can make the encode throw and downgrade a success to exception. The value path's entry name (_done_with_value) is bound into _run locals and its top-level _check_done_value/_encode_json_plain are def-time defaults, but the encoder's transitive deps (e.g. _dump_scalar/_dump_string/json/io — a non-exhaustive set) still resolve at call time. This is an accepted residual: under the bash-equivalent trust model a rebind here only harms the model's own run, and the verdict still reaches the host — send_done's fixed fallback frame delivers a done frame even when the error-path encode/write throws.

  • A wide binding REPLY expands host-side state per member. Resolutions cross through snapshotJsonValue in @deepseek-ai/dsh-session, whose walkJsonValue pushes one task frame per member, and binding resolution carries no seam-level byte cap. A legitimate reply of several million elements can therefore exhaust the host heap. The property belongs to that shared walk, not to this backend -- the worker-thread backend consumes the same function -- so the fix belongs in packages/core/session where every consumer benefits.

  • A cross-thread binding that the program joins with a synchronous t.join() can deadlock. This is specific to the process isolation backend: the reply pump runs on the child's main event loop, so when the program's main coroutine calls t.join() on a worker thread that is still awaiting a binding reply, the join blocks the main thread's event loop — the loop the pump needs to deliver that reply — and the worker's await never resumes until the wall clock. The worker-thread backend does not share this structure, so the fix belongs here, not in packages/core/session.