Files
deepseek-harness/packages/code-runtime/code-runtime-python/README.md
T
Chinesezjc 2a9a917853 fix(code-runtime-python): drop a late binding resolution before snapshotting it
`sendReply` already refuses to write after the run settled, but only after
`snapshotJsonValue` walked and copied the resolution. Binding resolution carries
no seam-level byte cap, so a binding resolving a wide value after `maxWallMs`,
an abort, or dispose settled the run spent host heap building a frame that was
then discarded. The check moves ahead of the snapshot.

Also in this change:

- `readProcessStart` moved after `messageOf`. Inserting it between `messageOf`'s
  JSDoc and its body left that function undocumented and the orphaned block
  reading as a second doc for the reader; `verify-export-jsdoc` does not catch it
  because `messageOf` is not exported.
- The README pair adds the disposed-runtime rejection to `run()`'s public
  contract, which `src/index.ts` has enforced all along.
- Known Limitations records three deferred constraints that until now existed
  only in review discussion: the combined log-and-value peak the load gate does
  not model, the host-side per-member expansion of a wide binding reply (owned by
  `packages/core/session`, and shared with the worker-thread backend), and the
  absence of fd-3 backpressure for concurrent replies.
- The Agent Note's same-group section records the teardown identity guard and its
  two rulings, including why an ABSENT start-time reading proceeds rather than
  withholding the signal, and that reading it as a mismatch is what turned the
  three same-group heartbeat cases red on Linux.
2026-08-31 14:31:48 +08:00

8.7 KiB

description, kind
description kind
CPython subprocess implementation of the DeepSeek Harness code-execution seam, with fd-3 bindings, resource limits, log capture, and process-group teardown. package-reference

@deepseek-ai/dsh-code-runtime-python

English | 中文

CPython-subprocess implementation of the @deepseek-ai/dsh-code-runtime seam. Companion to @deepseek-ai/dsh-code-runtime-worker-thread; trades the Node worker thread for a fresh python3 subprocess so model code is Python instead of TypeScript.

The package owns the wire protocol for that seam: the host-side frame codec and the Python-side mirror of the same message vocabulary. On top of that protocol it ships PythonCodeRuntime (the plugin's default export), which registers as codeRuntime with language: 'python' and isolation: 'process'. Each run() spawns a fresh python3 -I process, sends a boot frame and the program over fd 3, and resolves a CodeRunResult for every program outcome — run() rejects only for seam misuse, such as a malformed binding namespace or a call on a runtime whose fiber was already disposed. Configuration is rejected earlier, when the plugin loads: a non-Unix platform, a non-positive or non-integer budget, a timer value setTimeout would clamp, a budget larger than one fd-3 frame can carry, and an addressSpaceMb/output-budget pair whose worst-case peak would breach RLIMIT_AS all throw from the constructor, so a misconfiguration fails at assembly rather than on a later run. The child runs the program as the body of an async function, so top-level await and return both work; binding calls travel back over fd 3 as JSON-lines. Containment (not a security boundary — model code has bash-equivalent trust) comes from an empty environment, RLIMIT_CPU/RLIMIT_AS, a wall-clock ceiling, and a SIGTERM→grace→SIGKILL teardown on the child's process group.

Wire protocol

The host and the CPython subprocess exchange a versionless, JSON-lines protocol on the child's fd 3 — one JSON object per line, leaving stdout/stderr free for the program's own output. src/protocol.ts is the host side; py/protocol.py mirrors its message shapes and the shared truncation-marker text on the Python side.

  • fd 3, not stdout — Node pins the channel positionally with stdio: ['pipe','pipe','pipe','pipe']; the Python bootstrap reads the same PROTOCOL_FD constant. JSON-lines framing.
  • Host treats every inbound frame as hostile — model code has full access to fd 3 and can post anything through it, so validateChildFrame shape-validates and REBUILDS each frame before the host reads it: forged extra fields never ride along, a non-number call id can never be echoed into a reply, and junk drops to undefined rather than throwing in the host's message handler. The Python side trusts host replies (the host is not model-controlled).
  • Lossless-JSON crossing — completion values and binding arguments cross as exact JSON. encodeJsonPlain serializes a JSON.parse-produced value without recursion, so a deep value below the byte budget crosses intact instead of dying on JSON.stringify's stack limit; checkDoneValue meters a forged completion value's byte length AND number losslessness in one bounded traversal that rejects an over-budget payload before enqueuing its children; hasUnsafeIntegerToken reads the raw frame text to catch an integer token that JSON.parse would silently round; hasNonLosslessNumber rejects a non-finite or negative-zero number in unbounded call.args. Beyond-safe-range integral doubles serialize through BigInt digits so the exact integer crosses, not the rounded String() form.
  • Shared truncation markerlogTruncationMarker(maxBytes) produces byte-identical text on both sides, so a truncated log run reads the same however the cap was hit. The log frame's truncated flag distinguishes the child ledger's own marker from program output.

Configuration

Every cap is a validated Config field with a default, changeable from cordis.yml (no hardcoded tunables). cpuSeconds (default 60) is the RLIMIT_CPU whole-second budget; the child sets the soft limit to cpuSeconds and the hard limit to cpuSeconds + 1, so the kernel's SIGXCPU at the soft limit classifies as a timeout while the +1s hard limit is a SIGKILL backstop. maxWallMs (default 600000) is the wall-clock ceiling that backstops CPU time for a program awaiting a promise nobody resolves. addressSpaceMb (default 512) is the RLIMIT_AS cap, not applied on Darwin (the dyld shared cache mapped into every process exceeds any practical cap there; cpuSeconds and maxWallMs still bound the run). maxLogBytes (default 65536) is the shared captured-log byte budget; maxValueBytes (default 32768) caps the completion value; graceMs (default 3000) is the SIGTERMSIGKILL grace window; pythonBin (default python3) is the interpreter, resolved against PATH before the child spawns with an empty environment.

Model Experience

Indirectly, through Code Mode in dsh-tools, which renders this backend's exact completion value when it fits (or an explicit invalid-output / output-limit failure), plus the exact [dsh-code-runtime-python] log capture truncated at <maxLogBytes> bytes log marker, into a retained run_code result.

KV Cache effect

No direct invalidation; the named consumer owns any request-prefix changes.

Known Limitations and Deferred Work

  • The cross-language guard covers executed values and frame field sets, not field typestests/protocol-mirror.e2e.ts compares PROTOCOL_FD, the log truncation marker, and each TypedDict's required and optional fields against a real python3. Comparing field types across TypeScript and Python has no mechanical equivalent here, so review plus the backend's real-subprocess suite owns type-level drift.
  • RLIMIT_AS is not enforced on macOS — the dyld shared cache mapped into every process at exec exceeds any practical address-space cap, and the kernel rejects the setrlimit call, so addressSpaceMb is skipped there. cpuSeconds and maxWallMs still bound every run.
  • A descendant that calls setsid() / start_new_session=True escapes teardown. Termination signals the child's process group with kill(-pid); a descendant that moves itself into a fresh session is no longer in that group and no signal reaches it. If it also releases the inherited stdout/stderr/fd-3 pipes, the leader's close still settles the run, and after the closeDeadline bound the fiber goes quiescent while that orphan keeps running. This is the containment boundary, not a security one — model code has bash-equivalent trust, and a bash tool can setsid away just the same. Reaching such an orphan would require tracking every descendant pid (as the bash-local backend's process-inspector does) and is deferred; the process-group teardown reaps everything that stays in the group.
  • A combined log-and-value peak is not modelled by the load gate. Each budget is checked against addressSpaceMb on its own. A model daemon thread that keeps writing while the completion value is metered and framed can refill the log pending toward maxLogBytes during that window, so the two peaks add in a way no gate admits or rejects. A gate over (maxLogBytes + maxValueBytes) was considered and deferred: its discriminating case cannot be scheduled deterministically under RLIMIT_AS, so the gate would only prove its own arithmetic. When the combined peak is reached the run dies as worker-exit -- containment holds and only the failure classification is degraded.
  • A wide binding REPLY expands host-side state per member. Resolutions cross through snapshotJsonValue in @deepseek-ai/dsh-session, whose walkJsonValue pushes one task frame per member, and binding resolution carries no seam-level byte cap. A legitimate reply of several million elements can therefore exhaust the host heap. The property belongs to that shared walk, not to this backend -- the worker-thread backend consumes the same function -- so the fix belongs in packages/core/session where every consumer benefits.
  • Concurrent binding replies are not paced against fd 3. proto.write returns false once the pipe's buffer is full and this backend does not wait for drain, so several bindings resolving large values in one asyncio.gather round encode and queue together in host memory. Serializing the replies would bound it, at the cost of changing the concurrency the seam currently allows; the sibling worker-thread backend has no equivalent (it posts structured clones, which carry no stream backpressure), so there is no in-repo precedent to copy.