The skill carries the reliability rules, but nothing an agent loads by default said that specs run concurrently at all. The testing policy described tiers and evidence without ever stating how a spec is executed, and neither subtree AGENTS.md mentioned it — including scripts/, where the two suites that recently failed on unrelated branches live. docs/testing.md gains the execution model as the one home for the fact: forked workers, concurrent coverage partitions beside other gates, and self-hosted runners sharing a host and volume, with the rule that a spec passing only when run alone is a defect in the spec. It links the skill for the detailed rules. packages/AGENTS.md and scripts/AGENTS.md carry the short actionable form and link that section, so the rule is present in the context loaded while a test in either subtree is being written. Both ceilings are raised for the added words and the targets in docs/AGENTS.md move with them: docs/testing.md 1150 to 1300 (now 1237) and packages/AGENTS.md 675 to 750 (now 712), each keeping the 5% headroom the standard requires.
9.7 KiB
Testing policy
English | 中文
How this repo tests, tier by tier, and the rules that keep a green suite meaningful. Commands live in root AGENTS.md; linked Agent Notes carry the rationale.
Tiers
- Unit (
pnpm run test): vitest over package and example specs under theirtests/**directories plus repository script specs underscripts/**/*.spec.ts; tests stay with the code area they exercise. Every registry gets an HMR-safety test (dispose the contributing fiber, assert cleanup). Prefer edge cases, error paths, event ordering, concurrency races, and permanent tests for contract regressions (seepackages/core/agent-loop/tests/contract-regressions.spec.ts). - Coverage gate (
pnpm run test:coverage): the gating run, per-file 100% onpackages/*/*/src. An uncovered line is often dead code the gate flags for deletion, not a missing test to bolt on. Line coverage is necessary, never sufficient — it proves lines ran, not that the feature works as shipped. Per-file 100% onpackages/shell/pwsh-local/srcneeds a realpwsh: without one its executor suites self-skip andvitest.config.tsexempts the file so pwsh-less hosts stay green, while CI runners ship pwsh and enforce the full bar. - Real-API e2e (
pnpm run test:e2e): with-key tests against live provider APIs — the DeepSeek model plus provider-specific smokes that gate on their own keys (EXA_API_KEY,PERPLEXITY_API_KEY, …); each suite self-skips without its key so keyless CI stays green (real-API e2e Agent Note). - Owner-local expected output (
pnpm run test:expected): keyless assembled CLI/process expectations without a recorded-session round trip. Drivers use*.expected.e2e.tsbesidetests/expected/; CI runs built exports. Package/script expectations usetest, while browser expectations usetest:web. - Snapshot (
pnpm run test:snapshot): a top-level scenario's recordedsession.jsonlsupplies user input and model replay, then serves as the expected persisted result. Process scenarios start throughdsh: headless owns one-shot behavior, the SDK owns persistent control, ACP owns automation-protocol behavior, and Web retains browser/ARIA evidence beside the same session.snapshot.ymldeclares the profile, composition/header class, recording policy, exceptional replay or input metadata, and workspace facts. Typed tokens preserve parent/child identity relationships; only header pins own prompt/schema sidecars. A mutating scenario independently compares the completeworkspace.expected/tree, which record and refresh never rewrite. Usetest:snapshot:recordwhen a model transcript changes andtest:snapshot:refreshwhen replay input remains valid; review every resulting diff. - Web browser snapshot (
pnpm run test:web; required Linux PR gate): Chromium compares session-driven output undersnapshots/web/and UI-only output underapps/web/tests/expected/. CI forces read-onlyDSH_SNAPSHOT=replay, never writing expected outputs; record/refresh stay local and every diff is reviewed (web e2e lane, CI gate decision).test:webbuilds first for plugin CSS.
Session fixtures keep headers and payloads but omit body sequence/time envelopes. Replay synthesizes them. Fixtures use canonical packed rows; the migrator rewrites old layouts.
How specs execute
Forked workers run several spec files at once, the coverage gate splits into concurrent partitions beside the other gates in its job, and the self-hosted runners share one host and one volume. Only the process is isolated: ports, predictable paths, external namespaces, and inherited children are not. Own each acquired resource through its teardown, and read a spec that passes only when it runs alone as a defect in the spec rather than an unstable runner. dsh-ci-test-reliability owns the allocation, restoration, synchronization, timeout-budget, platform, and teardown rules; its flake diagnosis workflow classifies an existing probabilistic failure.
The with-key policy: inference is cheap here
We are DeepSeek — do not ration real-API tests. A no-key test proves plumbing; only a with-key run proves the agent works against a real model. Cover file-writing prompts, multi-turn conversations, tool use, and mid-stream cancellation. Highest-value are smoke tests that boot a shipped dsh profile, send one prompt, and check the world — they catch the "green unit tests, broken product" class that mocks cannot (postmortem 0001). Self-skip keeps secretless CI and keyless contributors unblocked; it is not a cost signal. Profile-level integration tests live under apps/cli/tests/profiles/; package-specific compositions stay with their package tests.
Prefer the real implementation over a mock
Mock only the expensive or non-deterministic boundary (LLM adapter, network, clock); keep everything downstream real. A hand-rolled stand-in proves the bridge moves bytes, not that the shipping tool behaves as asserted. Bridge tool-call tests keep the real tool registry and pipeline behind the scripted mock model: makeBridgeHarness() mounts the loop, session store, tool registry, and JSONL persistence with a MockAdapter as the only mock (packages/acp/acp/tests/harness.ts).
Recovery tests separate pre/post-chunk failures by step and prove failed chunks derive no message or tool side effect. Cover exhaustion, cancellation, policy composition, persistence, status, wire counts, transport-closing idle timeouts, and shipping Loader composition.
Verify the world, not the self-report
An e2e assertion re-runs the command or re-reads the file externally; a keyword probe on the agent's own output lets a cheating agent pass. Assert untouched files are byte-identical. e2e tests own their resources: create it in the test, dispose in afterEach (even on failure/retry/timeout); shared fixtures live in a plain tests/harness.ts, never another *.e2e.ts (importing a spec re-registers its describe and duplicates real API calls).
Test the real entry path
- Product-visible plugins require a non-unit REAL-composition test. Hand-built
ctx.plugin(...)suites are insufficient: boot test-onlycordis.ymlthrough Loader and app/process, mock only external services or nondeterministic inputs, and assert model-visible request/log, durable state, or user-visible output. Keep opt-ins out of shipped defaults. - A guard only guards if the regression fails it. For a plugin without
inject(bundle/composition plugins), a Loader smoke stays green when a default export replaces the required named exports — add an explicitexpect('default' in mod).toBe(false)plus anunwrapExportsround-trip assertion, and prove it: introduce the regression, watch red, revert. - "Real entry path" means the published artifact: a package
binruns builtlib/bin.jsunder plainnode, exposing failures tsx masks (settle races, module resolution, swallowed load failures). The same applies to non-index runtime entries (the worker-thread siblinglib/worker.cjs) and singleton modules shared across bundles (packages/sdk/server/tests/built-scope-carrier.e2e.ts). Keep the built-artifact smokes green (packages/examples/*/tests/built-bin.e2e.ts,packages/code-runtime/code-runtime-worker-thread/tests/built-lib.e2e.ts), and assert a genuinely-missing config exits non-zero.
Test resolution: source plane only
- Every vitest config points vite-tsconfig-paths at
tsconfig.base.json; bare workspace imports resolve tosrc(layout), never through packageexportsto builtlib/— stale artifacts there load a second copy of module singletons. Built artifacts are consumed only explicitly:lib-mode subprocesses and the built smokes below.
Test subprocess launch modes
- CI and build-having test lanes run every profile or Cordis-config subprocess from built
lib/through the shared dual-mode launcher. Do not hand-write--import tsxfor these subprocesses. - Protocol and operating-system fixtures that do not load Cordis run erasable
.tsdirectly with Node, without tsx or the root paths map. - Only a test whose subject is source-path resolution may select
src; state that contract in the test.
When a snapshot test is required
Every non-trivial model-, protocol-, or human-visible change adds or updates a keyless recorded-session scenario in the same PR; package, e2e, mock-only, and rationale evidence does not replace the assembled transcript. Headless, SDK, ACP, and Web recordings live under snapshots/session/, snapshots/sdk/, snapshots/acp/, and snapshots/web/; a Web rendering may explicitly borrow another scenario's canonical session. Expected output that is not driven by a recorded session stays with its owning app, package, or script under tests/expected/ and does not use the *.snapshot.ts suffix. dsh-session-snapshot owns the shared storage rules and profile adapters. Agent-loop, session-lifecycle, and SessionEventMap changes update both SDK projections: snapshots/sdk/ owns TypeScript, while required Python-runtime CI owns scripts/snapshots/python-sdk-single-exe/. New capability seams and lifecycle or transcript variants name every required tier at plan time.