mirror of
https://github.com/deepseek-ai/deepseek-harness.git
synced 2026-09-09 04:02:35 +00:00
Add a benchmark lane (`vitest.bench.config.ts`, `pnpm run test:bench`, gate mode `ci-bench`) and a required `node 24 / benchmarks` CI job that runs it alone. Benchmarks synthesize their input in-process from fixed parameters and fail on documented budgets: - `open-generation.bench.ts`: a 200-turn released-v0 log with 500 text and 125 reasoning deltas per reply (127,400 events, ~2.8 MB) encoded through the frozen v0 codec; the migrating first `open()` must finish within 2,000 ms in a child process capped at 128 MB of old space, and a fresh process must open the published current generation within 500 ms. - `conversation-fold.bench.client.ts`: 200 replies whose compact streams hold 2,000 text + 500 reasoning deltas each, folded through every Chat Definition by the real assembler; the fold must finish within 150 ms and stay within 3x the fold of the same window with 100 deltas per reply. On this commit both gates fail: the migration exhausts the 128 MB heap (4.8 s and 696 MB peak RSS without the cap; the pre-stack decode of the same bytes took 34 ms and 168 MB) and the fold scales 11x with the delta count. The stacked fixes bring both paths to O(records).