mirror of
https://github.com/deepseek-ai/deepseek-harness.git
synced 2026-08-29 04:26:38 +00:00
Merge commit '74a6e1e5cf520cdde8bf461d9051c81142fd6afe' into codex/dsh-sdk-minimal-diagnostics
This commit is contained in:
+2
-2
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.md
|
||||
2026-07-29-projected-token-usage-and-request-context.md: f2179885512bcb216ecb191ce98b535db571807a
|
||||
2026-07-29-projected-token-usage-and-request-context.zh.md: e4435b6245d1e20b51fc2cc1d73151ced8d94731
|
||||
2026-07-29-projected-token-usage-and-request-context.md: 063f2300f378f6f7763bce87b11add5da3093230
|
||||
2026-07-29-projected-token-usage-and-request-context.zh.md: 37b8741d09e9ec56f6b9f273e05460b2deb4f6f9
|
||||
|
||||
+4
-2
@@ -14,7 +14,9 @@ Context occupancy needs a numerator and a denominator that no existing surface c
|
||||
|
||||
Both values are ordinary durable session-projection state. `@deepseek-ai/dsh-token-meter` registers two units when `ctx.sessionProjections` is present.
|
||||
|
||||
`tokenUsage` folds the complete durable log into uncached input, output, cache-read, and cache-write buckets. An `assistant/chunk` usage sample survives a later failed request; an `assistant/message` usage value for the same `(turn, step)` replaces the earlier sample instead of double-counting it. Reasoning stays an output subdivision. Compaction and surface replacement do not erase earlier billing.
|
||||
`tokenUsage` folds the complete durable log into uncached input, output, cache-read, and cache-write buckets. An `assistant/chunk` usage sample survives a later failed request; an `assistant/message` usage value replaces the earlier sample from the same model attempt instead of double-counting it. A matching `llm/retry-started` boundary ends that replacement scope, so a retry with the same `(turn, step)` contributes a new attempt. Reasoning stays an output subdivision. Compaction and surface replacement do not erase earlier billing.
|
||||
|
||||
Token-meter also owns the shared pure attempt/Turn fold over durable events. It applies the same retry boundary while adding the stricter completeness and exact-total checks required by an exact per-Turn disclosure. A presentation consumer may select a complete Turn window and invoke that fold, but does not own or duplicate the accounting semantics.
|
||||
|
||||
`contextPressure` carries optional `pressureTokens` — the newest provider-reported prompt size, summing uncached input plus cache reads and writes, excluding output — and optional `contextWindow` from the newest `request/context` record. Neither field is synthesized before its source exists.
|
||||
|
||||
@@ -56,4 +58,4 @@ Token totals stay stable across pagination, compaction, replay, restart, and rec
|
||||
|
||||
Occupancy is approximate in the ways documented above. It is available immediately after restore or reconnect, since both fields are durable, at the cost of describing the last recorded request rather than an exact current boundary.
|
||||
|
||||
Each session log gains one small `request/context` record per route or advertised-capacity change. The token-meter projection is the canonical owner of durable session-projection usage semantics; the TUI retains its live per-step map because it does not mount the generic projection seam, and the standalone browser fixture mirrors the unit. ApiProxy carries no token-specific code, owns no per-session metrics cache, and performs no measurement. The browser keeps two generic projection values and no connection-local telemetry, and streaming text deltas still do not force the stats line to recompute.
|
||||
Each session log gains one small `request/context` record per route or advertised-capacity change. Token-meter is the canonical owner of durable usage semantics, including retry-attempt separation in the cumulative projection and the reusable exact attempt/Turn fold; Web Chat only selects a complete loaded Turn and renders the fold result. The TUI retains its live per-step map because it does not mount the generic projection seam, and the standalone browser fixture mirrors the unit. ApiProxy carries no token-specific code, owns no per-session metrics cache, and performs no measurement. The browser keeps two generic projection values and no connection-local telemetry, and streaming text deltas still do not force the stats line to recompute.
|
||||
|
||||
+4
-2
@@ -14,7 +14,9 @@ Web 统计行原先从当前已加载的会话节点推导 token 总量。该窗
|
||||
|
||||
这两个值都是普通的持久会话投影状态。当 `ctx.sessionProjections` 存在时,`@deepseek-ai/dsh-token-meter` 会注册两个单元。
|
||||
|
||||
`tokenUsage` 将完整持久日志归并为未缓存输入、输出、缓存读取和缓存写入四类计数项。即使后续请求失败,`assistant/chunk` 用量样本仍会保留;同一 `(turn, step)` 的 `assistant/message` 用量值会替换先前样本,不会重复计数。推理(reasoning)仍是输出的细分项。压缩和表层替换不会抹除先前的计费用量。
|
||||
`tokenUsage` 将完整持久日志归并为未缓存输入、输出、缓存读取和缓存写入四类计数项。即使后续请求失败,`assistant/chunk` 用量样本仍会保留;`assistant/message` 用量值会替换同一次模型 attempt 的先前样本,不会重复计数。匹配的 `llm/retry-started` 边界会结束该替换作用域,因此复用同一 `(turn, step)` 的重试会贡献一次新的 attempt。推理(reasoning)仍是输出的细分项。压缩和表层替换不会抹除先前的计费用量。
|
||||
|
||||
token-meter 还拥有在持久事件上运行的共享纯 attempt/Turn fold。它采用相同的重试边界,并增加精确单轮次 disclosure 所需的更严格完整性与精确总量检查。展示消费方可以选择完整 Turn 窗口并调用该 fold,但不拥有或复制记账语义。
|
||||
|
||||
`contextPressure` 携带可选的 `pressureTokens`(提供方报告的最新提示词规模,为未缓存输入加缓存读取与写入之和,不含输出),以及来自最新一条 `request/context` 记录的可选 `contextWindow`。在各自来源出现前,两个字段都不会被合成。
|
||||
|
||||
@@ -56,4 +58,4 @@ token 总量在分页、压缩、回放、重启和重连期间保持稳定,
|
||||
|
||||
占用率在上文记录的意义上是近似值。由于两个字段都是持久的,它在恢复或重连后立即可用;代价是它描述的是最后一条已记录的请求,而不是精确的当前边界。
|
||||
|
||||
每个会话日志会为每次路由或已公布容量变化增加一条小型 `request/context` 记录。token-meter 投影是持久会话投影用量语义的正典所有方;TUI 未挂载通用投影 seam,因此保留自己的实时逐步骤 map,而独立浏览器 fixture(测试前置数据)会镜像该单元。ApiProxy 不携带任何 token 专用代码,不拥有逐会话指标缓存,也不执行测量。浏览器只保留两个通用投影值,不保留连接本地的遥测数据;流式文本增量仍不会迫使统计行重新计算。
|
||||
每个会话日志会为每次路由或已公布容量变化增加一条小型 `request/context` 记录。token-meter 是持久用量语义的正典所有方,包括累计投影中的重试 attempt 分离,以及可复用的精确 attempt/Turn fold;Web Chat 只选择已完整加载的 Turn 并渲染 fold 结果。TUI 未挂载通用投影 seam,因此保留自己的实时逐步骤 map,而独立浏览器 fixture(测试前置数据)会镜像该单元。ApiProxy 不携带任何 token 专用代码,不拥有逐会话指标缓存,也不执行测量。浏览器只保留两个通用投影值,不保留连接本地的遥测数据;流式文本增量仍不会迫使统计行重新计算。
|
||||
|
||||
+2
-2
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-09-headless-direct-core-entry-point.md
|
||||
2026-08-09-headless-direct-core-entry-point.md: 8ed979794afa008588d1b849f0074e8696e6e43f
|
||||
2026-08-09-headless-direct-core-entry-point.zh.md: 512d4b88c921431fe26afd9f62c34a1939ac5bdd
|
||||
2026-08-09-headless-direct-core-entry-point.md: 9c17b8d418924c38174b4d958fd54b057b118019
|
||||
2026-08-09-headless-direct-core-entry-point.zh.md: d95978a832d52b26b1139cabb4b23ade93ce0da3
|
||||
|
||||
+5
-5
@@ -6,7 +6,7 @@ English | [中文](2026-08-09-headless-direct-core-entry-point.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The `headless` product contract is one local task with final assistant text on stdout, a success-sensitive exit code, empty stderr on success, and no listening port. A composition containing Workspace Host services, ApiProxy, HTTP, the Web runtime, or browser plugins contradicts that contract and makes local completion depend on an unrelated transport tree.
|
||||
The `headless` product contract is one local task with final assistant text on stdout, a success-sensitive exit code, no listening port, and the stderr reasoning projection owned by [headless reasoning progress](../feature/2026-08-21-headless-reasoning-progress.md). A composition containing Workspace Host services, ApiProxy, HTTP, the Web runtime, or browser plugins contradicts that contract and makes local completion depend on an unrelated transport tree.
|
||||
|
||||
The direct entry point still needs the same deployment model state as Web-created Agents. A separate provider/model default would give one deployment two answers, while deriving completion before the Agent and Session persistence are quiescent permits stdout and the exit code to observe incomplete state.
|
||||
|
||||
@@ -14,17 +14,17 @@ The direct entry point still needs the same deployment model state as Web-create
|
||||
|
||||
The shipped `headless` profile contains `dsh-base` and `dsh-headless`. The base supplies the disabled module-HMR default; the headless bundle supplies its persona and tool mode, mounts the Code Mode worker explicitly, and inserts `headless-runner` without overriding that policy. Its tree contains no `@deepseek-ai/dsh-host-*` package, ApiProxy, HTTP server, Web runtime, or browser client. Code Mode and Session persistence are one-shot Agent capabilities independent of Web presentation.
|
||||
|
||||
`headless-runner` is a direct core entry point. After Loader settlement, it reads `ctx.agentDefaultModel.currentSelection()`, creates a fresh persisted Agent through `ctx.agents.create`, installs that `ModelSelection` in the Agent scope, waits for startup quiescence, anchors the Session sequence, submits one ordinary user message, and waits for quiescence again. It awaits `ctx.sessions.flush`, folds its durable event interval for the last non-empty assistant text and final `turn/end` reason, writes the text plus one newline to stdout, and requests bounded launcher shutdown with exit 0 exactly when the reason is `completed`. A terminal `error` reason writes its durable code and message to stderr; unexpected driver failures also use stderr and exit 1.
|
||||
`headless-runner` is a direct core entry point. After Loader settlement, it reads `ctx.agentDefaultModel.currentSelection()`, creates a fresh persisted Agent through `ctx.agents.create`, installs that `ModelSelection` in the Agent scope, waits for startup quiescence, anchors the Session sequence, submits one ordinary user message, and waits for quiescence again. It awaits `ctx.sessions.flush`, folds its durable event interval for the last non-empty assistant text and final `turn/end` reason, writes the text plus one newline to stdout, and requests bounded launcher shutdown with exit 0 exactly when the reason is `completed`. [Headless reasoning progress](../feature/2026-08-21-headless-reasoning-progress.md) owns the live stderr projection; a terminal `error` reason writes its durable code and message there, and unexpected driver failures also use stderr and exit 1.
|
||||
|
||||
`@deepseek-ai/dsh-agent-default-model` owns the transport-independent default used for an Agent without a session-local selection. `AgentDefaultModelConfig` provides `ctx.agentDefaultModel` and registers the `agent-default-model` Settings section. Composition config supplies `{provider, model}`; user settings may also supply `reasoningEffort`. `currentSelection()` returns the live complete selection and `saveSelection()` writes it as a complete section, so a selection without an effort clears any stored effort. `dsh-base` supplies the composition entry. Direct and ApiProxy entry points consume this service; ApiProxy alone owns session-local precedence, model validation, and persistence of accepted Web selections.
|
||||
|
||||
`loadProfile` recognizes the exact installation-owned headless tuple (`dsh-base`, `dsh-web-app`, `dsh-headless`) and normalizes it to the shipped headless template while preserving every other manifest field. Extra, missing, or reordered bundle lists are user-owned and remain untouched.
|
||||
|
||||
This note owns the headless transport and completion contracts. [Apps own their command lines](2026-08-06-app-owned-command-line.md) owns the current `dsh --profile headless` grammar; the former [`dsh run` decision](../../archived/feature/2026-08-08-dsh-run-headless-command.md) records the superseded launcher-owned grammar, [GUI layering and RPC protocol](2026-07-19-gui-layering-and-rpc-protocol.md) owns browser gateway boundaries, [web config-tree boot and transport layering](2026-07-24-web-config-tree-boot-and-transport-layering.md) owns the Web tree, and [the default model follows the picker](../feature/2026-08-07-default-model-follows-the-picker.md) owns persistence of the shared Agent default.
|
||||
This note owns the headless transport and completion contracts; [headless reasoning progress](../feature/2026-08-21-headless-reasoning-progress.md) owns successful stderr output. [Apps own their command lines](2026-08-06-app-owned-command-line.md) owns the current `dsh --profile headless` grammar; the former [`dsh run` decision](../../archived/feature/2026-08-08-dsh-run-headless-command.md) records the superseded launcher-owned grammar, [GUI layering and RPC protocol](2026-07-19-gui-layering-and-rpc-protocol.md) owns browser gateway boundaries, [web config-tree boot and transport layering](2026-07-24-web-config-tree-boot-and-transport-layering.md) owns the Web tree, and [the default model follows the picker](../feature/2026-08-07-default-model-follows-the-picker.md) owns persistence of the shared Agent default.
|
||||
|
||||
## Verification
|
||||
|
||||
Package tests use the real Session store and Agent registry around a scripted Agent factory to pin idle-to-idle aggregation, late asynchronous completion, terminal model diagnostics, other non-completed exits, direct failures, Loader-time disposal, and flush-before-exit ordering. The keyless assembled snapshots drive `dsh --profile headless` through a replayed tool round trip, record a `user/message` with `source.kind: 'user'`, and expose a terminal model failure on stderr. Built-bin acceptance reaches a mock provider through the published entry and requires final text on stdout, exit 0, and empty stderr. Config-dump acceptance excludes every Host, Web, and Client package from the shipped headless tree; PTY shutdown coverage requires no observation line and bounded disposal.
|
||||
Package tests use the real Session store and Agent registry around a scripted Agent factory to pin idle-to-idle aggregation, late asynchronous completion, terminal model diagnostics, other non-completed exits, direct failures, Loader-time disposal, and flush-before-exit ordering. The keyless assembled snapshots drive `dsh --profile headless` through a replayed tool round trip, record a `user/message` with `source.kind: 'user'`, and expose both reasoning progress and a terminal model failure on stderr. Built-bin acceptance reaches a mock DeepSeek endpoint through the published entry and requires streamed reasoning on stderr, final text on stdout, and exit 0. Config-dump acceptance excludes every Host, Web, and Client package from the shipped headless tree; PTY shutdown coverage requires no observation line and bounded disposal.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
@@ -39,6 +39,6 @@ Package tests use the real Session store and Agent registry around a scripted Ag
|
||||
|
||||
## Consequences
|
||||
|
||||
`dsh --profile headless` provides a local Agent task rather than browser observation, Host APIs, or HTTP. Users who need those capabilities choose `dsh web`. Successful stderr is empty, completion follows durable flush, and the persisted Session remains available to later tooling. Its initial user message records `source.kind: 'user'` and therefore carries no ApiProxy `rpcId`.
|
||||
`dsh --profile headless` provides a local Agent task rather than browser observation, Host APIs, or HTTP. Users who need those capabilities choose `dsh web`. Text-only successful runs leave stderr empty, reasoned runs stream the provider-reported content there, completion follows durable flush, and the persisted Session remains available to later tooling. Its initial user message records `source.kind: 'user'` and therefore carries no ApiProxy `rpcId`.
|
||||
|
||||
ApiProxy carrier coverage stays in the ApiProxy package. Custom one-shot profiles may include Host or Web bundles explicitly, while the shipped profile and the recognized installation-owned tuple are Web-free.
|
||||
|
||||
+5
-5
@@ -6,7 +6,7 @@ Status: implemented
|
||||
|
||||
## 问题
|
||||
|
||||
`headless` 的产品约定是一个本地任务:最终 assistant 文本写入 stdout,退出状态反映成功与否,成功时 stderr 为空,并且不打开监听端口。包含 Workspace Host 服务、ApiProxy、HTTP、Web 运行时或浏览器插件的组合违背这一约定,也使本地完成状态依赖无关的传输树。
|
||||
`headless` 的产品约定是一个本地任务:最终 assistant 文本写入 stdout,退出状态反映成功与否,不打开监听端口,并由 [headless 推理进度](../feature/2026-08-21-headless-reasoning-progress.zh.md)负责 stderr 推理投影。包含 Workspace Host 服务、ApiProxy、HTTP、Web 运行时或浏览器插件的组合违背这一约定,也使本地完成状态依赖无关的传输树。
|
||||
|
||||
直接入口仍需要与 Web 所创建 Agent 相同的部署模型状态。独立的提供方/模型默认值会让同一部署产生两种答案,而在 Agent 与会话持久化完全停稳之前推导完成状态,会让 stdout 与退出状态观察到不完整状态。
|
||||
|
||||
@@ -14,17 +14,17 @@ Status: implemented
|
||||
|
||||
随附的 `headless` profile 包含 `dsh-base` 与 `dsh-headless`。base 提供默认禁用模块 HMR(热模块替换)的策略;headless 组合包提供自身的 persona 与工具模式、显式挂载 Code Mode worker,并在不覆盖该策略的情况下插入 `headless-runner`。其插件树不包含任何 `@deepseek-ai/dsh-host-*` 包、ApiProxy、HTTP server、Web 运行时或浏览器客户端。Code Mode 与会话持久化均为独立于 Web 呈现的一次性 Agent 能力。
|
||||
|
||||
`headless-runner` 是直接使用核心服务的入口。Loader 完全加载后,它读取 `ctx.agentDefaultModel.currentSelection()`,通过 `ctx.agents.create` 创建一个新的持久化 Agent,在 Agent 作用域中安装该 `ModelSelection`,等待启动工作完全停稳,锚定会话事件序号,提交一条普通用户消息,再次等待完全停稳。随后,它等待 `ctx.sessions.flush`,折叠自身持有的持久事件区间,以取得最后一条非空 assistant 文本和最终 `turn/end` 结束原因,将文本连同一个换行写入 stdout,并且仅在结束原因为 `completed` 时请求启动器以退出状态 0 有界关闭。结束原因为 `error` 时,其持久化错误码与消息写入 stderr;驱动器的意外失败也写入 stderr 并以 1 退出。
|
||||
`headless-runner` 是直接使用核心服务的入口。Loader 完全加载后,它读取 `ctx.agentDefaultModel.currentSelection()`,通过 `ctx.agents.create` 创建一个新的持久化 Agent,在 Agent 作用域中安装该 `ModelSelection`,等待启动工作完全停稳,锚定会话事件序号,提交一条普通用户消息,再次等待完全停稳。随后,它等待 `ctx.sessions.flush`,折叠自身持有的持久事件区间,以取得最后一条非空 assistant 文本和最终 `turn/end` 结束原因,将文本连同一个换行写入 stdout,并且仅在结束原因为 `completed` 时请求启动器以退出状态 0 有界关闭。[Headless 推理进度](../feature/2026-08-21-headless-reasoning-progress.zh.md)负责实时 stderr 投影;结束原因为 `error` 时,其持久化错误码与消息写入 stderr,驱动器的意外失败也写入 stderr 并以 1 退出。
|
||||
|
||||
`@deepseek-ai/dsh-agent-default-model` 拥有与传输无关的默认值,供没有会话级选择的 Agent 使用。`AgentDefaultModelConfig` 提供 `ctx.agentDefaultModel` 并注册 `agent-default-model` Settings 分节。组合配置提供 `{provider, model}`,用户设置还可以提供 `reasoningEffort`。`currentSelection()` 返回当前的完整选择,`saveSelection()` 则写入完整分节,因此不含强度的选择会清除已存强度。`dsh-base` 提供组合条目。直接入口与 ApiProxy 入口均消费该服务;只有 ApiProxy 负责会话级优先级、模型校验与已接受 Web 选择的持久化。
|
||||
|
||||
`loadProfile` 识别安装过程拥有的精确 headless 元组(`dsh-base`、`dsh-web-app`、`dsh-headless`),将其规范化为随附的 headless 模板,并保留 manifest(元数据清单)的其他所有字段。带额外项、缺少项或顺序不同的组合包列表归用户所有,保持不变。
|
||||
|
||||
本 Agent Note 负责 headless 的传输与完成约定。[应用持有自己的命令行](2026-08-06-app-owned-command-line.zh.md)负责当前的 `dsh --profile headless` 语法;原 [`dsh run` 决策](../../archived/feature/2026-08-08-dsh-run-headless-command.md)记录已被取代的启动器持有语法,[GUI 分层与 RPC 协议](2026-07-19-gui-layering-and-rpc-protocol.zh.md)负责浏览器网关边界,[Web 配置树启动与传输分层](2026-07-24-web-config-tree-boot-and-transport-layering.zh.md)负责 Web 插件树,[默认模型跟随选择器](../feature/2026-08-07-default-model-follows-the-picker.zh.md)负责共享 Agent 默认值的持久化。
|
||||
本 Agent Note 负责 headless 的传输与完成约定;[headless 推理进度](../feature/2026-08-21-headless-reasoning-progress.zh.md)负责成功运行时的 stderr 输出。[应用持有自己的命令行](2026-08-06-app-owned-command-line.zh.md)负责当前的 `dsh --profile headless` 语法;原 [`dsh run` 决策](../../archived/feature/2026-08-08-dsh-run-headless-command.md)记录已被取代的启动器持有语法,[GUI 分层与 RPC 协议](2026-07-19-gui-layering-and-rpc-protocol.zh.md)负责浏览器网关边界,[Web 配置树启动与传输分层](2026-07-24-web-config-tree-boot-and-transport-layering.zh.md)负责 Web 插件树,[默认模型跟随选择器](../feature/2026-08-07-default-model-follows-the-picker.zh.md)负责共享 Agent 默认值的持久化。
|
||||
|
||||
## 验证
|
||||
|
||||
包测试围绕脚本化 Agent 工厂使用真实的会话存储与 Agent 注册表,固定空闲态到空闲态的聚合、延迟异步完成、终止态模型诊断、其他未完成退出、直接失败、Loader 加载期间的 dispose(资源释放),以及退出前 flush 的顺序。组装后的无密钥快照通过回放的工具往返驱动 `dsh --profile headless`,记录一条带 `source.kind: 'user'` 的 `user/message`,并在 stderr 暴露终止态模型失败。构建后二进制验收通过已发布入口访问 mock 提供方,并要求最终文本出现在 stdout、退出状态为 0 且 stderr 为空。配置转储验收排除随附 headless 树中的所有 Host、Web 与 Client 包;PTY 关闭覆盖要求不出现观察行,并在有界时间内完成 dispose。
|
||||
包测试围绕脚本化 Agent 工厂使用真实的会话存储与 Agent 注册表,固定空闲态到空闲态的聚合、延迟异步完成、终止态模型诊断、其他未完成退出、直接失败、Loader 加载期间的 dispose(资源释放),以及退出前 flush 的顺序。组装后的无密钥快照通过回放的工具往返驱动 `dsh --profile headless`,记录一条带 `source.kind: 'user'` 的 `user/message`,并在 stderr 暴露推理进度与终止态模型失败。构建后二进制验收通过已发布入口访问 mock DeepSeek 端点,并要求推理流出现在 stderr、最终文本出现在 stdout 且退出状态为 0。配置转储验收排除随附 headless 树中的所有 Host、Web 与 Client 包;PTY 关闭覆盖要求不出现观察行,并在有界时间内完成 dispose。
|
||||
|
||||
## 考虑过的替代方案
|
||||
|
||||
@@ -39,6 +39,6 @@ Status: implemented
|
||||
|
||||
## 后果
|
||||
|
||||
`dsh --profile headless` 提供本地 Agent 任务,而不是浏览器观察、Host API 或 HTTP。需要这些能力的用户选择 `dsh web`。成功时 stderr 为空,完成结果在持久化 flush 后推导,持久化会话仍可供后续工具使用。初始用户消息记录 `source.kind: 'user'`,因此不携带 ApiProxy `rpcId`。
|
||||
`dsh --profile headless` 提供本地 Agent 任务,而不是浏览器观察、Host API 或 HTTP。需要这些能力的用户选择 `dsh web`。没有推理内容的成功运行会保持 stderr 为空,有推理内容的运行则在那里流式输出提供方报告的内容;完成结果在持久化 flush 后推导,持久化会话仍可供后续工具使用。初始用户消息记录 `source.kind: 'user'`,因此不携带 ApiProxy `rpcId`。
|
||||
|
||||
ApiProxy 载体覆盖保留在 ApiProxy 包中。自定义一次性 profile 可以显式包含 Host 或 Web 组合包;随附 profile 与可识别的安装过程所属元组均不含 Web。
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-21-headless-reasoning-progress.md
|
||||
2026-08-21-headless-reasoning-progress.md: 6c4a3574b63ef316fd456f24b473406054ed30f3
|
||||
2026-08-21-headless-reasoning-progress.zh.md: e19fffc33fb5a55432ba2b6cb50550a0cfbd00b8
|
||||
@@ -0,0 +1,37 @@
|
||||
# Agent Note: headless streams provider reasoning to stderr
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-08-21-headless-reasoning-progress.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The one-shot headless runner waits for complete Agent quiescence before printing the final assistant text. Reasoning-capable providers already expose their reasoning as durable `assistant/chunk` events, but a long reasoned response leaves the terminal silent until the run completes. The final answer must remain the only stdout payload so command substitution and other consumers keep a stable result channel.
|
||||
|
||||
The earlier [direct core entry-point decision](../architecture/2026-08-09-headless-direct-core-entry-point.md) required empty stderr on every successful run. That clause prevents live reasoning progress and is superseded by this note; its transport, durability, and completion decisions remain unchanged.
|
||||
|
||||
## Decision
|
||||
|
||||
`headless-runner` observes the exact Session it creates after startup quiescence and before submitting the task. Once the owned interval opens with `turn/start`, each non-empty `assistant/chunk.reasoning-delta` is written immediately to stderr. A contiguous reasoning phase starts with `dsh: reasoning:` on its own line; deltas retain provider order without token-boundary decoration. Reasoning block boundaries and usage metadata keep that phase open; a later non-reasoning block or output delta, stream finish, new turn, or listener disposal terminates it with one newline when the provider supplied none.
|
||||
|
||||
This output is a transient projection of the existing durable Session event stream. The runner still derives final text and exit status from the flushed log rather than from progress-presentation state. The LLM adapter, agent loop, Session event types, persistence format, and SDK projections do not change.
|
||||
|
||||
Reasoning progress is not TTY-gated and has no separate flag. A redirected stderr stream and a supervisor receive the same provider-reported content as an attached terminal. A successful run without reasoning still writes nothing to stderr; terminal model and driver errors keep their existing `dsh:` diagnostics after any open reasoning phase is terminated.
|
||||
|
||||
## Verification
|
||||
|
||||
The package test holds the Agent active after a reasoning delta and observes stderr before idle, then pins newline ownership for provider-terminated and unterminated phases plus terminal errors. The owner-local product expectation drives the shipped headless profile through a reasoning-plus-tool round and pins both stderr and the persisted Session. Recorded-session replay reconstructs expected stderr from scalar and packed chunk rows, closes sections on packed text and tool-call output, and uses the raw run log before fixture path tokenization in record modes. Built-bin acceptance sends `reasoning_content` through the native DeepSeek SSE adapter and requires reasoning on stderr while stdout remains the final answer.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Dump reasoning after quiescence.** Folding reasoning from the persisted log would preserve content but leave the terminal silent during the long-running interval that motivates the feature.
|
||||
|
||||
**Wrap the LLM stream.** Tapping `ctx.llm.stream()` would place a presentation concern in the request path and duplicate the authoritative chunks that the agent loop already appends to the Session.
|
||||
|
||||
**Print a spinner or periodic heartbeat.** A timer reports process liveness rather than provider progress, adds an interval policy, and still hides reasoning that the provider already supplies. Time before the first reasoning delta remains silent and can be addressed separately if providers buffer their first token.
|
||||
|
||||
**Enable output only on a TTY or explicit flag.** Headless runs under CI and supervisors need the same progress signal, while implicit TTY-dependent behavior makes redirected runs differ from interactive runs. Callers that do not want reasoning logs redirect stderr.
|
||||
|
||||
## Consequences
|
||||
|
||||
Reasoning-capable successful runs now write provider-reported content to stderr, so log collectors may retain substantially more and potentially sensitive model output. Stdout remains one final assistant result, text-only success keeps stderr empty, errors remain line-separated, and no new configuration or durable format is introduced. Silence before the provider emits its first non-empty reasoning delta remains an explicit limitation.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Agent Note: headless 将提供方推理流式写入 stderr
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-08-21-headless-reasoning-progress.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
一次性 headless runner 会等待 Agent(智能体)完全停稳,再打印最终 assistant 文本。具备推理能力的提供方已经把推理作为持久化的 `assistant/chunk` 事件暴露,但耗时较长的推理响应会让终端在运行完成前始终保持静默。最终答案必须继续作为 stdout 中唯一的载荷,使命令替换和其他消费方保持稳定的结果通道。
|
||||
|
||||
此前的[直接使用核心服务入口决策](../architecture/2026-08-09-headless-direct-core-entry-point.zh.md)要求每次成功运行都保持 stderr 为空。该条款会阻止实时推理进度,因此由本 Agent Note 取代;其中关于传输、持久性与完成状态的其他决策保持不变。
|
||||
|
||||
## 决策
|
||||
|
||||
`headless-runner` 在启动工作完全停稳后、提交任务前,观察其创建的精确 Session。自身持有的区间以 `turn/start` 打开后,每个非空的 `assistant/chunk.reasoning-delta` 都会立即写入 stderr。一段连续推理以独占一行的 `dsh: reasoning:` 开始;各分片保持提供方顺序,不添加 token 边界装饰。推理块边界与用量元数据会保持该段打开;之后出现非推理块或输出分片、流结束、新轮次或 listener dispose(资源释放)时,如果提供方没有输出末尾换行,runner 会用一个换行终止该段。
|
||||
|
||||
该输出是既有持久化会话事件流的瞬时投影。runner 仍从 flush 后的日志而不是进度呈现状态推导最终文本与退出状态。LLM(大语言模型)适配器、agent loop(智能体循环)、Session 事件类型、持久化格式与 SDK 投影均不改变。
|
||||
|
||||
推理进度不按 TTY 启用,也没有单独 flag。重定向的 stderr 流与监督进程会收到和已连接终端相同的提供方报告内容。没有推理内容的成功运行仍不会写入 stderr;终止态模型错误与驱动器错误继续在任何已打开推理段终止后输出既有的 `dsh:` 诊断。
|
||||
|
||||
## 验证
|
||||
|
||||
包测试在推理分片后保持 Agent 活跃,并在 idle 前观察 stderr;测试同时固定由提供方终止和未终止的推理段换行归属,以及终止态错误。产品自有期望通过包含推理与工具调用的轮次驱动随附 headless profile,并固定 stderr 与持久化 Session。录制会话回放从标量及压缩分片记录重建预期 stderr,在压缩文本或工具调用输出处关闭推理段,并在录制模式下于 fixture 路径标记化之前使用原始运行日志。构建后二进制验收通过原生 DeepSeek SSE(Server-Sent Events)适配器发送 `reasoning_content`,要求推理出现在 stderr,同时 stdout 仍只包含最终答案。
|
||||
|
||||
## 考虑过的替代方案
|
||||
|
||||
**完全停稳后再输出推理。** 从持久化日志折叠推理能够保留内容,但在导致本功能产生的长时间运行区间内,终端仍会保持静默。
|
||||
|
||||
**包装 LLM 流。** 截取 `ctx.llm.stream()` 会把呈现职责放入请求路径,并重复处理 agent loop 已经追加到 Session 的权威分片。
|
||||
|
||||
**打印 spinner 或周期性心跳。** 定时器报告的是进程存活状态,而不是提供方进度;它还会新增间隔策略,并继续隐藏提供方已经给出的推理。首个推理分片前的时间仍保持静默;如果提供方会缓冲首个 token,可以另行处理。
|
||||
|
||||
**仅在 TTY 或显式 flag 下启用输出。** CI 与监督进程中的 headless 运行需要相同的进度信号,而隐式依赖 TTY 会让重定向运行与交互式运行产生差异。不需要推理日志的调用方可以重定向 stderr。
|
||||
|
||||
## 后果
|
||||
|
||||
具备推理能力的成功运行会把提供方报告的内容写入 stderr,因此日志收集器可能保留明显更多且可能敏感的模型输出。stdout 仍只包含一个最终 assistant 结果,没有推理内容的成功运行保持 stderr 为空,错误继续与推理内容分行,并且本决策不引入新配置或持久化格式。提供方发出首个非空推理分片前保持静默,这是明确的限制。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-24-web-per-turn-token-usage.md
|
||||
2026-08-24-web-per-turn-token-usage.md: 91aab0f2c261e2141ee964c828e7209ba2b3f72f
|
||||
2026-08-24-web-per-turn-token-usage.zh.md: f9c424fa0f84b802e98280bea9eaf4038bf31f19
|
||||
@@ -0,0 +1,31 @@
|
||||
# Agent Note: Exact Web per-Turn token usage
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-08-24-web-per-turn-token-usage.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
Web Chat exposes cumulative session token usage near the composer, but that value cannot explain the cost of one completed Turn. A paged history window may begin inside a Turn, retries may consume several model calls, streaming and final events may repeat one attempt's usage, and optional cache fields do not prove an exact total. Displaying a partial subtotal as Turn usage would make recorded provider facts look more complete than they are.
|
||||
|
||||
## Decision
|
||||
|
||||
The shared `TokenUsage` value carries optional `totalTokens` for one model call. Adapters publish it only from an exact provider total or authoritative aggregate prompt and output counters. DeepSeek checks its prompt-plus-completion aggregate against any wire total, and pi-ai preserves its provided total.
|
||||
|
||||
Token-meter owns a browser-safe pure Turn-local fold over durable session events, shared with its retry-aware cumulative usage projection. `step/start` and `llm/retry-started` open actual attempts; a final assistant message replaces the same attempt's streaming sample; terminal failures, retries, and step boundaries close attempts without double counting. Every started attempt must close with safe non-negative integer usage and an exact total. Optional cache, reasoning, and route aggregates appear only when every contributing attempt reports them, and reasoning remains a subset of output.
|
||||
|
||||
Web Chat selects a Turn only when its loaded match window includes `turn/start`, passes that complete durable-event window to the token-meter fold, and renders the result. A complete, exact result appears through a local-state `DisclosureRow` above the existing actions; incomplete or contradictory evidence produces no row. Chat owns no token-accounting state machine.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Subtract neighboring cumulative session values.** Rejected because pagination, compaction, retry coverage, and projection completeness can make adjacent values incomparable; subtraction would infer data that no call reported.
|
||||
|
||||
**Publish historical per-Turn values through a new client session projection.** Rejected because the loaded per-Turn view already has the durable attempt events it needs, while a history-growing projection would add transport, persistence, and versioning costs. Reusing token-meter's pure fold keeps one accounting owner without adding another wire value.
|
||||
|
||||
**Show known buckets without an exact total.** Rejected because a lower-bound subtotal presented in a completed Turn footer is indistinguishable from a complete bill.
|
||||
|
||||
## Consequences
|
||||
|
||||
New provider records can expose exact per-Turn accounting without a new transport or persisted UI state. Older sessions and adapters without enough evidence simply omit the disclosure. Model routes disappear as a group when any billed attempt lacks attribution, while trustworthy token totals remain visible.
|
||||
|
||||
Focused adapter, token-meter fold/projection, component, pagination, and assembled Web replay tests pin total preservation, retry-attempt separation, fail-closed validation, optional-field omission, interaction, and full-window publication. The cumulative projection and exact Turn fold now share token-meter ownership; the projection remains a whole-log bucket view, while the fold alone makes the stricter exactness and completeness claim required by the disclosure.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Agent Note: Web 单轮次精确 token 用量
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-08-24-web-per-turn-token-usage.md) | 中文
|
||||
|
||||
## Problem
|
||||
|
||||
Web Chat 在编辑框附近显示会话累计 token 用量,但该值无法解释一个已完成轮次的消耗。分页历史窗口可能从轮次中间开始,重试可能消耗多次模型调用,流式事件与最终事件可能重复携带同一次 attempt 的用量,而可选 cache 字段也不能证明精确总量。将局部小计显示成轮次用量,会让已记录的提供方事实显得比实际更完整。
|
||||
|
||||
## Decision
|
||||
|
||||
共享 `TokenUsage` 值为一次模型调用携带可选的 `totalTokens`。适配器只从提供方精确总量,或权威的提示词与输出聚合计数发布该字段。DeepSeek 会将提示词加输出的聚合值与协议提供的总量核对,pi-ai 则保留其提供的总量。
|
||||
|
||||
token-meter 拥有一份可安全用于浏览器的纯轮次局部 fold,并与其具备重试感知能力的累计用量投影共享记账所有权。`step/start` 与 `llm/retry-started` 打开真实 attempt;最终 assistant 消息替换同一 attempt 的流式样本;终止失败、重试与步骤边界关闭 attempt,且不会重复计数。每个已开始的 attempt 都必须以安全的非负整数用量和精确总量关闭。只有每个参与聚合的 attempt 都报告时,才会显示可选的 cache、推理与路由聚合值;推理仍是输出的子集。
|
||||
|
||||
Web Chat 只选择已加载匹配窗口包含 `turn/start` 的 Turn,将该完整的持久事件窗口交给 token-meter fold,再渲染结果。完整且精确的结果通过现有 actions 上方、仅保留本地状态的 `DisclosureRow` 显示;证据不完整或矛盾时不显示该行。Chat 不拥有 token 记账状态机。
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**对相邻的会话累计值做减法。** 不采用,因为分页、压缩、重试覆盖范围与投影完整性可能让相邻值无法比较;减法会推断任何调用都未报告的数据。
|
||||
|
||||
**通过新的客户端会话投影发布历史单轮次值。** 不采用,因为已加载的单轮次视图已经拥有所需的持久 attempt 事件,而随历史增长的投影会增加传输、持久化与版本成本。复用 token-meter 的纯 fold,可以在不新增 wire 值的前提下保持唯一记账所有方。
|
||||
|
||||
**缺少精确总量时仍显示已知 bucket。** 不采用,因为在已完成轮次 footer 中展示的下界小计与完整账单无法区分。
|
||||
|
||||
## Consequences
|
||||
|
||||
新的提供方记录无需新增传输接口或持久化 UI 状态,即可显示精确的单轮次记账。证据不足的旧会话与适配器只会省略 disclosure。任一计费 attempt 缺少归属时,模型路由会整体消失,可信 token 总量仍可显示。
|
||||
|
||||
定向的适配器、token-meter fold/投影、组件、分页与组装 Web 回放测试固定了总量保留、重试 attempt 分离、fail-closed 校验、可选字段省略、交互与完整窗口发布。累计投影与精确 Turn fold 现在同归 token-meter 所有;投影仍是完整日志的 bucket 视图,只有 fold 会作出 disclosure 所需的更严格精确性与完整性声明。
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write apps/cli/reference/README.md
|
||||
README.md: 6aec47ab2b3b7650bb86886c0201238daeaa4502
|
||||
README.zh.md: 59bd4617de77ab2809bc6d9cdf554d92716ae016
|
||||
README.md: fe8d6ef0bb296f0807de4a3ec2756016bbb510c2
|
||||
README.zh.md: e8c353f33bc9760fd6da74af33a85111cf9012aa
|
||||
|
||||
@@ -30,7 +30,7 @@ The shipped apps own these command lines:
|
||||
| `sdk-minimal` | no options; stdio carries the same JSON-RPC protocol |
|
||||
| `acp` | no options; stdio carries Agent Client Protocol |
|
||||
|
||||
A one-shot task (`dsh --profile headless "run the tests"`) creates one fresh persisted Agent through the core registry, submits the task, waits for quiescence, and flushes the Session before deriving the last non-empty assistant text and final `turn/end` reason from its durable interval. It prints the text on stdout and exits 0 for `completed`, else 1. An invocation with no task is a usage error from that app. The shipped headless profile mounts no ApiProxy, Host, HTTP server, Web runtime, or browser client; a successful run writes nothing to stderr and opens no listening port.
|
||||
A one-shot task (`dsh --profile headless "run the tests"`) creates one fresh persisted Agent through the core registry, submits the task, waits for quiescence, and flushes the Session before deriving the last non-empty assistant text and final `turn/end` reason from its durable interval. It streams non-empty provider reasoning deltas to stderr under a `dsh: reasoning:` heading, prints only the final text on stdout, and exits 0 for `completed`, else 1; a successful response with no reasoning leaves stderr empty. An invocation with no task is a usage error from that app. The shipped headless profile mounts no ApiProxy, Host, HTTP server, Web runtime, or browser client, and opens no listening port.
|
||||
|
||||
Inspect the composed tree without booting it:
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@
|
||||
| `sdk-minimal` | 无选项;stdio 携带相同的 JSON-RPC 协议 |
|
||||
| `acp` | 无选项;stdio 携带 Agent Client Protocol |
|
||||
|
||||
一次性任务(`dsh --profile headless "run the tests"`)通过核心注册表创建一个全新的持久化 Agent(智能体),提交任务、等待完全停稳并对会话执行 flush,再从其持久化事件区间中推导最后一个非空 assistant 文本与最终 `turn/end` 原因。它在 stdout 打印文本,并在原因为 `completed` 时以 0 退出,否则以 1 退出。没有任务的调用是该应用的用法错误。随附 headless profile 不挂载 ApiProxy、Host、HTTP 服务器、Web 运行时或浏览器客户端;成功运行不会向 stderr 写入任何内容,也不会打开监听端口。
|
||||
一次性任务(`dsh --profile headless "run the tests"`)通过核心注册表创建一个全新的持久化 Agent(智能体),提交任务、等待完全停稳并对会话执行 flush,再从其持久化事件区间中推导最后一个非空 assistant 文本与最终 `turn/end` 原因。它在 `dsh: reasoning:` 标题下将非空的提供方推理分片流式写入 stderr,只在 stdout 打印最终文本,并在原因为 `completed` 时以 0 退出,否则以 1 退出;没有推理内容的成功响应会保持 stderr 为空。没有任务的调用是该应用的用法错误。随附 headless profile 不挂载 ApiProxy、Host、HTTP 服务器、Web 运行时或浏览器客户端,也不会打开监听端口。
|
||||
|
||||
可在不启动的情况下检查组合出的配置树:
|
||||
|
||||
|
||||
@@ -560,8 +560,9 @@ describe.skipIf(!existsSync(dshBin))('dsh BUILT bin (node lib/bin.js, no tsx)',
|
||||
it('runs the headless profile through its app-owned task positional', async () => {
|
||||
const apiKey = 'built-dsh-headless-key'
|
||||
const server = await startMockLlmServer({
|
||||
sequence: ['success'],
|
||||
sequence: ['reasoning_success'],
|
||||
apiKey,
|
||||
reasoningText: 'Inspecting the published entry.',
|
||||
successText: 'published headless profile reached the mock',
|
||||
})
|
||||
const home = mkdtempSync(join(tmpdir(), 'dsh-built-headless-'))
|
||||
@@ -574,7 +575,7 @@ describe.skipIf(!existsSync(dshBin))('dsh BUILT bin (node lib/bin.js, no tsx)',
|
||||
})
|
||||
expect(result.code, result.stderr).toBe(0)
|
||||
expect(result.stdout).toBe('published headless profile reached the mock')
|
||||
expect(result.stderr).toBe('')
|
||||
expect(result.stderr).toBe('dsh: reasoning:\nInspecting the published entry.')
|
||||
expect(server.requests.length).toBeGreaterThan(0)
|
||||
expect(server.requests.every(request => request.path === '/chat/completions')).toBe(true)
|
||||
expect(JSON.stringify(server.requests.map(request => request.body))).toContain('answer from the published entry')
|
||||
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
dsh: reasoning:
|
||||
Inspecting the task before the tool call.
|
||||
+9
-6
@@ -12,14 +12,17 @@
|
||||
{"type":"request/header","data":{"header":{"config":{"provider":"cli-mock","model":"cli-mock","reasoningEffort":"high"},"adapterDefaults":{"reasoningEffort":true},"system":"{{system}}","tools":"{{tools}}"},"reason":"initial"}}
|
||||
{"type":"request/context","data":{"provider":"cli-mock","model":"cli-mock"}}
|
||||
{"type":"session/title-llm-request","data":{"titleProvider":"session-title-first-prompt-llm","messageSeqs":[7],"route":{"provider":"cli-mock","model":"cli-mock"},"system":"Create a concise title for an AI coding-assistant session from the supplied human messages.\nReturn only the title on one line, **in plain text of natural language**, with no quotes, prefix, explanation, Markdown, XML, or terminal control codes. No code is allowed.\nUse the language of the messages.\nAim for about 5 words in non-CJK languages or 10 CJK characters.","messages":[{"content":[{"type":"text","text":"Generate the session title from this JSON array of human messages:\n[{\"seq\":7,\"text\":\"Prove the product headless profile path with one real tool round trip.\"}]"}],"source":{"kind":"plugin","plugin":"dsh-session-title-llm"},"role":"user","id":"{{sessionId}}"}],"maxTokens":64}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":0,"id":"cli-smoke-call","name":"bash","argumentsDelta":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"cli-smoke-call","name":"bash","arguments":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"reasoning"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":"Inspecting the task before the tool call."}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"reasoning","text":"Inspecting the task before the tool call."}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":1,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"cli-smoke-call","name":"bash","argumentsDelta":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":1,"block":{"type":"tool-call","id":"cli-smoke-call","name":"bash","arguments":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":11,"outputTokens":3,"cacheReadTokens":2}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"tool-call","id":"cli-smoke-call","name":"bash","arguments":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}],"source":{"kind":"model","provider":"cli-mock","model":"cli-mock"},"id":"{{sessionId}}"},"usage":{"inputTokens":11,"outputTokens":3,"cacheReadTokens":2}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"reasoning","text":"Inspecting the task before the tool call."},{"type":"tool-call","id":"cli-smoke-call","name":"bash","arguments":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}],"source":{"kind":"model","provider":"cli-mock","model":"cli-mock"},"id":"{{sessionId}}"},"usage":{"inputTokens":11,"outputTokens":3,"cacheReadTokens":2}},"sourceEventSeqs":[13,14,15,16,17,18,19,20],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":1,"callId":"cli-smoke-call","name":"bash","arguments":"{\"command\":\"printf CLI_TOOL_ROUND_TRIP\",\"description\":\"Prove the CLI tool round trip.\"}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"cli-smoke-call"},"content":[{"type":"tool-result","toolCallId":"cli-smoke-call","content":[{"type":"text","text":"CLI_TOOL_ROUND_TRIP"}],"isError":false}],"role":"user","id":"{{sessionId}}"}},"sourceEventSeqs":[19],"surfaceOp":"append"}
|
||||
{"type":"tool/result","data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"cli-smoke-call"},"content":[{"type":"tool-result","toolCallId":"cli-smoke-call","content":[{"type":"text","text":"CLI_TOOL_ROUND_TRIP"}],"isError":false}],"role":"user","id":"{{sessionId}}"}},"sourceEventSeqs":[22],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
{"type":"step/start","data":{"turn":1,"step":2}}
|
||||
{"type":"request/header","data":{"header":{"config":{"provider":"cli-mock","model":"cli-mock","reasoningEffort":"off"},"system":"{{system}}","tools":"{{tools}}"},"reason":"change"}}
|
||||
@@ -28,6 +31,6 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"CLI tool round trip complete: CLI_TOOL_ROUND_TRIP"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":7,"outputTokens":5,"reasoningTokens":1}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"text","text":"CLI tool round trip complete: CLI_TOOL_ROUND_TRIP"}],"source":{"kind":"model","provider":"cli-mock","model":"cli-mock"},"id":"{{sessionId}}"},"usage":{"inputTokens":7,"outputTokens":5,"reasoningTokens":1}},"sourceEventSeqs":[24,25,26,27,28],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"text","text":"CLI tool round trip complete: CLI_TOOL_ROUND_TRIP"}],"source":{"kind":"model","provider":"cli-mock","model":"cli-mock"},"id":"{{sessionId}}"},"usage":{"inputTokens":7,"outputTokens":5,"reasoningTokens":1}},"sourceEventSeqs":[27,28,29,30,31],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":2}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -41,6 +41,7 @@ const deepseekDefaultsConfigPath = fileURLToPath(new URL('./fixtures/deepseek-de
|
||||
const piAiDefaultsConfigPath = fileURLToPath(new URL('./fixtures/pi-ai-defaults.cordis.yml', import.meta.url))
|
||||
const headlessOverlayPath = fileURLToPath(new URL('./fixtures/headless-profile.cordis.yml', import.meta.url))
|
||||
const headlessSessionExpected = join(goldensDir, 'headless-profile', 'session.expected.jsonl')
|
||||
const headlessReasoningExpected = join(goldensDir, 'headless-profile', 'reasoning.stderr.expected.txt')
|
||||
const headlessFailureExpected = join(goldensDir, 'headless-profile', 'stderr.expected.txt')
|
||||
const refreshing = process.env.DSH_SNAPSHOT === 'refresh'
|
||||
|
||||
@@ -223,7 +224,8 @@ describe('headless stream-json snapshots', () => {
|
||||
})
|
||||
|
||||
expect(result.stdout).toBe('CLI tool round trip complete: CLI_TOOL_ROUND_TRIP\n')
|
||||
expect(result.stderr).toBe('')
|
||||
if (refreshing) await writeFile(headlessReasoningExpected, result.stderr)
|
||||
expect(result.stderr).toBe(await readFile(headlessReasoningExpected, 'utf8'))
|
||||
}, LOADER_SMOKE_TEST_TIMEOUT_MS)
|
||||
|
||||
it('prints a terminal model failure through the product headless profile command', async () => {
|
||||
|
||||
@@ -27,6 +27,7 @@ const FIXTURE = join(SNAPSHOT_DIR, 'session.jsonl')
|
||||
// Two goldens for the same message: parked mid-turn, then settled.
|
||||
const RUNNING_EXPECTED = join(SNAPSHOT_DIR, 'running.expected.md')
|
||||
const SETTLED_EXPECTED = join(SNAPSHOT_DIR, 'settled.expected.md')
|
||||
const USAGE_EXPANDED_EXPECTED = join(SNAPSHOT_DIR, 'usage-expanded.expected.md')
|
||||
const MODE = webSnapshotMode()
|
||||
|
||||
// The recording must carry text in the SAME assistant message as the tool
|
||||
@@ -156,7 +157,34 @@ describe('web e2e: assistant IconActions wait for the turn to end', () => {
|
||||
expect(tripwire.warnings).toEqual([])
|
||||
}, 120_000)
|
||||
|
||||
it.skipIf(MODE === 'record')('shows exact completed-Turn usage and expands its available facts', async () => {
|
||||
await launch()
|
||||
onTestFailed(() => saveFailureShot(page, 'web-e2e-turn-usage-expanded'))
|
||||
const { settled } = await sendPrompt(120_000)
|
||||
await settled
|
||||
|
||||
const disclosure = page.getByRole('button', { name: /Turn usage/ })
|
||||
await expect.poll(() => disclosure.count(), { timeout: 10_000 }).toBe(1)
|
||||
expect(await disclosure.getAttribute('aria-expanded')).toBe('false')
|
||||
expect(await page.getByText('15.8K tok · Cache hit 49.7%', { exact: true }).count()).toBe(1)
|
||||
|
||||
await disclosure.click()
|
||||
expect(await disclosure.getAttribute('aria-expanded')).toBe('true')
|
||||
expect(await page.getByText('deepseek-official/deepseek-v4-flash', { exact: true }).count()).toBe(1)
|
||||
expect(await page.getByText('7,891 tok', { exact: true }).count()).toBe(1)
|
||||
expect(await page.getByText('7,808 tok', { exact: true }).count()).toBe(1)
|
||||
expect(await page.getByText('112 tok (42 tok reasoning)', { exact: true }).count()).toBe(1)
|
||||
expect(await page.getByText('15,811 tok', { exact: true }).count()).toBe(1)
|
||||
|
||||
const expanded = await captureStableAria(page, '[class*="centerCol"]', scaffold!.workspaceCwd)
|
||||
await compareOrRefreshGolden(USAGE_EXPANDED_EXPECTED, expanded, MODE)
|
||||
expect(tripwire.pageErrors).toEqual([])
|
||||
expect(tripwire.warnings).toEqual([])
|
||||
}, 120_000)
|
||||
|
||||
it.skipIf(MODE === 'record')('keeps a closed fixture inventory', async () => {
|
||||
await assertFixtureInventory(SNAPSHOT_DIR, ['running.expected.md', 'session.jsonl', 'settled.expected.md'])
|
||||
await assertFixtureInventory(SNAPSHOT_DIR, [
|
||||
'running.expected.md', 'session.jsonl', 'settled.expected.md', 'usage-expanded.expected.md',
|
||||
])
|
||||
})
|
||||
})
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/config-catalog.md
|
||||
config-catalog.md: ef735e98ad7e61701b4a5cf5945ed200d3a1fa5e
|
||||
config-catalog.zh.md: 25447462df05de9087e082a728a95f77b749279f
|
||||
config-catalog.md: 70a6afd44a037e8cb797a7450274c387f9e79880
|
||||
config-catalog.zh.md: 442397a27c4edfd95b86e93f21813edd9236152f
|
||||
|
||||
@@ -703,7 +703,7 @@ export interface Config {
|
||||
}
|
||||
```
|
||||
|
||||
Source: [`packages/bundle/headless/src/index.ts:31`](../packages/bundle/headless/src/index.ts)
|
||||
Source: [`packages/bundle/headless/src/index.ts:32`](../packages/bundle/headless/src/index.ts)
|
||||
|
||||
<a id="deepseek-aidsh-hooks-claude-code"></a>
|
||||
|
||||
|
||||
@@ -705,7 +705,7 @@ export interface Config {
|
||||
}
|
||||
```
|
||||
|
||||
来源:[`packages/bundle/headless/src/index.ts:31`](../packages/bundle/headless/src/index.ts)
|
||||
来源:[`packages/bundle/headless/src/index.ts:32`](../packages/bundle/headless/src/index.ts)
|
||||
|
||||
<a id="deepseek-aidsh-hooks-claude-code"></a>
|
||||
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/event-producer-consumer.md
|
||||
event-producer-consumer.md: de2a94abb5e4d16433eae71e34e329fcf0042ede
|
||||
event-producer-consumer.zh.md: 7a9e825750213b2d0a67d9c022bffe031194c8ba
|
||||
event-producer-consumer.md: baeddb7b0171b4347fa1748e0adf87df5c15e4d0
|
||||
event-producer-consumer.zh.md: baee416d8c0bf1b4102f839cdcd98654d0acd63e
|
||||
|
||||
@@ -47,7 +47,7 @@ This matrix shows which packages dispatch each harness-owned event and which pac
|
||||
| `session-telemetry/record` | `waterfall` | [`packages/session/session-telemetry/src/index.ts:43`](../packages/session/session-telemetry/src/index.ts) | [`session-telemetry`](../packages/session/session-telemetry) (`waterfall`) | - |
|
||||
| `session/created` | `emit` | [`packages/core/session/src/index.ts:54`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`compaction`](../packages/compaction/compaction), [`goal`](../packages/goal/goal), [`hook-protocol`](../packages/hooks/hook-protocol), [`llm-retry`](../packages/llm/llm-retry), [`permission-presets`](../packages/interaction/permission-presets), [`plan-mode`](../packages/plan/plan-mode), [`schedule`](../packages/schedule/schedule), `server`, [`session`](../packages/core/session), `session-controller`, [`session-log-deepseek`](../packages/session/session-log-deepseek), [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-telemetry`](../packages/session/session-telemetry), [`time-context`](../packages/context/time-context), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/disposed` | `emit` | [`packages/core/session/src/index.ts:64`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`agent-loop`](../packages/core/agent-loop), `agent-team`, `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-title`](../packages/session/session-title) |
|
||||
| `session/event` | `emit` | [`packages/core/session/src/index.ts:76`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`acp`](../packages/acp/acp), [`agent-instructions`](../packages/context/agent-instructions), [`agent-loop`](../packages/core/agent-loop), [`agent-presets`](../packages/preset/agent-presets), `agent-team`, [`compaction`](../packages/compaction/compaction), [`compaction-basic`](../packages/compaction/compaction-basic), [`file-reference-local`](../packages/context/file-reference-local), [`goal`](../packages/goal/goal), [`goal-round-driver`](../packages/goal/goal-round-driver), [`hook-protocol`](../packages/hooks/hook-protocol), [`loader-smoke`](../packages/test-support/loader-smoke), `server`, [`session`](../packages/core/session), `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-telemetry-otel`](../packages/session/session-telemetry-otel), [`session-title`](../packages/session/session-title), [`token-meter`](../packages/llm/token-meter), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/event` | `emit` | [`packages/core/session/src/index.ts:76`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`acp`](../packages/acp/acp), [`agent-instructions`](../packages/context/agent-instructions), [`agent-loop`](../packages/core/agent-loop), [`agent-presets`](../packages/preset/agent-presets), `agent-team`, [`compaction`](../packages/compaction/compaction), [`compaction-basic`](../packages/compaction/compaction-basic), [`file-reference-local`](../packages/context/file-reference-local), [`goal`](../packages/goal/goal), [`goal-round-driver`](../packages/goal/goal-round-driver), [`headless`](../packages/bundle/headless), [`hook-protocol`](../packages/hooks/hook-protocol), [`loader-smoke`](../packages/test-support/loader-smoke), `server`, [`session`](../packages/core/session), `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-telemetry-otel`](../packages/session/session-telemetry-otel), [`session-title`](../packages/session/session-title), [`token-meter`](../packages/llm/token-meter), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/flush` | `parallel` | [`packages/core/session/src/index.ts:85`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`session-persistence`](../packages/session/session-persistence), [`session-telemetry`](../packages/session/session-telemetry) |
|
||||
| `settings/document-updated` | `emit` | [`packages/settings/settings/src/types.ts:48`](../packages/settings/settings/src/types.ts) | [`settings`](../packages/settings/settings) (`events.dispatch`) | `remotes` |
|
||||
| `settings/updated` | `emit` | [`packages/settings/settings/src/types.ts:35`](../packages/settings/settings/src/types.ts) | [`settings`](../packages/settings/settings) (`events.dispatch`) | [`settings`](../packages/settings/settings) |
|
||||
|
||||
@@ -49,7 +49,7 @@
|
||||
| `session-telemetry/record` | `waterfall` | [`packages/session/session-telemetry/src/index.ts:43`](../packages/session/session-telemetry/src/index.ts) | [`session-telemetry`](../packages/session/session-telemetry) (`waterfall`) | - |
|
||||
| `session/created` | `emit` | [`packages/core/session/src/index.ts:54`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`compaction`](../packages/compaction/compaction), [`goal`](../packages/goal/goal), [`hook-protocol`](../packages/hooks/hook-protocol), [`llm-retry`](../packages/llm/llm-retry), [`permission-presets`](../packages/interaction/permission-presets), [`plan-mode`](../packages/plan/plan-mode), [`schedule`](../packages/schedule/schedule), `server`, [`session`](../packages/core/session), `session-controller`, [`session-log-deepseek`](../packages/session/session-log-deepseek), [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-telemetry`](../packages/session/session-telemetry), [`time-context`](../packages/context/time-context), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/disposed` | `emit` | [`packages/core/session/src/index.ts:64`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`agent-loop`](../packages/core/agent-loop), `agent-team`, `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-title`](../packages/session/session-title) |
|
||||
| `session/event` | `emit` | [`packages/core/session/src/index.ts:76`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`acp`](../packages/acp/acp), [`agent-instructions`](../packages/context/agent-instructions), [`agent-loop`](../packages/core/agent-loop), [`agent-presets`](../packages/preset/agent-presets), `agent-team`, [`compaction`](../packages/compaction/compaction), [`compaction-basic`](../packages/compaction/compaction-basic), [`file-reference-local`](../packages/context/file-reference-local), [`goal`](../packages/goal/goal), [`goal-round-driver`](../packages/goal/goal-round-driver), [`hook-protocol`](../packages/hooks/hook-protocol), [`loader-smoke`](../packages/test-support/loader-smoke), `server`, [`session`](../packages/core/session), `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-telemetry-otel`](../packages/session/session-telemetry-otel), [`session-title`](../packages/session/session-title), [`token-meter`](../packages/llm/token-meter), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/event` | `emit` | [`packages/core/session/src/index.ts:76`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`acp`](../packages/acp/acp), [`agent-instructions`](../packages/context/agent-instructions), [`agent-loop`](../packages/core/agent-loop), [`agent-presets`](../packages/preset/agent-presets), `agent-team`, [`compaction`](../packages/compaction/compaction), [`compaction-basic`](../packages/compaction/compaction-basic), [`file-reference-local`](../packages/context/file-reference-local), [`goal`](../packages/goal/goal), [`goal-round-driver`](../packages/goal/goal-round-driver), [`headless`](../packages/bundle/headless), [`hook-protocol`](../packages/hooks/hook-protocol), [`loader-smoke`](../packages/test-support/loader-smoke), `server`, [`session`](../packages/core/session), `session-controller`, [`session-persistence`](../packages/session/session-persistence), [`session-projection`](../packages/session/session-projection), [`session-projection-cache`](../packages/session/session-projection-cache), [`session-telemetry`](../packages/session/session-telemetry), [`session-telemetry-otel`](../packages/session/session-telemetry-otel), [`session-title`](../packages/session/session-title), [`token-meter`](../packages/llm/token-meter), [`tool-todo`](../packages/todo/tool-todo), [`tool-workflow`](../packages/workflow/tool-workflow), [`tools`](../packages/core/tools), [`user-approval`](../packages/interaction/user-approval) |
|
||||
| `session/flush` | `parallel` | [`packages/core/session/src/index.ts:85`](../packages/core/session/src/index.ts) | [`session`](../packages/core/session) (`events.dispatch`) | [`session-persistence`](../packages/session/session-persistence), [`session-telemetry`](../packages/session/session-telemetry) |
|
||||
| `settings/document-updated` | `emit` | [`packages/settings/settings/src/types.ts:48`](../packages/settings/settings/src/types.ts) | [`settings`](../packages/settings/settings) (`events.dispatch`) | `remotes` |
|
||||
| `settings/updated` | `emit` | [`packages/settings/settings/src/types.ts:35`](../packages/settings/settings/src/types.ts) | [`settings`](../packages/settings/settings) (`events.dispatch`) | [`settings`](../packages/settings/settings) |
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/module-graph.md
|
||||
module-graph.md: cc8eaf8b49dc95568d34a4d57e9cbff8d30656c1
|
||||
module-graph.zh.md: 73ba9081455a265194aae943fb96efc0ec95d38f
|
||||
module-graph.md: b407080d634c0e70a00f494c686f55f85998046e
|
||||
module-graph.zh.md: 542ea6be5a1f1f41c5b39e4e83b2c49c97439332
|
||||
|
||||
@@ -741,6 +741,7 @@ flowchart TD
|
||||
pkg_token_meter --> pkg_compaction
|
||||
pkg_token_meter --> pkg_invariants
|
||||
pkg_token_meter --> pkg_llm
|
||||
pkg_token_meter --> pkg_llm_retry
|
||||
pkg_token_meter --> pkg_session
|
||||
pkg_token_meter --> pkg_session_projection
|
||||
pkg_agent_loop --> pkg_agent
|
||||
@@ -1778,7 +1779,7 @@ flowchart TD
|
||||
| [`bash-local`](../packages/shell/bash-local) | `shell` | [`invariants`](../packages/runtime-diagnostics/invariants), [`settings`](../packages/settings/settings), [`shell`](../packages/shell/shell), [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
|
||||
| [`pwsh-local`](../packages/shell/pwsh-local) | `shell` | [`invariants`](../packages/runtime-diagnostics/invariants), [`settings`](../packages/settings/settings), [`shell`](../packages/shell/shell), [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
|
||||
| [`terminal-bash`](../packages/terminal/terminal-bash) | `terminal` | [`agent`](../packages/core/agent), [`invariants`](../packages/runtime-diagnostics/invariants), [`sandbox`](../packages/sandbox/sandbox), [`sandbox-policy`](../packages/sandbox/sandbox-policy), [`session`](../packages/core/session), [`subprocess`](../packages/subprocess/subprocess), [`terminal`](../packages/terminal/terminal) |
|
||||
| [`token-meter`](../packages/llm/token-meter) | `llm` | [`compaction`](../packages/compaction/compaction), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`session`](../packages/core/session), [`session-projection`](../packages/session/session-projection) |
|
||||
| [`token-meter`](../packages/llm/token-meter) | `llm` | [`compaction`](../packages/compaction/compaction), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`llm-retry`](../packages/llm/llm-retry), [`session`](../packages/core/session), [`session-projection`](../packages/session/session-projection) |
|
||||
| [`agent-loop`](../packages/core/agent-loop) | `core` | [`agent`](../packages/core/agent), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`scope`](../packages/core/scope), [`session`](../packages/core/session), [`session-persistence`](../packages/session/session-persistence), [`settings`](../packages/settings/settings), [`system-prompt`](../packages/core/system-prompt), [`tools`](../packages/core/tools) |
|
||||
| [`agent-tool-presentation`](../packages/core/agent-tool-presentation) | `core` | [`invariants`](../packages/runtime-diagnostics/invariants), [`tools`](../packages/core/tools) |
|
||||
| [`tool-goal`](../packages/goal/tool-goal) | `goal` | [`agent`](../packages/core/agent), [`goal`](../packages/goal/goal), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`session`](../packages/core/session), [`system-prompt`](../packages/core/system-prompt), [`tools`](../packages/core/tools) |
|
||||
|
||||
@@ -743,6 +743,7 @@ flowchart TD
|
||||
pkg_token_meter --> pkg_compaction
|
||||
pkg_token_meter --> pkg_invariants
|
||||
pkg_token_meter --> pkg_llm
|
||||
pkg_token_meter --> pkg_llm_retry
|
||||
pkg_token_meter --> pkg_session
|
||||
pkg_token_meter --> pkg_session_projection
|
||||
pkg_agent_loop --> pkg_agent
|
||||
@@ -1780,7 +1781,7 @@ flowchart TD
|
||||
| [`bash-local`](../packages/shell/bash-local) | `shell` | [`invariants`](../packages/runtime-diagnostics/invariants), [`settings`](../packages/settings/settings), [`shell`](../packages/shell/shell), [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
|
||||
| [`pwsh-local`](../packages/shell/pwsh-local) | `shell` | [`invariants`](../packages/runtime-diagnostics/invariants), [`settings`](../packages/settings/settings), [`shell`](../packages/shell/shell), [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
|
||||
| [`terminal-bash`](../packages/terminal/terminal-bash) | `terminal` | [`agent`](../packages/core/agent), [`invariants`](../packages/runtime-diagnostics/invariants), [`sandbox`](../packages/sandbox/sandbox), [`sandbox-policy`](../packages/sandbox/sandbox-policy), [`session`](../packages/core/session), [`subprocess`](../packages/subprocess/subprocess), [`terminal`](../packages/terminal/terminal) |
|
||||
| [`token-meter`](../packages/llm/token-meter) | `llm` | [`compaction`](../packages/compaction/compaction), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`session`](../packages/core/session), [`session-projection`](../packages/session/session-projection) |
|
||||
| [`token-meter`](../packages/llm/token-meter) | `llm` | [`compaction`](../packages/compaction/compaction), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`llm-retry`](../packages/llm/llm-retry), [`session`](../packages/core/session), [`session-projection`](../packages/session/session-projection) |
|
||||
| [`agent-loop`](../packages/core/agent-loop) | `core` | [`agent`](../packages/core/agent), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`scope`](../packages/core/scope), [`session`](../packages/core/session), [`session-persistence`](../packages/session/session-persistence), [`settings`](../packages/settings/settings), [`system-prompt`](../packages/core/system-prompt), [`tools`](../packages/core/tools) |
|
||||
| [`agent-tool-presentation`](../packages/core/agent-tool-presentation) | `core` | [`invariants`](../packages/runtime-diagnostics/invariants), [`tools`](../packages/core/tools) |
|
||||
| [`tool-goal`](../packages/goal/tool-goal) | `goal` | [`agent`](../packages/core/agent), [`goal`](../packages/goal/goal), [`invariants`](../packages/runtime-diagnostics/invariants), [`llm`](../packages/llm/llm), [`session`](../packages/core/session), [`system-prompt`](../packages/core/system-prompt), [`tools`](../packages/core/tools) |
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/subsystems/llm-streaming.md
|
||||
llm-streaming.md: bdc830a5d387cde6967575551ec9b0a9b2626f46
|
||||
llm-streaming.zh.md: b602336bc06cd88a2634f5259eff117da3dcd986
|
||||
llm-streaming.md: 29efabd2b01659bdf2cc798ceadb4bb495e1731e
|
||||
llm-streaming.zh.md: 21ad56e526b9a507644b436b41ad063c5310b2ce
|
||||
|
||||
@@ -278,7 +278,7 @@ interface AppIdentity {
|
||||
|
||||
## `TokenUsage`
|
||||
|
||||
Per-call token accounting. Counts are **disjoint**: `inputTokens` is uncached input only; cached input is reported separately, and billed input is the sum of the three. Adapters whose providers fold cache hits into a single prompt total (DeepSeek's `prompt_tokens`) subtract them back out. `reasoningTokens`, when present, is informational detail already included in `outputTokens`; totals must not add it again.
|
||||
Per-call token accounting. Counts are **disjoint**: `inputTokens` is uncached input only; cached input is reported separately, and billed input is the sum of the three. Adapters whose providers fold cache hits into a single prompt total (DeepSeek's `prompt_tokens`) subtract them back out. Optional `totalTokens` is an exact aggregate prompt-plus-output count preserved from the provider or reconstructed from authoritative aggregate counters; adapters omit it when unavailable or inconsistent. `reasoningTokens`, when present, is informational detail already included in `outputTokens`; totals must not add it again.
|
||||
|
||||
```ts type-equiv
|
||||
/**
|
||||
@@ -292,6 +292,14 @@ Per-call token accounting. Counts are **disjoint**: `inputTokens` is uncached in
|
||||
interface TokenUsage {
|
||||
inputTokens: number
|
||||
outputTokens: number
|
||||
/**
|
||||
* Exact full-call total including aggregate prompt and output tokens.
|
||||
*
|
||||
* Adapters preserve a provider total or derive it from authoritative
|
||||
* aggregate prompt/output counters; they omit it when unavailable or
|
||||
* inconsistent.
|
||||
*/
|
||||
totalTokens?: number
|
||||
cacheReadTokens?: number
|
||||
cacheWriteTokens?: number
|
||||
reasoningTokens?: number
|
||||
|
||||
@@ -282,7 +282,7 @@ interface AppIdentity {
|
||||
|
||||
## `TokenUsage`
|
||||
|
||||
逐调用 token 记账。各计数**互不重叠**:`inputTokens` 只包含未缓存输入;缓存输入单独报告,计费输入是三者之和。若提供方把缓存命中折入单一提示词总数(如 DeepSeek 的 `prompt_tokens`),适配器会再将其扣除。`reasoningTokens` 存在时只是信息性细节,已经包含在 `outputTokens` 中;汇总时不得重复相加。
|
||||
逐调用 token 记账。各计数**互不重叠**:`inputTokens` 只包含未缓存输入;缓存输入单独报告,计费输入是三者之和。若提供方把缓存命中折入单一提示词总数(如 DeepSeek 的 `prompt_tokens`),适配器会再将其扣除。可选的 `totalTokens` 是精确的提示词与输出聚合计数,由适配器保留提供方原值或从权威聚合计数重建;不可用或不一致时省略。`reasoningTokens` 存在时只是信息性细节,已经包含在 `outputTokens` 中;汇总时不得重复相加。
|
||||
|
||||
```ts type-equiv
|
||||
/**
|
||||
@@ -296,6 +296,14 @@ interface AppIdentity {
|
||||
interface TokenUsage {
|
||||
inputTokens: number
|
||||
outputTokens: number
|
||||
/**
|
||||
* Exact full-call total including aggregate prompt and output tokens.
|
||||
*
|
||||
* Adapters preserve a provider total or derive it from authoritative
|
||||
* aggregate prompt/output counters; they omit it when unavailable or
|
||||
* inconsistent.
|
||||
*/
|
||||
totalTokens?: number
|
||||
cacheReadTokens?: number
|
||||
cacheWriteTokens?: number
|
||||
reasoningTokens?: number
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/bundle/headless/README.md
|
||||
README.md: 22b4ac8ecbbaabc1d5268230ea99a5d3a89aff14
|
||||
README.zh.md: a57e29dc947c0c368165af0ad4342a748711500b
|
||||
README.md: 373e50c515ef45d09c32e7dd1b011f149d921250
|
||||
README.zh.md: 443945ced7e5e41899e91e8c169918f4e186de67
|
||||
|
||||
@@ -4,7 +4,9 @@ English | [中文](README.zh.md)
|
||||
|
||||
The dsh one-shot bundle. [`cordis.patch.yml`](cordis.patch.yml) rides directly over [`dsh-base`](../base/README.md): it inherits the base's disabled module-HMR policy, supplies the coding persona and tool mode, mounts Code Mode's worker as a core execution capability, and inserts this package's `headless-runner` plugin (config `{task}`, resolved from the injected `headlessStartup` provider). It mounts no Host, HTTP server, Web runtime, or browser plugin.
|
||||
|
||||
After the Loader settles, the runner reads the shared [`ctx.agentDefaultModel`](../../core/agent-default-model/README.md), creates one fresh persisted Agent through `ctx.agents`, submits the task as an ordinary user message, and waits for quiescence. It flushes the Session before folding the owned durable event interval, writes the last non-empty assistant text to stdout, and requests exit through the launcher-provided `ctx.appExit` host hook ([`dsh-cmdline`](../../boot/cmdline/README.md)) (final `turn/end` completed → 0, otherwise 1). A terminal `error` reason also writes its code and message to stderr; successful runs keep stderr empty. The process opens no listening port. The task text is this app's command line: the ordinary `headless-startup` provider ([`src/startup.ts`](src/startup.ts)) injects `ctx.cmdlineArgs` ([`dsh-cmdline`](../../boot/cmdline/README.md)), reads the positional argument of `dsh --profile headless "task"`, prints the app's `--help`, and provides `headlessStartup`; the runner injects that service and reads its task from lazy config. A missing or whitespace-only task is rejected before the runner activates.
|
||||
After the Loader settles, the runner reads the shared [`ctx.agentDefaultModel`](../../core/agent-default-model/README.md), creates one fresh persisted Agent through `ctx.agents`, submits the task as an ordinary user message, and waits for quiescence. Each non-empty provider reasoning delta from that Agent is written to stderr as it arrives under a `dsh: reasoning:` heading; consecutive deltas remain one section, and the runner terminates the section before later output when the provider supplied no trailing newline. It then flushes the Session before folding the owned durable event interval, writes the last non-empty assistant text to stdout, and requests exit through the launcher-provided `ctx.appExit` host hook ([`dsh-cmdline`](../../boot/cmdline/README.md)) (final `turn/end` completed → 0, otherwise 1). A terminal `error` reason also writes its code and message to stderr; a successful run with no reasoning keeps stderr empty. The process opens no listening port.
|
||||
|
||||
The task text is this app's command line: the ordinary `headless-startup` provider ([`src/startup.ts`](src/startup.ts)) injects `ctx.cmdlineArgs` ([`dsh-cmdline`](../../boot/cmdline/README.md)), reads the positional argument of `dsh --profile headless "task"`, prints the app's `--help`, and provides `headlessStartup`; the runner injects that service and reads its task from lazy config. A missing or whitespace-only task is rejected before the runner activates.
|
||||
|
||||
## Model Experience
|
||||
|
||||
@@ -17,4 +19,6 @@ None; the runner adds nothing to the request prefix.
|
||||
## Known Limitations and Deferred Work
|
||||
|
||||
- **One submitted task only** — the runner has no interactive follow-up surface; it waits through any work the Agent completes before returning to idle and prints the last non-empty assistant message in that interval.
|
||||
- **No pre-token heartbeat** — stderr remains silent until the provider emits a non-empty reasoning delta; a provider that delays its first streamed token exposes no earlier progress signal.
|
||||
- **Reasoning enters stderr logs** — redirection and supervisors may retain substantially more and potentially sensitive model output; route stderr to a controlled sink when that content must not be collected.
|
||||
- **`ctx.appExit` is launcher-owned** — booting the headless profile outside the `dsh` launcher fails loud at activation until the host provides the exit request.
|
||||
|
||||
@@ -4,7 +4,9 @@
|
||||
|
||||
dsh 一次性任务组合包。[`cordis.patch.yml`](cordis.patch.yml) 直接叠加在 [`dsh-base`](../base/README.zh.md) 之上:继承 base 默认禁用模块 HMR(热模块替换)的策略,提供编码 persona 和工具模式,将 Code Mode 的 worker 作为核心执行能力挂载,并插入本包的 `headless-runner` 插件(配置为 `{task}`,从注入的 `headlessStartup` 提供方解析)。它不挂载任何 Host、HTTP server、Web runtime 或浏览器插件。
|
||||
|
||||
Loader 结算后,runner 读取共享的 [`ctx.agentDefaultModel`](../../core/agent-default-model/README.zh.md),通过 `ctx.agents` 创建一个全新的持久化 Agent(智能体),将任务作为普通用户消息提交,并等待完全停稳。它对 Session 执行 flush 后再汇总自身持有的持久化事件区间,将最后一条非空 assistant 文本写入 stdout,再经启动器提供的 `ctx.appExit` 宿主钩子([`dsh-cmdline`](../../boot/cmdline/README.zh.md))请求退出(最终 `turn/end` 完成 → 0,否则为 1)。最终结束原因为 `error` 时,还会将 code 与 message 写入 stderr;成功运行时 stderr 保持为空。进程不会打开监听端口。任务文本就是这个应用的命令行:普通 `headless-startup` 提供方([`src/startup.ts`](src/startup.ts))注入 `ctx.cmdlineArgs`([`dsh-cmdline`](../../boot/cmdline/README.zh.md)),读取 `dsh --profile headless "task"` 的位置参数、打印应用自己的 `--help`,并提供 `headlessStartup`;runner 注入该服务,再从惰性配置中读取任务。缺失或只有空白的任务会在 runner 激活前被拒绝。
|
||||
Loader 结算后,runner 读取共享的 [`ctx.agentDefaultModel`](../../core/agent-default-model/README.zh.md),通过 `ctx.agents` 创建一个全新的持久化 Agent(智能体),将任务作为普通用户消息提交,并等待完全停稳。该 Agent 每次产生非空的提供方推理分片时,runner 都会在 `dsh: reasoning:` 标题下将其即时写入 stderr;连续分片保留在同一段中,提供方没有输出末尾换行时,runner 会在后续输出前终止该段。随后,它对 Session 执行 flush,再汇总自身持有的持久化事件区间,将最后一条非空 assistant 文本写入 stdout,并经启动器提供的 `ctx.appExit` 宿主钩子([`dsh-cmdline`](../../boot/cmdline/README.zh.md))请求退出(最终 `turn/end` 完成 → 0,否则为 1)。最终结束原因为 `error` 时,还会将 code 与 message 写入 stderr;没有推理内容的成功运行会保持 stderr 为空。进程不会打开监听端口。
|
||||
|
||||
任务文本就是这个应用的命令行:普通 `headless-startup` 提供方([`src/startup.ts`](src/startup.ts))注入 `ctx.cmdlineArgs`([`dsh-cmdline`](../../boot/cmdline/README.zh.md)),读取 `dsh --profile headless "task"` 的位置参数、打印应用自己的 `--help`,并提供 `headlessStartup`;runner 注入该服务,再从惰性配置中读取任务。缺失或只有空白的任务会在 runner 激活前被拒绝。
|
||||
|
||||
## 模型体验
|
||||
|
||||
@@ -17,4 +19,6 @@ Loader 结算后,runner 读取共享的 [`ctx.agentDefaultModel`](../../core/a
|
||||
## 已知限制与暂缓事项
|
||||
|
||||
- **只提交一个任务**:runner 没有用于交互式后续输入的 surface;它会等待 Agent 在返回 idle 前完成的所有工作,并打印该区间内最后一条非空 assistant 消息。
|
||||
- **首个 token 前没有心跳**:在提供方发出非空推理分片前,stderr 保持静默;如果提供方延迟首个流式 token,系统不会提供更早的进度信号。
|
||||
- **推理会进入 stderr 日志**:重定向与监督进程可能保留明显更多且可能敏感的模型输出;不得收集该内容时,应将 stderr 送往受控目标。
|
||||
- **`ctx.appExit` 由启动器持有**:在 `dsh` 启动器之外启动 headless profile 会在激活时明确报错,直到宿主提供该退出请求。
|
||||
|
||||
@@ -2,7 +2,8 @@
|
||||
* @deepseek-ai/dsh-headless — one-shot direct Agent driver. The bundle patch
|
||||
* rides over dsh-base without Host, HTTP, or browser plugins; this runner
|
||||
* creates one Agent through the core registry, drives the task to quiescence,
|
||||
* flushes its Session, prints the final assistant text, and exits.
|
||||
* streams provider reasoning to stderr, flushes its Session, prints the final
|
||||
* assistant text to stdout, and exits.
|
||||
*
|
||||
* @module @deepseek-ai/dsh-headless
|
||||
*/
|
||||
@@ -11,9 +12,9 @@ import { randomUUID } from 'node:crypto'
|
||||
import type { Context } from '@deepseek-ai/cordis'
|
||||
import z from '@deepseek-ai/schemastery'
|
||||
import { installModelSelection } from '@deepseek-ai/dsh-agent'
|
||||
import type { ModelSelectionRef } from '@deepseek-ai/dsh-agent'
|
||||
import type { Agent, ModelSelectionRef } from '@deepseek-ai/dsh-agent'
|
||||
import type {} from '@deepseek-ai/dsh-agent-default-model'
|
||||
import { createUserMessage } from '@deepseek-ai/dsh-llm'
|
||||
import { assertNever, createUserMessage } from '@deepseek-ai/dsh-llm'
|
||||
import { SessionId } from '@deepseek-ai/dsh-session'
|
||||
import type { SessionEvent } from '@deepseek-ai/dsh-session'
|
||||
// Empty type imports carry the loader Context merge for the settlement await
|
||||
@@ -81,6 +82,71 @@ function summarize(events: readonly SessionEvent[], firstSeq: number): RunOutcom
|
||||
return { text, reason }
|
||||
}
|
||||
|
||||
/**
|
||||
* Project provider-reported reasoning from one owned run to stderr as it is
|
||||
* appended, while keeping final outcome derivation on the durable log.
|
||||
* @param ctx - plugin context carrying the Session event feed.
|
||||
* @param agent - the exact Agent whose reasoning belongs to this invocation.
|
||||
* @param stderr - progress output sink.
|
||||
* @returns a disposer that also terminates an unterminated reasoning line.
|
||||
*/
|
||||
function streamReasoning(
|
||||
ctx: Context,
|
||||
agent: Agent,
|
||||
stderr: HeadlessIo['stderr'],
|
||||
): () => void {
|
||||
let started = false
|
||||
let open = false
|
||||
let endsWithNewline = true
|
||||
const close = (): void => {
|
||||
if (!open) return
|
||||
if (!endsWithNewline) stderr.write('\n')
|
||||
open = false
|
||||
endsWithNewline = true
|
||||
}
|
||||
const dispose = ctx.on('session/event', (session, event) => {
|
||||
if (session !== agent.session) return
|
||||
if (event.type === 'turn/start') {
|
||||
close()
|
||||
started = true
|
||||
return
|
||||
}
|
||||
if (!started || event.type !== 'assistant/chunk') return
|
||||
const chunk = event.data.chunk
|
||||
switch (chunk.type) {
|
||||
case 'reasoning-delta':
|
||||
if (chunk.text === '') return
|
||||
if (!open) {
|
||||
stderr.write('dsh: reasoning:\n')
|
||||
open = true
|
||||
}
|
||||
stderr.write(chunk.text)
|
||||
endsWithNewline = chunk.text.endsWith('\n')
|
||||
return
|
||||
case 'block-start':
|
||||
if (chunk.blockType !== 'reasoning') close()
|
||||
return
|
||||
case 'block-end':
|
||||
if (chunk.block.type !== 'reasoning') close()
|
||||
return
|
||||
case 'usage':
|
||||
return
|
||||
case 'text-delta':
|
||||
case 'tool-call-delta':
|
||||
case 'finish':
|
||||
close()
|
||||
return
|
||||
/* v8 ignore next -- closed-union exhaustiveness guard */
|
||||
default:
|
||||
return assertNever(chunk, 'headless reasoning stream')
|
||||
}
|
||||
})
|
||||
return () => {
|
||||
dispose()
|
||||
close()
|
||||
}
|
||||
}
|
||||
|
||||
/** Report an unexpected direct-driver failure and request a failing exit. */
|
||||
function fail(io: HeadlessIo, error: unknown): void {
|
||||
io.stderr.write(`dsh: ${error instanceof Error ? error.message : String(error)}\n`)
|
||||
@@ -119,11 +185,16 @@ async function run(ctx: Context, task: string, io: HeadlessIo): Promise<void> {
|
||||
})
|
||||
await agent.whenIdle()
|
||||
const firstSeq = agent.session.seq
|
||||
agent.followup(createUserMessage({
|
||||
content: [{ type: 'text', text: task }],
|
||||
source: { kind: 'user' },
|
||||
}))
|
||||
await agent.whenIdle()
|
||||
const stopReasoning = streamReasoning(ctx, agent, io.stderr)
|
||||
try {
|
||||
agent.followup(createUserMessage({
|
||||
content: [{ type: 'text', text: task }],
|
||||
source: { kind: 'user' },
|
||||
}))
|
||||
await agent.whenIdle()
|
||||
} finally {
|
||||
stopReasoning()
|
||||
}
|
||||
await sessions.flush(agent.session)
|
||||
const outcome = summarize(agent.session.events, firstSeq)
|
||||
io.stdout.write(outcome.text + '\n')
|
||||
|
||||
@@ -14,10 +14,10 @@ export const name = 'headless-invariant'
|
||||
export const inject = ['invariants']
|
||||
|
||||
/**
|
||||
* No runtime invariant: the runner is a one-shot driver over the API carrier
|
||||
* whose observable contract (final text on stdout, exit code by turn-end
|
||||
* reason) is process-level and owned by the launcher e2e; it registers
|
||||
* nothing and holds no mutable relation to audit inside the tree.
|
||||
* No runtime invariant: the runner's observable contract (provider reasoning
|
||||
* on stderr, final text on stdout, exit code by turn-end reason) is
|
||||
* process-level and owned by the launcher e2e; it registers nothing and holds
|
||||
* no mutable relation to audit inside the tree.
|
||||
*/
|
||||
const install: InvariantInstaller = () => {}
|
||||
|
||||
|
||||
@@ -31,7 +31,7 @@ export interface HeadlessStartupValues {
|
||||
function headlessCommand(): Command {
|
||||
return new Command()
|
||||
.name('dsh --profile headless')
|
||||
.description('Answer one task, print the final assistant message, and exit.')
|
||||
.description('Answer one task, stream reasoning to stderr, print the final assistant message, and exit.')
|
||||
.helpOption('-h, --help', 'show this help')
|
||||
.argument('[task...]', 'the task text; multiple words are joined by spaces')
|
||||
.addHelpText('after', `
|
||||
|
||||
@@ -50,9 +50,13 @@ function appendTurn(
|
||||
/** Mount the real registries around a small scripted Agent factory. */
|
||||
async function bench(script: Script): Promise<{
|
||||
ctx: Context
|
||||
output(): { out: string; err: string; order: string[] }
|
||||
run(): Promise<{ code: number; out: string; err: string; order: string[] }>
|
||||
}> {
|
||||
const ctx = new Context()
|
||||
let out = ''
|
||||
let err = ''
|
||||
const order: string[] = []
|
||||
await ctx.plugin(SessionStore)
|
||||
await ctx.plugin(AgentRegistry)
|
||||
await ctx.plugin(AgentDefaultModelConfig, { provider: 'test-provider', model: 'test-model' })
|
||||
@@ -91,10 +95,8 @@ async function bench(script: Script): Promise<{
|
||||
})
|
||||
return {
|
||||
ctx,
|
||||
output: () => ({ out, err, order: [...order] }),
|
||||
run: async () => {
|
||||
let out = ''
|
||||
let err = ''
|
||||
const order: string[] = []
|
||||
ctx.on('session/flush', () => { order.push('flush') })
|
||||
internals.stdout = { write: (chunk: string) => { out += chunk; return true } }
|
||||
internals.stderr = { write: (chunk: string) => { err += chunk; return true } }
|
||||
@@ -143,6 +145,110 @@ describe('headless runner', () => {
|
||||
await test.ctx.fiber.dispose()
|
||||
})
|
||||
|
||||
it('streams reasoning before the Agent becomes idle and terminates its stderr line', async () => {
|
||||
const reasoningAppended = Promise.withResolvers<undefined>()
|
||||
const release = Promise.withResolvers<undefined>()
|
||||
const test = await bench({
|
||||
async afterPrompt(session, message) {
|
||||
session.append('turn/start', { turn: 1 })
|
||||
session.append('step/start', { turn: 1, step: 1 })
|
||||
session.append('user/message', message, { surfaceOp: 'append' })
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'block-start', index: 0, blockType: 'reasoning' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 0, text: '' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 0, text: 'checking the workspace' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 0, text: ' safely\n' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'block-end', index: 0, block: { type: 'reasoning', text: 'checking the workspace safely\n' } },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'usage', usage: { inputTokens: 1, outputTokens: 2, reasoningTokens: 2 } },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'block-start', index: 1, blockType: 'reasoning' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 1, text: 'second pass\n' },
|
||||
})
|
||||
reasoningAppended.resolve(undefined)
|
||||
await release.promise
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'block-start', index: 2, blockType: 'text' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'text-delta', index: 2, text: 'done' },
|
||||
})
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'block-end', index: 2, block: { type: 'text', text: 'done' } },
|
||||
})
|
||||
session.append('assistant/message', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
message: createAssistantMessage({
|
||||
content: [{ type: 'text', text: 'done' }],
|
||||
source: { provider: 'test-provider', model: 'test-model' },
|
||||
}),
|
||||
}, { surfaceOp: 'append' })
|
||||
session.append('step/end', { turn: 1, step: 1 })
|
||||
session.append('turn/end', { turn: 1, reason: { kind: 'completed' } })
|
||||
},
|
||||
})
|
||||
const running = test.run()
|
||||
await reasoningAppended.promise
|
||||
const other = test.ctx.sessions.create()
|
||||
other.append('turn/start', { turn: 1 })
|
||||
other.append('step/start', { turn: 1, step: 1 })
|
||||
other.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 0, text: 'other session' },
|
||||
})
|
||||
const streamed = test.output()
|
||||
release.resolve(undefined)
|
||||
const result = await running
|
||||
expect(streamed).toEqual({
|
||||
out: '',
|
||||
err: 'dsh: reasoning:\nchecking the workspace safely\nsecond pass\n',
|
||||
order: [],
|
||||
})
|
||||
expect(result).toEqual({
|
||||
code: 0,
|
||||
out: 'done\n',
|
||||
err: 'dsh: reasoning:\nchecking the workspace safely\nsecond pass\n',
|
||||
order: ['flush', 'exit'],
|
||||
})
|
||||
await test.ctx.fiber.dispose()
|
||||
})
|
||||
|
||||
it('exits 1 when the final turn does not complete', async () => {
|
||||
const test = await bench({
|
||||
afterPrompt(session, message) { appendTurn(session, 1, message, undefined, false) },
|
||||
@@ -172,6 +278,32 @@ describe('headless runner', () => {
|
||||
await test.ctx.fiber.dispose()
|
||||
})
|
||||
|
||||
it('separates an unterminated reasoning prefix from the terminal model failure', async () => {
|
||||
const test = await bench({
|
||||
afterPrompt(session, message) {
|
||||
session.append('turn/start', { turn: 1 })
|
||||
session.append('step/start', { turn: 1, step: 1 })
|
||||
session.append('user/message', message, { surfaceOp: 'append' })
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'reasoning-delta', index: 0, text: 'trying recovery' },
|
||||
})
|
||||
session.append('step/end', { turn: 1, step: 1 })
|
||||
session.append('turn/end', {
|
||||
turn: 1,
|
||||
reason: { kind: 'error', error: { code: 'SERVER', message: 'provider unavailable' } },
|
||||
})
|
||||
},
|
||||
})
|
||||
expect(await test.run()).toMatchObject({
|
||||
code: 1,
|
||||
out: '\n',
|
||||
err: 'dsh: reasoning:\ntrying recovery\ndsh: SERVER: provider unavailable\n',
|
||||
})
|
||||
await test.ctx.fiber.dispose()
|
||||
})
|
||||
|
||||
it('exits 1 when the owned interval contains no turn', async () => {
|
||||
const test = await bench({ afterPrompt: () => {} })
|
||||
expect(await test.run()).toMatchObject({ code: 1, out: '\n', err: '' })
|
||||
|
||||
@@ -99,6 +99,7 @@ describe('headless command-line provider', () => {
|
||||
it('prints its own help and leaves the runner pending', async () => {
|
||||
const { task, observed } = await bootStartup(['--help'])
|
||||
expect(observed.out).toContain('dsh --profile headless')
|
||||
expect(observed.out).toContain('stream reasoning to stderr')
|
||||
expect(task).toBeUndefined()
|
||||
expect(observed.runnerConfig).toBeUndefined()
|
||||
expect(observed.exits).toEqual([0])
|
||||
|
||||
@@ -53,12 +53,12 @@ function styleInjectionModule(
|
||||
}
|
||||
|
||||
/**
|
||||
* Wire/type layers a client bundle may inline: browser-safe contracts
|
||||
* with no runtime identity to share (no Symbol/instanceof/singleton state).
|
||||
* Contract layers and pure folds a client bundle may inline: browser-safe
|
||||
* values with no runtime identity to share (no Symbol/instanceof/singleton state).
|
||||
* Everything else under @deepseek-ai/* is either a module-table entry
|
||||
* (external) or a leak the purity gate rejects.
|
||||
*/
|
||||
export const INLINE_SAFE = /^@deepseek-ai\/dsh-(?:host-apiproxy|file-reference|session|llm|tools|brand|util-crypto|util-workspace-path)(?:\/|$)/
|
||||
export const INLINE_SAFE = /^(?:@deepseek-ai\/dsh-(?:host-apiproxy|file-reference|session|llm|tools|brand|util-crypto|util-workspace-path)(?:\/|$)|@deepseek-ai\/dsh-token-meter\/client$)/
|
||||
|
||||
/**
|
||||
* Vendored framework libraries: rescoped into @deepseek-ai, so the gate below
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/client/ui-chat/README.md
|
||||
README.md: ef9dc65de0d6b990fd0066c387518dc932bd4d2e
|
||||
README.zh.md: c4de06b18077485d7d65734b9bb38ff7745a4d67
|
||||
README.md: cc79de10289069ef94105397bd77a5194b4e6808
|
||||
README.zh.md: 3d4eb91492a497ff4544bd6378ae212810342c64
|
||||
|
||||
@@ -19,3 +19,4 @@ None; Chat presentation does not assemble or mutate provider requests.
|
||||
## Known Limitations and Deferred Work
|
||||
|
||||
- **The view reflects the loaded Session window** — older transcript nodes become available only after Session Controller loads the preceding event page.
|
||||
- **Per-Turn token usage is fail-closed** — a completed Turn shows its disclosure only when the loaded window includes `turn/start` and every started model attempt has safe, exact usage. Missing buckets are omitted, and incomplete or contradictory accounting hides the whole disclosure.
|
||||
|
||||
@@ -19,3 +19,4 @@ Chat 会为非空的初始或恢复请求、显式序列起点,或 system 字
|
||||
## 已知限制与暂缓事项
|
||||
|
||||
- **视图只反映已加载的 Session 窗口**——只有 Session Controller 加载前一页 event 后,更早的 transcript node 才会出现。
|
||||
- **单轮次 token 用量采用 fail-closed 方式**——只有已加载窗口包含 `turn/start`,且每个已开始的模型 attempt 都具有安全、精确的用量时,已完成轮次才显示 disclosure。缺失的 bucket 会被省略,记账不完整或矛盾时则隐藏整条 disclosure。
|
||||
|
||||
@@ -13,7 +13,6 @@ import { formatRunDuration } from './message-chrome.ts'
|
||||
import css from './ChatView.module.css'
|
||||
|
||||
const FOLLOW_THRESHOLD = 24
|
||||
const MAX_PAGING_ANCHOR_PROBES = 64
|
||||
|
||||
/** Active column host when present; otherwise the view-local scroller. */
|
||||
function scrollerOf(from: HTMLElement): HTMLElement {
|
||||
@@ -46,38 +45,30 @@ function pagingAnchor(list: HTMLElement, scrollport: HTMLElement): HTMLElement |
|
||||
const viewport = scrollport.getBoundingClientRect()
|
||||
const composer = scrollport.querySelector<HTMLElement>('[data-composer-seat]')
|
||||
const visibleBottom = composer?.getBoundingClientRect().top ?? viewport.bottom
|
||||
// Scroll events are hot: walk down one hit-test line and stop at the first
|
||||
// hit row with layout before considering the full mounted set. Starting at the
|
||||
// viewport edge preserves the reader's leading row when a later row is
|
||||
// inserted between already-visible messages. The fallback keeps jsdom and
|
||||
// pre-layout states deterministic; a virtualizer naturally bounds it.
|
||||
// The leading edge preserves nested call identity when it hits a row.
|
||||
// Chrome/gap misses use logarithmic layout reads over the ordered flex rows.
|
||||
if (typeof document.elementsFromPoint === 'function' && visibleBottom > viewport.top) {
|
||||
const content = list.getBoundingClientRect()
|
||||
const left = Math.max(viewport.left, content.left)
|
||||
const right = Math.min(viewport.right, content.right)
|
||||
const x = left + Math.max(0, right - left) / 2
|
||||
const height = visibleBottom - viewport.top
|
||||
let probes = 0
|
||||
for (
|
||||
let offset = 1;
|
||||
offset < height && probes < MAX_PAGING_ANCHOR_PROBES;
|
||||
offset = offset === 1 ? 16 : offset + 16
|
||||
) {
|
||||
probes++
|
||||
for (const element of document.elementsFromPoint(x, viewport.top + offset)) {
|
||||
const row = element instanceof HTMLElement
|
||||
? element.closest<HTMLElement>('[data-chat-anchor-key]')
|
||||
: null
|
||||
if (row !== null && list.contains(row)) return row
|
||||
}
|
||||
for (const element of document.elementsFromPoint(x, viewport.top + 1)) {
|
||||
const row = element instanceof HTMLElement
|
||||
? element.closest<HTMLElement>('[data-chat-anchor-key]')
|
||||
: null
|
||||
if (row !== null && list.contains(row)) return row
|
||||
}
|
||||
}
|
||||
const rows = [...list.querySelectorAll<HTMLElement>('[data-chat-anchor-key]')]
|
||||
const visibleRows = rows.filter((row) => {
|
||||
const rect = row.getBoundingClientRect()
|
||||
return rect.bottom > viewport.top && rect.top < visibleBottom
|
||||
})
|
||||
return visibleRows[0] ?? rows[0] ?? null
|
||||
const rows = list.querySelectorAll<HTMLElement>('[data-chat-flow] > [data-chat-flow-key]:not(:empty)')
|
||||
let low = 0
|
||||
let high = rows.length
|
||||
while (low < high) {
|
||||
const middle = (low + high) >>> 1
|
||||
if (rows.item(middle).getBoundingClientRect().bottom > viewport.top) high = middle
|
||||
else low = middle + 1
|
||||
}
|
||||
const row = rows[low]
|
||||
return row !== undefined && row.getBoundingClientRect().top < visibleBottom ? row : rows[0] ?? null
|
||||
}
|
||||
|
||||
type ChatScrollPosition = NonNullable<ReturnType<ChatViewSlotProps['chatScroll']['read']>>
|
||||
|
||||
@@ -13,6 +13,7 @@ import type { ChatViewSlotProps } from '../contract/slots.ts'
|
||||
import type { ChatSnapshot } from '../contract/snapshot.ts'
|
||||
import { formatTokensPerSecond } from './message-chrome.ts'
|
||||
import { assistantStepReading } from '../contract/turn-metrics.ts'
|
||||
import { formatCacheHitPercent, formatTokens } from './token-format.ts'
|
||||
import css from './StatsLine.module.css'
|
||||
|
||||
interface WindowStats {
|
||||
@@ -77,19 +78,6 @@ export function deriveStats(nodes: ChatSnapshot['legacy']['nodes']): WindowStats
|
||||
return { turns: turns.size, steps, llmMs, toolMs, ttftMs, ttftSteps, decodeMs, decodeTokens }
|
||||
}
|
||||
|
||||
/**
|
||||
* Compact token count: 517 / 12.2K / 517K / 1.2M (one decimal under three digits).
|
||||
* @param n - token count.
|
||||
* @returns display string.
|
||||
*/
|
||||
export function formatTokens(n: number, t: ChatViewSlotProps['t']): string {
|
||||
const scaled = (v: number): string =>
|
||||
v >= 100 ? String(Math.round(v)) : String(Math.round(v * 10) / 10)
|
||||
if (n < 1_000) return String(n)
|
||||
if (n < 1_000_000) return t('number.thousand', { value: scaled(n / 1_000) })
|
||||
return t('number.million', { value: scaled(n / 1_000_000) })
|
||||
}
|
||||
|
||||
/**
|
||||
* Compact duration: 45.2s under a minute, 2m42s from there on.
|
||||
* @param ms - duration in milliseconds.
|
||||
@@ -105,26 +93,6 @@ export function formatDuration(ms: number, t: ChatViewSlotProps['t']): string {
|
||||
})
|
||||
}
|
||||
|
||||
/** Round a cache-read ratio to an integer percentage, with positive ties rounded up. */
|
||||
function roundedIntegerPercent(cacheReadTokens: number, denominator: number): number {
|
||||
const denominatorQuotient = Math.floor(denominator / 200)
|
||||
const denominatorRemainder = denominator % 200
|
||||
let lower = 0
|
||||
let upper = 100
|
||||
while (lower < upper) {
|
||||
const candidate = Math.floor((lower + upper + 1) / 2)
|
||||
const factor = candidate * 2 - 1
|
||||
const threshold = factor * denominatorQuotient
|
||||
+ Math.ceil(factor * denominatorRemainder / 200)
|
||||
if (cacheReadTokens >= threshold) {
|
||||
lower = candidate
|
||||
} else {
|
||||
upper = candidate - 1
|
||||
}
|
||||
}
|
||||
return lower
|
||||
}
|
||||
|
||||
/**
|
||||
* Display-ready cache-hit share of prompt-side input over the whole durable log.
|
||||
* @param usage - the session's token-usage projection value.
|
||||
@@ -134,35 +102,7 @@ function roundedIntegerPercent(cacheReadTokens: number, denominator: number): nu
|
||||
*/
|
||||
export function cacheHitPercent(usage: TokenUsageProjection): string | null {
|
||||
const denominator = billedInputTokens(usage)
|
||||
if (denominator === 0) return null
|
||||
const missedInputTokens = usage.uncachedInputTokens + usage.cacheWriteTokens
|
||||
if (missedInputTokens === 0) return '100'
|
||||
|
||||
const integerPercent = roundedIntegerPercent(usage.cacheReadTokens, denominator)
|
||||
if (integerPercent < 100) return String(integerPercent)
|
||||
|
||||
// At the first distinguishing precision, the rounded result is 100 minus
|
||||
// one to five units in the final decimal place. Scale only while the next
|
||||
// multiplication remains at or below the denominator, then derive that
|
||||
// final digit through exact small-factor comparisons.
|
||||
let decimalPlaces = 1
|
||||
let scaledDoubleGap = missedInputTokens * 200
|
||||
const denominatorTens = Math.floor(denominator / 10)
|
||||
while (scaledDoubleGap <= denominatorTens) {
|
||||
scaledDoubleGap *= 10
|
||||
decimalPlaces += 1
|
||||
}
|
||||
const denominatorOnes = denominator % 10
|
||||
let roundedLoss = 5
|
||||
for (let loss = 1; loss < 5; loss += 1) {
|
||||
const factor = loss * 2 + 1
|
||||
const threshold = factor * denominatorTens + Math.floor(factor * denominatorOnes / 10)
|
||||
if (scaledDoubleGap <= threshold) {
|
||||
roundedLoss = loss
|
||||
break
|
||||
}
|
||||
}
|
||||
return `99.${'9'.repeat(decimalPlaces - 1)}${10 - roundedLoss}`
|
||||
return formatCacheHitPercent(usage.cacheReadTokens, denominator)
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -4,6 +4,13 @@
|
||||
gap: 16px;
|
||||
}
|
||||
|
||||
.footer {
|
||||
display: flex;
|
||||
min-width: 0;
|
||||
flex-direction: column;
|
||||
gap: 4px;
|
||||
}
|
||||
|
||||
.actions {
|
||||
margin-left: -6px;
|
||||
}
|
||||
|
||||
@@ -2,6 +2,7 @@ import { memo } from 'react'
|
||||
import type { PropsRenderSlots } from '@deepseek-ai/dsh-client-ui-slots'
|
||||
import type { ChatNodeViewProps, TurnTailOwnerProps } from '../contract/slots.ts'
|
||||
import { MessageIconActions } from './MessageIconActions.tsx'
|
||||
import { TurnUsageDisclosure } from './TurnUsageDisclosure.tsx'
|
||||
import { assistantText } from './turn-assistant.ts'
|
||||
import css from './TurnTailNodeView.module.css'
|
||||
|
||||
@@ -35,19 +36,22 @@ export const TurnTailNodeView = memo(function TurnTailNodeView({
|
||||
return (
|
||||
<div className={css.root} data-turn-tail={data.turn} data-time-hover-root>
|
||||
{tail}
|
||||
<MessageIconActions
|
||||
text={assistantText(closing.blocks)}
|
||||
time={closing.time}
|
||||
runMs={runMs}
|
||||
ttftMs={data.ttftMs}
|
||||
tokensPerSecond={data.tokensPerSecond}
|
||||
clock="end"
|
||||
onBranch={() => { forkAt(closing.finalNode.seq) }}
|
||||
branchUnavailable={data.branchUnavailable || hasLaterChatNode}
|
||||
className={css.actions}
|
||||
extraActions={assistantActions}
|
||||
t={t}
|
||||
/>
|
||||
<div className={css.footer}>
|
||||
{data.tokenUsage === undefined ? null : <TurnUsageDisclosure usage={data.tokenUsage} t={t} />}
|
||||
<MessageIconActions
|
||||
text={assistantText(closing.blocks)}
|
||||
time={closing.time}
|
||||
runMs={runMs}
|
||||
ttftMs={data.ttftMs}
|
||||
tokensPerSecond={data.tokensPerSecond}
|
||||
clock="end"
|
||||
onBranch={() => { forkAt(closing.finalNode.seq) }}
|
||||
branchUnavailable={data.branchUnavailable || hasLaterChatNode}
|
||||
className={css.actions}
|
||||
extraActions={assistantActions}
|
||||
t={t}
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
})
|
||||
|
||||
@@ -0,0 +1,87 @@
|
||||
.root {
|
||||
min-width: 0;
|
||||
}
|
||||
|
||||
.root[data-open] {
|
||||
padding-bottom: 4px;
|
||||
}
|
||||
|
||||
.root [data-disclosure-row]:focus-visible {
|
||||
border-radius: 6px;
|
||||
outline: 2px solid var(--dsw-alias-label-tertiary);
|
||||
outline-offset: -2px;
|
||||
}
|
||||
|
||||
.chevron {
|
||||
color: var(--dsw-alias-label-secondary);
|
||||
}
|
||||
|
||||
.separator {
|
||||
flex: none;
|
||||
width: 2px;
|
||||
height: 2px;
|
||||
margin: 0 8px;
|
||||
border-radius: 1px;
|
||||
background: var(--dsw-alias-label-caption);
|
||||
}
|
||||
|
||||
.summary {
|
||||
min-width: 0;
|
||||
overflow: hidden;
|
||||
color: var(--dsw-alias-label-tertiary);
|
||||
font-size: 14px;
|
||||
font-variant-numeric: tabular-nums;
|
||||
line-height: 24px;
|
||||
text-overflow: ellipsis;
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
.details {
|
||||
display: grid;
|
||||
grid-template-columns: minmax(76px, auto) minmax(0, 1fr);
|
||||
gap: 6px 16px;
|
||||
box-sizing: border-box;
|
||||
width: calc(100% - 22px);
|
||||
margin: 4px 0 0 22px;
|
||||
padding: 10px 16px 12px 12px;
|
||||
border-radius: 8px;
|
||||
background: var(--dsw-alias-markdown-code-block);
|
||||
color: var(--dsw-alias-label-tertiary);
|
||||
font-size: 12px;
|
||||
line-height: 18px;
|
||||
}
|
||||
|
||||
.details dt,
|
||||
.details dd {
|
||||
min-width: 0;
|
||||
margin: 0;
|
||||
}
|
||||
|
||||
.details dd {
|
||||
color: var(--dsw-alias-label-secondary);
|
||||
font-variant-numeric: tabular-nums;
|
||||
text-align: right;
|
||||
}
|
||||
|
||||
.details .route {
|
||||
overflow-wrap: anywhere;
|
||||
}
|
||||
|
||||
.reasoning {
|
||||
color: var(--dsw-alias-label-tertiary);
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
.totalLabel,
|
||||
.details .totalValue {
|
||||
padding-top: 6px;
|
||||
border-top: 1px solid var(--dsw-alias-separator-primary);
|
||||
color: var(--dsw-alias-label-primary);
|
||||
}
|
||||
|
||||
@media (max-width: 480px) {
|
||||
.details {
|
||||
grid-template-columns: minmax(72px, auto) minmax(0, 1fr);
|
||||
gap-inline: 10px;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,86 @@
|
||||
import { useState } from 'react'
|
||||
import { DisclosureRow, IconDataOutline16 } from '@deepseek-ai/dsh-client-ui-primitives'
|
||||
import type { TurnTokenUsage } from '../contract/chat-nodes.ts'
|
||||
import type { ChatViewSlotProps } from '../contract/slots.ts'
|
||||
import { formatCacheHitPercent, formatExactTokens, formatTokens } from './token-format.ts'
|
||||
import css from './TurnUsageDisclosure.module.css'
|
||||
|
||||
export interface TurnUsageDisclosureProps {
|
||||
usage: TurnTokenUsage
|
||||
t: ChatViewSlotProps['t']
|
||||
}
|
||||
|
||||
function formatCompactCount(value: number, t: ChatViewSlotProps['t']): string {
|
||||
return t('message.turnUsage.count', { count: formatTokens(value, t) })
|
||||
}
|
||||
|
||||
function formatExactCount(value: number, t: ChatViewSlotProps['t']): string {
|
||||
return t('message.turnUsage.count', { count: formatExactTokens(value, t) })
|
||||
}
|
||||
|
||||
/** Compact per-Turn usage summary with an opt-in bucket breakdown. */
|
||||
export function TurnUsageDisclosure({ usage, t }: TurnUsageDisclosureProps) {
|
||||
const [open, setOpen] = useState(false)
|
||||
const cacheHit = usage.cacheReadTokens === undefined
|
||||
? null
|
||||
: formatCacheHitPercent(usage.cacheReadTokens, usage.totalTokens - usage.outputTokens, 1)
|
||||
const total = formatCompactCount(usage.totalTokens, t)
|
||||
const summary = cacheHit === null
|
||||
? total
|
||||
: t('message.turnUsage.summaryWithCache', { total, percent: cacheHit })
|
||||
const routes = usage.routes?.map(route => `${route.provider}/${route.model}`).join(', ') ?? ''
|
||||
|
||||
return (
|
||||
<DisclosureRow
|
||||
icon={<IconDataOutline16 />}
|
||||
title={t('message.turnUsage.title')}
|
||||
open={open}
|
||||
expandable
|
||||
onToggle={() => { setOpen(value => !value) }}
|
||||
expandOnRowClick
|
||||
keepContentWhenOpen
|
||||
collapsedContent={(
|
||||
<>
|
||||
<span className={css.separator} aria-hidden />
|
||||
<span className={css.summary}>{summary}</span>
|
||||
</>
|
||||
)}
|
||||
className={css.root}
|
||||
chevronClassName={css.chevron}
|
||||
>
|
||||
<dl className={css.details} data-turn-usage-details>
|
||||
{routes !== '' && (
|
||||
<>
|
||||
<dt>{t('message.turnUsage.model')}</dt>
|
||||
<dd className={css.route}>{routes}</dd>
|
||||
</>
|
||||
)}
|
||||
<dt>{t('message.turnUsage.input')}</dt>
|
||||
<dd>{formatExactCount(usage.uncachedInputTokens, t)}</dd>
|
||||
{usage.cacheReadTokens !== undefined && (
|
||||
<>
|
||||
<dt>{t('message.turnUsage.cacheRead')}</dt>
|
||||
<dd>{formatExactCount(usage.cacheReadTokens, t)}</dd>
|
||||
</>
|
||||
)}
|
||||
{usage.cacheWriteTokens !== undefined && (
|
||||
<>
|
||||
<dt>{t('message.turnUsage.cacheWrite')}</dt>
|
||||
<dd>{formatExactCount(usage.cacheWriteTokens, t)}</dd>
|
||||
</>
|
||||
)}
|
||||
<dt>{t('message.turnUsage.output')}</dt>
|
||||
<dd>
|
||||
{formatExactCount(usage.outputTokens, t)}
|
||||
{usage.reasoningTokens !== undefined && (
|
||||
<span className={css.reasoning}>
|
||||
{t('message.turnUsage.reasoning', { tokens: formatExactCount(usage.reasoningTokens, t) })}
|
||||
</span>
|
||||
)}
|
||||
</dd>
|
||||
<dt className={css.totalLabel}>{t('message.turnUsage.total')}</dt>
|
||||
<dd className={css.totalValue}>{formatExactCount(usage.totalTokens, t)}</dd>
|
||||
</dl>
|
||||
</DisclosureRow>
|
||||
)
|
||||
}
|
||||
@@ -0,0 +1,98 @@
|
||||
import type { ChatViewSlotProps } from '../contract/slots.ts'
|
||||
|
||||
/**
|
||||
* Compact token count: 517 / 12.2K / 517K / 1.2M.
|
||||
* @param value - non-negative token count.
|
||||
* @param t - Chat locale seat.
|
||||
* @returns locale-owned compact display string.
|
||||
*/
|
||||
export function formatTokens(value: number, t: ChatViewSlotProps['t']): string {
|
||||
const scaled = (candidate: number): string =>
|
||||
candidate >= 100 ? String(Math.round(candidate)) : String(Math.round(candidate * 10) / 10)
|
||||
if (value < 1_000) return String(value)
|
||||
if (value < 1_000_000) return t('number.thousand', { value: scaled(value / 1_000) })
|
||||
return t('number.million', { value: scaled(value / 1_000_000) })
|
||||
}
|
||||
|
||||
/**
|
||||
* Exact integer token count with locale-owned digit grouping.
|
||||
* @param value - non-negative safe integer token count.
|
||||
* @param t - Chat locale seat.
|
||||
* @returns an unrounded display string.
|
||||
*/
|
||||
export function formatExactTokens(value: number, t: ChatViewSlotProps['t']): string {
|
||||
const digits = String(value)
|
||||
const groups: string[] = []
|
||||
for (let end = digits.length; end > 0; end -= 3) {
|
||||
groups.unshift(digits.slice(Math.max(0, end - 3), end))
|
||||
}
|
||||
return groups.join(t('number.groupSeparator'))
|
||||
}
|
||||
|
||||
/** Round a cache-read ratio to exact percentage units, with positive ties rounded up. */
|
||||
function roundedPercentUnits(cacheReadTokens: number, denominator: number, decimalPlaces: 0 | 1): number {
|
||||
const unitsPerPercent = decimalPlaces === 0 ? 1 : 10
|
||||
const scale = unitsPerPercent * 100
|
||||
const doubledScale = scale * 2
|
||||
const denominatorQuotient = Math.floor(denominator / doubledScale)
|
||||
const denominatorRemainder = denominator % doubledScale
|
||||
let lower = 0
|
||||
let upper = scale
|
||||
while (lower < upper) {
|
||||
const candidate = Math.floor((lower + upper + 1) / 2)
|
||||
const factor = candidate * 2 - 1
|
||||
const threshold = factor * denominatorQuotient
|
||||
+ Math.ceil(factor * denominatorRemainder / doubledScale)
|
||||
if (cacheReadTokens >= threshold) lower = candidate
|
||||
else upper = candidate - 1
|
||||
}
|
||||
return lower
|
||||
}
|
||||
|
||||
function displayPercentUnits(units: number, decimalPlaces: 0 | 1): string {
|
||||
if (decimalPlaces === 0) return String(units)
|
||||
const whole = Math.floor(units / 10)
|
||||
const tenths = units % 10
|
||||
return tenths === 0 ? String(whole) : `${whole}.${tenths}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Display-ready cache-hit share without rounding a partial hit to 100%.
|
||||
* @param cacheReadTokens - exact prompt tokens served from cache.
|
||||
* @param promptTokens - exact aggregate prompt tokens.
|
||||
* @param decimalPlaces - ordinary-ratio precision; partial hits that would
|
||||
* round to 100 automatically use enough additional precision to stay honest.
|
||||
* @returns percentage text, or null when there was no prompt input.
|
||||
*/
|
||||
export function formatCacheHitPercent(
|
||||
cacheReadTokens: number,
|
||||
promptTokens: number,
|
||||
decimalPlaces: 0 | 1 = 0,
|
||||
): string | null {
|
||||
if (promptTokens === 0) return null
|
||||
const missedInputTokens = promptTokens - cacheReadTokens
|
||||
if (missedInputTokens === 0) return '100'
|
||||
|
||||
const roundedUnits = roundedPercentUnits(cacheReadTokens, promptTokens, decimalPlaces)
|
||||
const fullHitUnits = decimalPlaces === 0 ? 100 : 1_000
|
||||
if (roundedUnits < fullHitUnits) return displayPercentUnits(roundedUnits, decimalPlaces)
|
||||
|
||||
let distinguishingPlaces = 1
|
||||
let scaledDoubleGap = missedInputTokens * 200
|
||||
const denominatorTens = Math.floor(promptTokens / 10)
|
||||
while (scaledDoubleGap <= denominatorTens) {
|
||||
scaledDoubleGap *= 10
|
||||
distinguishingPlaces += 1
|
||||
}
|
||||
const denominatorOnes = promptTokens % 10
|
||||
let roundedLoss = 5
|
||||
for (let loss = 1; loss < 5; loss += 1) {
|
||||
const factor = loss * 2 + 1
|
||||
const threshold = factor * denominatorTens + Math.floor(factor * denominatorOnes / 10)
|
||||
if (scaledDoubleGap <= threshold) {
|
||||
roundedLoss = loss
|
||||
break
|
||||
}
|
||||
}
|
||||
return `99.${'9'.repeat(distinguishingPlaces - 1)}${10 - roundedLoss}`
|
||||
}
|
||||
@@ -59,6 +59,29 @@ export interface RetryChatData {
|
||||
readonly current: ModelRetryNode
|
||||
}
|
||||
|
||||
/** One provider/model route that contributed a billed request attempt. */
|
||||
export interface TurnTokenUsageRoute {
|
||||
readonly provider: string
|
||||
readonly model: string
|
||||
}
|
||||
|
||||
/** Exact provider-reported token accounting for every attempt in one completed Turn. */
|
||||
export interface TurnTokenUsage {
|
||||
/** Sum of uncached prompt input across all attempts. */
|
||||
readonly uncachedInputTokens: number
|
||||
readonly outputTokens: number
|
||||
/** Exact aggregate prompt plus output total across all attempts. */
|
||||
readonly totalTokens: number
|
||||
/** Present only when every attempt reported the bucket. */
|
||||
readonly cacheReadTokens?: number
|
||||
/** Present only when every attempt reported the bucket. */
|
||||
readonly cacheWriteTokens?: number
|
||||
/** Output subset, present only when every attempt reported it. */
|
||||
readonly reasoningTokens?: number
|
||||
/** Present only when every billed attempt has provider/model attribution. */
|
||||
readonly routes?: readonly TurnTokenUsageRoute[]
|
||||
}
|
||||
|
||||
/** Turn-local footer row that owns actions and optional feature contributions. */
|
||||
export interface TurnTailChatData {
|
||||
readonly turn: number
|
||||
@@ -70,6 +93,8 @@ export interface TurnTailChatData {
|
||||
readonly branchUnavailable: boolean
|
||||
readonly ttftMs?: number
|
||||
readonly tokensPerSecond?: number
|
||||
/** Exact per-Turn accounting; absent when the loaded evidence is incomplete. */
|
||||
readonly tokenUsage?: TurnTokenUsage
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -4,6 +4,7 @@ import type {
|
||||
} from '@deepseek-ai/dsh-client-ui-conversation/client'
|
||||
import type {} from '@deepseek-ai/dsh-llm-retry/types'
|
||||
import { isAppendSurfaceEvent } from '@deepseek-ai/dsh-session/surface'
|
||||
import { deriveTurnTokenUsage } from '@deepseek-ai/dsh-token-meter/client'
|
||||
import type {
|
||||
AssistantChatData, FinalAssistantChatData, TurnTailChatData,
|
||||
} from '../contract/chat-nodes.ts'
|
||||
@@ -57,10 +58,13 @@ function turnCoordinates(event: Parameters<ConversationNodeDefinition['match']>[
|
||||
} | undefined {
|
||||
if (event.type === 'assistant/message'
|
||||
|| event.type === 'assistant/chunk'
|
||||
|| event.type === 'step/start'
|
||||
|| event.type === 'step/end') {
|
||||
return { turn: event.data.turn, step: event.data.step }
|
||||
}
|
||||
if (event.type === 'llm/retry') return { turn: event.data.turn, step: event.data.step }
|
||||
if (event.type === 'llm/retry' || event.type === 'llm/retry-started') {
|
||||
return { turn: event.data.turn, step: event.data.step }
|
||||
}
|
||||
return undefined
|
||||
}
|
||||
|
||||
@@ -138,6 +142,9 @@ function tailData(context: ConversationNodeContext<TurnTailState>): TurnTailChat
|
||||
}
|
||||
}
|
||||
const metrics = deriveTurnMetrics(finalized.map(candidate => candidate.finalNode)).get(end.event.data.turn)
|
||||
const tokenUsage = context.start?.event.type === 'turn/start'
|
||||
? deriveTurnTokenUsage(context.matches.map(match => match.event))
|
||||
: undefined
|
||||
return {
|
||||
turn: end.event.data.turn,
|
||||
seq: end.event.seq,
|
||||
@@ -146,6 +153,7 @@ function tailData(context: ConversationNodeContext<TurnTailState>): TurnTailChat
|
||||
branchUnavailable: closing === null || latestTranscriptSeq !== closing.finalNode.seq,
|
||||
...metrics?.ttftMs === undefined ? {} : { ttftMs: metrics.ttftMs },
|
||||
...metrics?.tokensPerSecond === undefined ? {} : { tokensPerSecond: metrics.tokensPerSecond },
|
||||
...tokenUsage === undefined ? {} : { tokenUsage },
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -6,6 +6,7 @@ export const NS = 'chat'
|
||||
/** Simplified Chinese dictionary and key-set source of truth. */
|
||||
export const zh = {
|
||||
'view.chat': '对话',
|
||||
'number.groupSeparator': ',',
|
||||
'duration.compactSeconds': '{seconds}秒',
|
||||
'duration.compactMinutes': '{minutes}分{seconds}秒',
|
||||
'duration.milliseconds': '{milliseconds}毫秒',
|
||||
@@ -74,6 +75,16 @@ export const zh = {
|
||||
'message.ranFor': '用时 {duration}',
|
||||
'message.ttft': '首 token {seconds}秒',
|
||||
'message.tokensPerSecond': '{tps} tok/s',
|
||||
'message.turnUsage.title': '本轮用量',
|
||||
'message.turnUsage.summaryWithCache': '{total} · 缓存命中率 {percent}%',
|
||||
'message.turnUsage.model': '提供方 / 模型',
|
||||
'message.turnUsage.input': '未缓存输入',
|
||||
'message.turnUsage.cacheRead': '缓存读取',
|
||||
'message.turnUsage.cacheWrite': '缓存写入',
|
||||
'message.turnUsage.output': '输出',
|
||||
'message.turnUsage.reasoning': '(其中推理 {tokens})',
|
||||
'message.turnUsage.total': '总计',
|
||||
'message.turnUsage.count': '{count} tok',
|
||||
'duration.seconds': '{seconds}秒',
|
||||
'duration.minutes': '{minutes}分{seconds}秒',
|
||||
'command.running': '执行中…',
|
||||
@@ -93,6 +104,7 @@ export type ChatKey = keyof typeof zh
|
||||
/** English dictionary, checked against the Chinese key set. */
|
||||
export const en = {
|
||||
'view.chat': 'Chat',
|
||||
'number.groupSeparator': ',',
|
||||
'duration.compactSeconds': '{seconds}s',
|
||||
'duration.compactMinutes': '{minutes}m{seconds}s',
|
||||
'duration.milliseconds': '{milliseconds}ms',
|
||||
@@ -161,6 +173,16 @@ export const en = {
|
||||
'message.ranFor': 'Ran for {duration}',
|
||||
'message.ttft': 'TTFT {seconds}s',
|
||||
'message.tokensPerSecond': '{tps} tok/s',
|
||||
'message.turnUsage.title': 'Turn usage',
|
||||
'message.turnUsage.summaryWithCache': '{total} · Cache hit {percent}%',
|
||||
'message.turnUsage.model': 'Provider / model',
|
||||
'message.turnUsage.input': 'Uncached input',
|
||||
'message.turnUsage.cacheRead': 'Cached input',
|
||||
'message.turnUsage.cacheWrite': 'Cache write',
|
||||
'message.turnUsage.output': 'Output',
|
||||
'message.turnUsage.reasoning': ' ({tokens} reasoning)',
|
||||
'message.turnUsage.total': 'Total',
|
||||
'message.turnUsage.count': '{count} tok',
|
||||
'duration.seconds': '{seconds}s',
|
||||
'duration.minutes': '{minutes}m {seconds}s',
|
||||
'command.running': 'Running…',
|
||||
|
||||
@@ -9,7 +9,8 @@ import { bindSnapshotSelector } from '@deepseek-ai/dsh-client-test-runtime'
|
||||
import { makeTranslate } from '@deepseek-ai/dsh-client-test-runtime'
|
||||
import { en as commonEn } from '@deepseek-ai/dsh-client-locale/src/locales/en.ts'
|
||||
import { zh as commonZh } from '@deepseek-ai/dsh-client-locale/src/locales/zh.ts'
|
||||
import { StatsLine, deriveStats, formatDuration, formatTokens, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
|
||||
import { StatsLine, deriveStats, formatDuration, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
|
||||
import { formatTokens } from '../src/client/chat/token-format.ts'
|
||||
import { en, zh } from '../src/client/locale.ts'
|
||||
import { chatSnapshotFixture } from './chat-snapshot-fixture.client.ts'
|
||||
|
||||
|
||||
@@ -474,7 +474,7 @@ describe('ChatView', () => {
|
||||
|
||||
readerScroll(scroller, 100)
|
||||
|
||||
expect(hitTest).toHaveBeenCalledTimes(64)
|
||||
expect(hitTest).toHaveBeenCalledTimes(1)
|
||||
expect(h.chatScroll.read()?.anchorKey).toBe('fixture:user:1')
|
||||
} finally {
|
||||
if (originalHitTest !== undefined) {
|
||||
@@ -485,6 +485,60 @@ describe('ChatView', () => {
|
||||
}
|
||||
})
|
||||
|
||||
it('falls back to the first visible row when the viewport top hit-test misses', () => {
|
||||
const originalHitTest = Object.getOwnPropertyDescriptor(document, 'elementsFromPoint')
|
||||
const nodes = Array.from({ length: 16 }, (_, index) => user(20 + index, `row ${String(index)}`))
|
||||
const h = makeHarness(
|
||||
{ nodes },
|
||||
{ hasMore: true },
|
||||
)
|
||||
const view = render(<h.ChatView {...h.props} />)
|
||||
const scroller = view.container.querySelector('[class*="scroll"]') as HTMLDivElement
|
||||
const rows = [...view.container.querySelectorAll<HTMLElement>('[data-chat-flow-key]')]
|
||||
let prepended = false
|
||||
let rowRectCalls = 0
|
||||
vi.spyOn(scroller, 'getBoundingClientRect').mockImplementation(
|
||||
() => ({ top: 0, bottom: 200 } as DOMRect),
|
||||
)
|
||||
rows.forEach((row, index) => {
|
||||
vi.spyOn(row, 'getBoundingClientRect').mockImplementation(() => {
|
||||
rowRectCalls += 1
|
||||
const shift = prepended ? (index === 8 ? 400 : 500) : 0
|
||||
const top = 20 + (index - 8) * 60 + shift
|
||||
return { top, bottom: top + 40 } as DOMRect
|
||||
})
|
||||
})
|
||||
Object.defineProperty(scroller, 'scrollHeight', { value: 800, writable: true })
|
||||
Object.defineProperty(scroller, 'clientHeight', { value: 200, writable: true })
|
||||
readerScroll(scroller, 50)
|
||||
|
||||
const hitTest = vi.fn((_x: number, _y: number): Element[] => [])
|
||||
Object.defineProperty(document, 'elementsFromPoint', {
|
||||
configurable: true,
|
||||
value: hitTest,
|
||||
})
|
||||
try {
|
||||
rowRectCalls = 0
|
||||
fireEvent.click(view.getByText('加载更早'))
|
||||
expect(hitTest).toHaveBeenCalledTimes(1)
|
||||
expect(hitTest.mock.calls[0]?.[1]).toBe(1)
|
||||
expect(rowRectCalls).toBeLessThanOrEqual(6)
|
||||
|
||||
Object.defineProperty(scroller, 'scrollHeight', { value: 1_300, writable: true })
|
||||
prepended = true
|
||||
act(() => {
|
||||
h.setChat({ nodes: [assistant(2, 'older'), ...nodes] })
|
||||
})
|
||||
expect(scroller.scrollTop).toBe(450) // reader offset 50 + first visible row's 400px shift
|
||||
} finally {
|
||||
if (originalHitTest !== undefined) {
|
||||
Object.defineProperty(document, 'elementsFromPoint', originalHitTest)
|
||||
} else {
|
||||
Reflect.deleteProperty(document, 'elementsFromPoint')
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
it('renders the fixture main line as independently keyed business nodes', () => {
|
||||
const h = makeHarness({
|
||||
nodes: [user(1, 'do the thing'), assistant(2, 'running tools'), toolResult(3, 'a'), toolResult(4, 'b')],
|
||||
|
||||
@@ -500,6 +500,44 @@ describe('built-in conversation node Definitions', () => {
|
||||
expect(tail.branchUnavailable).toBe(true)
|
||||
})
|
||||
|
||||
it('publishes exact Turn usage only after pagination supplies the full lifecycle window', () => {
|
||||
const value = assembler([
|
||||
at(3, 'assistant/message', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
message: assistantMessage('usage-assistant', 'done'),
|
||||
usage: {
|
||||
inputTokens: 10,
|
||||
outputTokens: 4,
|
||||
totalTokens: 17,
|
||||
cacheReadTokens: 2,
|
||||
cacheWriteTokens: 1,
|
||||
reasoningTokens: 1,
|
||||
},
|
||||
}, { surfaceOp: 'append' }),
|
||||
at(4, 'step/end', { turn: 1, step: 1 }),
|
||||
at(5, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
], true)
|
||||
|
||||
expect((node(snapshot(value), 'turn-tail')?.data as TurnTailChatData).tokenUsage).toBeUndefined()
|
||||
|
||||
value.prepend([
|
||||
at(1, 'turn/start', { turn: 1 }),
|
||||
at(2, 'step/start', { turn: 1, step: 1 }),
|
||||
], false)
|
||||
value.flush()
|
||||
|
||||
expect((node(snapshot(value), 'turn-tail')?.data as TurnTailChatData).tokenUsage).toEqual({
|
||||
uncachedInputTokens: 10,
|
||||
outputTokens: 4,
|
||||
totalTokens: 17,
|
||||
cacheReadTokens: 2,
|
||||
cacheWriteTokens: 1,
|
||||
reasoningTokens: 1,
|
||||
routes: [{ provider: 'fake', model: 'fake' }],
|
||||
})
|
||||
})
|
||||
|
||||
it('replays inbox predecessors after prepend and reclassifies the dependent message as steering', () => {
|
||||
const value = assembler([
|
||||
at(3, 'user/message', textMessage('steer-1', 'change direction'), { surfaceOp: 'append' }),
|
||||
|
||||
@@ -6,6 +6,7 @@ import type {
|
||||
} from '@deepseek-ai/dsh-client-ui-chat/client'
|
||||
import { assistantStepReading, deriveTurnMetrics } from '../src/client/contract/turn-metrics.ts'
|
||||
import { formatLatencySeconds, formatTokensPerSecond } from '../src/client/chat/message-chrome.ts'
|
||||
import { formatCacheHitPercent } from '../src/client/chat/token-format.ts'
|
||||
|
||||
interface StepSpec {
|
||||
seq: number
|
||||
@@ -139,6 +140,10 @@ describe('deriveTurnMetrics', () => {
|
||||
})
|
||||
|
||||
describe('footer figure formatters', () => {
|
||||
it('omits a redundant decimal zero in cache-hit percentages', () => {
|
||||
expect(formatCacheHitPercent(1, 2, 1)).toBe('50')
|
||||
})
|
||||
|
||||
it('formats latency with one decimal under ten seconds and whole seconds beyond', () => {
|
||||
expect(formatLatencySeconds(840)).toBe('0.8')
|
||||
expect(formatLatencySeconds(1_000)).toBe('1')
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
// @vitest-environment jsdom
|
||||
|
||||
import { afterEach, describe, expect, it } from 'vitest'
|
||||
import { cleanup, fireEvent, render } from '@testing-library/react'
|
||||
import { makeTranslate } from '@deepseek-ai/dsh-client-test-runtime'
|
||||
import { en as commonEn } from '@deepseek-ai/dsh-client-locale/src/locales/en.ts'
|
||||
import { TurnUsageDisclosure } from '../src/client/chat/TurnUsageDisclosure.tsx'
|
||||
import type { TurnTokenUsage } from '../src/client/contract/chat-nodes.ts'
|
||||
import { en } from '../src/client/locale.ts'
|
||||
|
||||
const t = makeTranslate(en, commonEn)
|
||||
|
||||
afterEach(cleanup)
|
||||
|
||||
describe('TurnUsageDisclosure', () => {
|
||||
it('shows the exact compact summary and expands into provider facts', () => {
|
||||
const usage: TurnTokenUsage = {
|
||||
uncachedInputTokens: 5_060,
|
||||
cacheReadTokens: 4_940,
|
||||
cacheWriteTokens: 0,
|
||||
outputTokens: 5_800,
|
||||
reasoningTokens: 42,
|
||||
totalTokens: 15_800,
|
||||
routes: [{ provider: 'deepseek', model: 'deepseek-chat' }],
|
||||
}
|
||||
const view = render(<TurnUsageDisclosure usage={usage} t={t} />)
|
||||
|
||||
expect(view.getByText('15.8K tok · Cache hit 49.4%')).toBeTruthy()
|
||||
expect(view.queryByRole('definition')).toBeNull()
|
||||
|
||||
fireEvent.click(view.getByRole('button'))
|
||||
const details = view.container.querySelector('[data-turn-usage-details]') as HTMLElement
|
||||
expect(details).toBeTruthy()
|
||||
expect(details.textContent).toContain('Provider / modeldeepseek/deepseek-chat')
|
||||
expect(details.textContent).toContain('Uncached input5,060 tok')
|
||||
expect(details.textContent).toContain('Cached input4,940 tok')
|
||||
expect(details.textContent).toContain('Cache write0 tok')
|
||||
expect(details.textContent).toContain('Output5,800 tok (42 tok reasoning)')
|
||||
expect(details.textContent).toContain('Total15,800 tok')
|
||||
})
|
||||
|
||||
it('omits unavailable optional facts instead of inventing values', () => {
|
||||
const usage: TurnTokenUsage = {
|
||||
uncachedInputTokens: 120,
|
||||
outputTokens: 30,
|
||||
totalTokens: 150,
|
||||
}
|
||||
const view = render(<TurnUsageDisclosure usage={usage} t={t} />)
|
||||
|
||||
expect(view.getByText('150 tok')).toBeTruthy()
|
||||
expect(view.queryByText(/Cache hit/)).toBeNull()
|
||||
fireEvent.click(view.getByRole('button'))
|
||||
expect(view.queryByText('Provider / model')).toBeNull()
|
||||
expect(view.queryByText('Cached input')).toBeNull()
|
||||
expect(view.queryByText('Cache write')).toBeNull()
|
||||
expect(view.queryByText(/reasoning/)).toBeNull()
|
||||
})
|
||||
|
||||
it('keeps a partial cache hit below 100 and supports keyboard toggling', () => {
|
||||
const usage: TurnTokenUsage = {
|
||||
uncachedInputTokens: 1,
|
||||
cacheReadTokens: 999,
|
||||
outputTokens: 100,
|
||||
totalTokens: 1_100,
|
||||
}
|
||||
const view = render(<TurnUsageDisclosure usage={usage} t={t} />)
|
||||
expect(view.getByText('1.1K tok · Cache hit 99.9%')).toBeTruthy()
|
||||
|
||||
const disclosure = view.getByRole('button')
|
||||
expect(disclosure.getAttribute('aria-expanded')).toBe('false')
|
||||
fireEvent.keyDown(disclosure, { key: ' ' })
|
||||
expect(disclosure.getAttribute('aria-expanded')).toBe('true')
|
||||
fireEvent.keyDown(disclosure, { key: 'Enter' })
|
||||
expect(disclosure.getAttribute('aria-expanded')).toBe('false')
|
||||
})
|
||||
})
|
||||
@@ -5243,7 +5243,7 @@ export const TYPE_API: readonly TypeApiEntry[] = [
|
||||
},
|
||||
{
|
||||
name: 'TokenUsage',
|
||||
declaration: 'export interface TokenUsage {\n inputTokens: number;\n outputTokens: number;\n cacheReadTokens?: number;\n cacheWriteTokens?: number;\n reasoningTokens?: number;\n}',
|
||||
declaration: 'export interface TokenUsage {\n inputTokens: number;\n outputTokens: number;\n totalTokens?: number;\n cacheReadTokens?: number;\n cacheWriteTokens?: number;\n reasoningTokens?: number;\n}',
|
||||
},
|
||||
{
|
||||
name: 'ToolCallKind',
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/llm/llm-deepseek/README.md
|
||||
README.md: 7433bb75104506ec2409c659f3d30058abc6f9a4
|
||||
README.zh.md: 7dcdfeac17b0bfca70a293760061182292edb531
|
||||
README.md: 11ee4c775c6565e0842707928683587a1e2f1eb8
|
||||
README.zh.md: 86da6c75891d7e458b870b630db877c799c33127
|
||||
|
||||
@@ -103,7 +103,7 @@ DeepSeek request identity is separate from app attribution. After credential res
|
||||
- The first thinking-mode chunk carries `reasoning_content: ""` — handled (no spurious reasoning block).
|
||||
- **Reasoning passback rule**: every assistant turn that carried reasoning serializes `reasoning_content` back in history. Thinking mode requires it on tool-call turns; DeepSeek ignores it elsewhere, while a gateway re-encoding the conversation for another vendor recovers that turn's upstream thinking signature by hashing the replayed text.
|
||||
- Image-capable user messages preserve text/image order. Tool-role content remains a string; consecutive tool-result images are grouped into the following user message with `Attached image(s) from tool result:`.
|
||||
- Cache accounting: `cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`; DeepSeek reports no cache-write metric.
|
||||
- Token accounting: `cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`; DeepSeek reports no cache-write metric. `totalTokens` is the exact `prompt_tokens + completion_tokens` aggregate and is omitted if a supplied `total_tokens` disagrees.
|
||||
|
||||
## Errors
|
||||
|
||||
|
||||
@@ -103,7 +103,7 @@ DeepSeek 请求身份独立于应用归因。凭据解析成功后,每个提
|
||||
- 第一个思考模式分片携带 `reasoning_content: ""`,系统会处理它(不会产生多余 reasoning 块)。
|
||||
- **推理回传规则**:每个携带推理内容的 assistant 轮次都会将 `reasoning_content` 序列化回历史。思考模式在工具调用轮次上必需它;DeepSeek 在其他轮次上会忽略它,而将该对话重新编码转发给其他厂商的网关,要靠对回传原文取哈希来恢复该轮次上游的思考签名。
|
||||
- 支持图片的 user 消息会保留文本/图片顺序。Tool role 内容仍为字符串;连续工具结果中的图片会用 `Attached image(s) from tool result:` 汇总到随后一条 user 消息。
|
||||
- Cache 计量:`cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`;DeepSeek 不报告 cache-write 指标。
|
||||
- Token 计量:`cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`;DeepSeek 不报告 cache-write 指标。`totalTokens` 是精确的 `prompt_tokens + completion_tokens` 聚合值;提供的 `total_tokens` 若不一致,则省略该字段。
|
||||
|
||||
## 错误
|
||||
|
||||
|
||||
@@ -48,14 +48,23 @@ export function mapFinishReason(reason: string): FinishReason {
|
||||
* api/create-chat-completion); the harness TokenUsage convention is
|
||||
* DISJOINT counts, so cache reads are subtracted out of `inputTokens`.
|
||||
* @param usage - wire usage from the finish chunk or the trailing usage-only chunk.
|
||||
* @returns disjoint harness counts; cache/reasoning fields present only when the wire reported them.
|
||||
* @returns disjoint harness counts; an exact total is present only when the
|
||||
* aggregate prompt/completion counters are valid and agree with any wire total.
|
||||
*/
|
||||
export function mapUsage(usage: WireUsage): TokenUsage {
|
||||
const cacheRead = usage.prompt_tokens_details?.cached_tokens ?? usage.prompt_cache_hit_tokens
|
||||
const reasoning = usage.completion_tokens_details?.reasoning_tokens
|
||||
const combined = usage.prompt_tokens + usage.completion_tokens
|
||||
const hasExactTotal = Number.isSafeInteger(usage.prompt_tokens)
|
||||
&& usage.prompt_tokens >= 0
|
||||
&& Number.isSafeInteger(usage.completion_tokens)
|
||||
&& usage.completion_tokens >= 0
|
||||
&& Number.isSafeInteger(combined)
|
||||
&& (usage.total_tokens === undefined || usage.total_tokens === combined)
|
||||
return {
|
||||
inputTokens: usage.prompt_tokens - (cacheRead ?? 0),
|
||||
outputTokens: usage.completion_tokens,
|
||||
...hasExactTotal ? { totalTokens: combined } : {},
|
||||
...cacheRead !== undefined ? { cacheReadTokens: cacheRead } : {},
|
||||
...reasoning !== undefined ? { reasoningTokens: reasoning } : {},
|
||||
}
|
||||
|
||||
@@ -166,6 +166,8 @@ export interface WireToolCallDelta {
|
||||
export interface WireUsage {
|
||||
prompt_tokens: number
|
||||
completion_tokens: number
|
||||
/** Provider-reported aggregate across prompt and completion tokens. */
|
||||
total_tokens?: number
|
||||
prompt_cache_hit_tokens?: number
|
||||
prompt_cache_miss_tokens?: number
|
||||
prompt_tokens_details?: { cached_tokens?: number }
|
||||
|
||||
@@ -318,7 +318,7 @@ describe('DeepSeekAdapter against a mock server', () => {
|
||||
})
|
||||
expect(result.message.content).toEqual([{ type: 'text', text: 'hello' }])
|
||||
expect(result.finish).toEqual({ kind: 'stop' })
|
||||
expect(result.usage).toEqual({ inputTokens: 3, outputTokens: 1 })
|
||||
expect(result.usage).toEqual({ inputTokens: 3, outputTokens: 1, totalTokens: 4 })
|
||||
|
||||
// The wire request carried the auth header contents we configured.
|
||||
expect(server.requests[0]).toMatchObject({
|
||||
|
||||
@@ -33,7 +33,7 @@ describe('translate: text', () => {
|
||||
{ type: 'text-delta', index: 0, text: 'Hel' },
|
||||
{ type: 'text-delta', index: 0, text: 'lo' },
|
||||
{ type: 'block-end', index: 0, block: { type: 'text', text: 'Hello' } },
|
||||
{ type: 'usage', usage: { inputTokens: 5, outputTokens: 2 } },
|
||||
{ type: 'usage', usage: { inputTokens: 5, outputTokens: 2, totalTokens: 7 } },
|
||||
{ type: 'finish', reason: { kind: 'stop' } },
|
||||
])
|
||||
})
|
||||
@@ -118,7 +118,7 @@ describe('translate: tool calls', () => {
|
||||
index: 0,
|
||||
block: { type: 'tool-call', id: 'call_00_x', name: 'get_weather', arguments: '{"city": "Paris"}' },
|
||||
},
|
||||
{ type: 'usage', usage: { inputTokens: 28, outputTokens: 6 } },
|
||||
{ type: 'usage', usage: { inputTokens: 28, outputTokens: 6, totalTokens: 34 } },
|
||||
{ type: 'finish', reason: { kind: 'tool-calls' } },
|
||||
])
|
||||
})
|
||||
@@ -172,7 +172,7 @@ describe('translate: finish and usage handling', () => {
|
||||
{ choices: [], usage: { prompt_tokens: 9, completion_tokens: 1 } },
|
||||
DONE,
|
||||
)))
|
||||
expect(chunks.at(-2)).toEqual({ type: 'usage', usage: { inputTokens: 9, outputTokens: 1 } })
|
||||
expect(chunks.at(-2)).toEqual({ type: 'usage', usage: { inputTokens: 9, outputTokens: 1, totalTokens: 10 } })
|
||||
expect(chunks.at(-1)).toEqual({ type: 'finish', reason: { kind: 'stop' } })
|
||||
})
|
||||
|
||||
@@ -184,7 +184,7 @@ describe('translate: finish and usage handling', () => {
|
||||
DONE,
|
||||
)))
|
||||
const usage = chunks.find(chunk => chunk.type === 'usage')
|
||||
expect(usage).toEqual({ type: 'usage', usage: { inputTokens: 2, outputTokens: 2 } })
|
||||
expect(usage).toEqual({ type: 'usage', usage: { inputTokens: 2, outputTokens: 2, totalTokens: 4 } })
|
||||
})
|
||||
|
||||
it('defaults to finish stop when no finish_reason ever arrives', async () => {
|
||||
@@ -219,7 +219,7 @@ describe('translate: finish and usage handling', () => {
|
||||
DONE,
|
||||
)))
|
||||
expect(chunks).toEqual([
|
||||
{ type: 'usage', usage: { inputTokens: 7, outputTokens: 0 } },
|
||||
{ type: 'usage', usage: { inputTokens: 7, outputTokens: 0, totalTokens: 7 } },
|
||||
{
|
||||
type: 'finish',
|
||||
reason: {
|
||||
@@ -286,6 +286,7 @@ describe('mapUsage', () => {
|
||||
expect(mapUsage({
|
||||
prompt_tokens: 283,
|
||||
completion_tokens: 69,
|
||||
total_tokens: 352,
|
||||
prompt_cache_hit_tokens: 256,
|
||||
prompt_cache_miss_tokens: 27,
|
||||
prompt_tokens_details: { cached_tokens: 256 },
|
||||
@@ -295,6 +296,7 @@ describe('mapUsage', () => {
|
||||
// (TokenUsage counts are disjoint).
|
||||
inputTokens: 27,
|
||||
outputTokens: 69,
|
||||
totalTokens: 352,
|
||||
cacheReadTokens: 256,
|
||||
reasoningTokens: 24,
|
||||
})
|
||||
@@ -302,12 +304,26 @@ describe('mapUsage', () => {
|
||||
|
||||
it('falls back to prompt_cache_hit_tokens when details are absent', () => {
|
||||
expect(mapUsage({ prompt_tokens: 10, completion_tokens: 2, prompt_cache_hit_tokens: 8 }))
|
||||
.toEqual({ inputTokens: 2, outputTokens: 2, cacheReadTokens: 8 })
|
||||
.toEqual({ inputTokens: 2, outputTokens: 2, totalTokens: 12, cacheReadTokens: 8 })
|
||||
})
|
||||
|
||||
it('omits optional fields when the wire omits them', () => {
|
||||
it('reconstructs an exact total when the wire omits it', () => {
|
||||
expect(mapUsage({ prompt_tokens: 10, completion_tokens: 2 }))
|
||||
.toEqual({ inputTokens: 10, outputTokens: 2 })
|
||||
.toEqual({ inputTokens: 10, outputTokens: 2, totalTokens: 12 })
|
||||
})
|
||||
|
||||
it.each([
|
||||
['contradictory total', { prompt_tokens: 10, completion_tokens: 2, total_tokens: 99 }],
|
||||
['negative prompt', { prompt_tokens: -1, completion_tokens: 2 }],
|
||||
['fractional prompt', { prompt_tokens: 1.5, completion_tokens: 2 }],
|
||||
['negative completion', { prompt_tokens: 2, completion_tokens: -1 }],
|
||||
['fractional completion', { prompt_tokens: 2, completion_tokens: 1.5 }],
|
||||
['unsafe aggregate', { prompt_tokens: Number.MAX_SAFE_INTEGER, completion_tokens: 1 }],
|
||||
])('omits the exact total for %s without changing existing buckets', (_name, wire) => {
|
||||
expect(mapUsage(wire)).toEqual({
|
||||
inputTokens: wire.prompt_tokens,
|
||||
outputTokens: wire.completion_tokens,
|
||||
})
|
||||
})
|
||||
})
|
||||
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/llm/llm-pi-ai/README.md
|
||||
README.md: 31e40e5f0fa3c1e7e0ae0df05aa0a76d54d120b0
|
||||
README.zh.md: cd40804ce5908aebd0c35011ad1d56879834164d
|
||||
README.md: dc17ec8be163d4c4d2b991afe53fdb15e455b61d
|
||||
README.zh.md: 007c9606cbf921c0f4d490ca7a5ef23713af87b2
|
||||
|
||||
@@ -155,7 +155,7 @@ Durable content is the authoritative record; replay state only restores native f
|
||||
|
||||
- pi-ai tool-call arguments are parsed objects; the harness stores raw JSON strings. The adapter parses input and re-stringifies output.
|
||||
- pi-ai reports failures as in-stream error events; these map to `finish {kind:'error'|'aborted', failure}` chunks. Provider-specific error text distinguishes terminal `QUOTA` from transient `RATE_LIMIT`, while text and usage signals evaluated against the resolved model's context window normalize overflow to `CONTEXT_WINDOW_EXCEEDED`. A terminal `stop` whose message carries no content blocks maps to a `finish {kind:'error'}` with code `EMPTY_RESPONSE` (retried by default policy) instead of a successful empty message.
|
||||
- pi-ai folds reasoning tokens into output usage; there is no separate reasoning count to map.
|
||||
- pi-ai folds reasoning tokens into output usage; there is no separate reasoning count to map. Its exact `totalTokens` value is preserved unchanged.
|
||||
- pi-ai's `off` thinking level crosses the Harness capability seam unchanged and becomes an omitted pi-ai common `reasoning` option at dispatch.
|
||||
- `GenerateOptions.stop` is rejected with `UNSUPPORTED_OPTION` because pi-ai's common streaming UI cannot guarantee it across providers.
|
||||
|
||||
|
||||
@@ -156,7 +156,7 @@ pi-ai 依据提供方 id 与 baseURL 决定每个请求的形状:系统提示
|
||||
|
||||
- pi-ai 工具调用参数是已解析对象;harness 存储原始 JSON 字符串。适配器会解析输入,并将输出重新字符串化。
|
||||
- pi-ai 将失败报告为流内错误事件;它们会映射到 `finish {kind:'error'|'aborted', failure}` 分片。提供方特定错误文本会区分终止型 `QUOTA` 与暂时型 `RATE_LIMIT`,针对已解析模型上下文窗口评估的文本与 usage 信号则将溢出规范化为 `CONTEXT_WINDOW_EXCEEDED`。终止时的 `stop` 若消息不含内容块,则会映射为 `finish {kind:'error'}`,code 为 `EMPTY_RESPONSE`(默认策略会重试),而非成功空消息。
|
||||
- pi-ai 将推理 token 折叠到输出 usage 中;没有可映射的独立推理计数。
|
||||
- pi-ai 将推理 token 折叠到输出 usage 中;没有可映射的独立推理计数。它的精确 `totalTokens` 值会原样保留。
|
||||
- pi-ai 的 `off` 思考级别会原样穿过 Harness 能力 seam,并在分派时变为被省略的 pi-ai 通用 `reasoning` 选项。
|
||||
- `GenerateOptions.stop` 会以 `UNSUPPORTED_OPTION` 被拒绝,因为 pi-ai 的通用流式输出接口无法保证所有提供方都支持它。
|
||||
|
||||
|
||||
@@ -17,12 +17,14 @@ import { toPiReplayState } from './replay.ts'
|
||||
/**
|
||||
* Map pi-ai usage (reasoning folded into output by pi-ai).
|
||||
* @param usage - cumulative usage from the terminal pi-ai event.
|
||||
* @returns harness counts; cache fields appear only when non-zero (pi-ai reports zeros, not absence).
|
||||
* @returns harness counts with pi-ai's exact total; cache fields appear only
|
||||
* when non-zero (pi-ai reports zeros, not absence).
|
||||
*/
|
||||
export function mapUsage(usage: PiUsage): TokenUsage {
|
||||
return {
|
||||
inputTokens: usage.input,
|
||||
outputTokens: usage.output,
|
||||
totalTokens: usage.totalTokens,
|
||||
...usage.cacheRead > 0 ? { cacheReadTokens: usage.cacheRead } : {},
|
||||
...usage.cacheWrite > 0 ? { cacheWriteTokens: usage.cacheWrite } : {},
|
||||
}
|
||||
|
||||
@@ -85,7 +85,7 @@ describe('PiAiAdapter provider routing', () => {
|
||||
})
|
||||
expect(result.message.content).toEqual([{ type: 'text', text: 'hello' }])
|
||||
expect(result.finish).toEqual({ kind: 'stop' })
|
||||
expect(result.usage).toEqual({ inputTokens: 3, outputTokens: 1 })
|
||||
expect(result.usage).toEqual({ inputTokens: 3, outputTokens: 1, totalTokens: 4 })
|
||||
expect(server.paths).toEqual(['/chat/completions'])
|
||||
})
|
||||
|
||||
|
||||
@@ -643,7 +643,7 @@ describe('toStreamChunks', () => {
|
||||
{ type: 'block-start', index: 0, blockType: 'text' },
|
||||
{ type: 'text-delta', index: 0, text: 'hi' },
|
||||
{ type: 'block-end', index: 0, block: { type: 'text', text: 'hi' } },
|
||||
{ type: 'usage', usage: { inputTokens: 3, outputTokens: 2 } },
|
||||
{ type: 'usage', usage: { inputTokens: 3, outputTokens: 2, totalTokens: 5 } },
|
||||
{
|
||||
type: 'finish',
|
||||
reason: { kind: 'stop' },
|
||||
@@ -694,7 +694,7 @@ describe('toStreamChunks', () => {
|
||||
{ type: 'tool-call-delta', index: 0, id: 'call-1', name: 'f', argumentsDelta: '{"a"' },
|
||||
{ type: 'tool-call-delta', index: 0, id: 'call-1', name: 'f', argumentsDelta: ':1}' },
|
||||
{ type: 'block-end', index: 0, block: { type: 'tool-call', id: 'call-1', name: 'f', arguments: '{"a":1}' } },
|
||||
{ type: 'usage', usage: { inputTokens: 0, outputTokens: 0 } },
|
||||
{ type: 'usage', usage: { inputTokens: 0, outputTokens: 0, totalTokens: 0 } },
|
||||
{
|
||||
type: 'finish',
|
||||
reason: { kind: 'tool-calls' },
|
||||
@@ -728,7 +728,7 @@ describe('toStreamChunks', () => {
|
||||
{ type: 'error', reason: 'error', error },
|
||||
)))
|
||||
expect(chunks).toEqual([
|
||||
{ type: 'usage', usage: { inputTokens: 1, outputTokens: 0 } },
|
||||
{ type: 'usage', usage: { inputTokens: 1, outputTokens: 0, totalTokens: 1 } },
|
||||
{ type: 'finish', reason: { kind: 'error', failure: { message: 'boom', code: 'PI_AI_ERROR' } } },
|
||||
])
|
||||
})
|
||||
@@ -890,10 +890,11 @@ describe('mapStopReason / mapUsage', () => {
|
||||
expect(mapUsage(usage(10, 5, 8, 2))).toEqual({
|
||||
inputTokens: 10,
|
||||
outputTokens: 5,
|
||||
totalTokens: 25,
|
||||
cacheReadTokens: 8,
|
||||
cacheWriteTokens: 2,
|
||||
})
|
||||
expect(mapUsage(usage(10, 5))).toEqual({ inputTokens: 10, outputTokens: 5 })
|
||||
expect(mapUsage(usage(10, 5))).toEqual({ inputTokens: 10, outputTokens: 5, totalTokens: 15 })
|
||||
})
|
||||
})
|
||||
|
||||
|
||||
@@ -135,6 +135,14 @@ export type FinishReason = FinishReasonMap[keyof FinishReasonMap]
|
||||
export interface TokenUsage {
|
||||
inputTokens: number
|
||||
outputTokens: number
|
||||
/**
|
||||
* Exact full-call total including aggregate prompt and output tokens.
|
||||
*
|
||||
* Adapters preserve a provider total or derive it from authoritative
|
||||
* aggregate prompt/output counters; they omit it when unavailable or
|
||||
* inconsistent.
|
||||
*/
|
||||
totalTokens?: number
|
||||
cacheReadTokens?: number
|
||||
cacheWriteTokens?: number
|
||||
reasoningTokens?: number
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/llm/token-meter/README.md
|
||||
README.md: 9cc56c0ac5e445f2de63cb71aa0b0e9354ae8492
|
||||
README.zh.md: eb2cfa9b1130c1ff227a284e84ad9afc979cee60
|
||||
README.md: ee80412476c4730e409e6a854d3a78922912bba7
|
||||
README.zh.md: 332cc4df33e3d4da5c786fbaf88af210b02cbc85
|
||||
|
||||
@@ -25,7 +25,9 @@ Usage accounting sums disjoint input, cache-read, cache-write, and output bucket
|
||||
|
||||
When the composition provides `ctx.sessionProjections`, token-meter registers three units through an optional child fiber.
|
||||
|
||||
`tokenUsage` carries the complete durable log's `uncachedInputTokens`, `outputTokens`, `cacheReadTokens`, and `cacheWriteTokens`. Usage chunks are counted even when a request later fails; a final assistant-message usage for the same `(turn, step)` replaces that sample instead of double-counting it. Reasoning remains an output subdivision. The single last-sample slot relies on a session-log ordering property: once a later step reports usage, a legal log never reports usage for an earlier step again.
|
||||
`tokenUsage` carries the complete durable log's `uncachedInputTokens`, `outputTokens`, `cacheReadTokens`, and `cacheWriteTokens`. Usage chunks are counted even when a request later fails; a final assistant-message usage replaces the streaming sample from the same model attempt instead of double-counting it. A matching `llm/retry-started` boundary ends that replacement scope, so a retry with the same `(turn, step)` contributes a new billed attempt. Reasoning remains an output subdivision. The single last-sample slot relies on a session-log ordering property: once a later step reports usage, a legal log never reports usage for an earlier step again.
|
||||
|
||||
Token-meter also owns the browser-safe pure fold from one complete Turn's durable events to exact attempt and Turn usage. `step/start` and `llm/retry-started` open real attempts; final message usage replaces that attempt's streaming sample; terminal failures, retries, and step boundaries close it. Missing lifecycle evidence, unsafe counts, or contradictory exact totals fail closed. Presentation consumers select a complete Turn window and render the result; they do not define a second accounting state machine.
|
||||
|
||||
`contextPressure` carries optional `pressureTokens` — the newest provider-reported prompt size, summing uncached input plus cache reads and writes — optional `projectedTokens`, and optional `contextWindow` from the newest `request/context` record. Both figures stay absent until a provider reports usage; capacity stays absent for a route whose adapter advertises none. Output is excluded, so `pressureTokens` holds still while a turn streams and steps forward when the next request reports its usage.
|
||||
|
||||
|
||||
@@ -25,7 +25,9 @@ fold 跟踪完整请求标头快照、步骤边界、表层追加与替换、成
|
||||
|
||||
当组合提供 `ctx.sessionProjections` 时,token-meter 会通过一个可选子 fiber 注册三个单元。
|
||||
|
||||
`tokenUsage` 携带完整持久日志中的 `uncachedInputTokens`、`outputTokens`、`cacheReadTokens` 和 `cacheWriteTokens`。即使请求随后失败,用量分片仍会计入;同一 `(turn, step)` 的最终 assistant 消息用量会替换该样本,而不是重复计数。推理仍是输出的一个细分项。只保留单个最新样本,依赖的是会话日志的一条顺序性质:一旦某个更晚的步骤报告了用量,合法日志就绝不会再为更早的步骤报告用量。
|
||||
`tokenUsage` 携带完整持久日志中的 `uncachedInputTokens`、`outputTokens`、`cacheReadTokens` 和 `cacheWriteTokens`。即使请求随后失败,用量分片仍会计入;最终 assistant 消息用量会替换同一次模型 attempt 的流式样本,而不是重复计数。匹配的 `llm/retry-started` 边界会结束该替换作用域,因此复用同一 `(turn, step)` 的重试会贡献一次新的计费 attempt。推理仍是输出的一个细分项。只保留单个最新样本,依赖的是会话日志的一条顺序性质:一旦某个更晚的步骤报告了用量,合法日志就绝不会再为更早的步骤报告用量。
|
||||
|
||||
token-meter 还拥有一份可安全用于浏览器的纯 fold,将一个完整 Turn 的持久事件归并为精确的 attempt 与 Turn 用量。`step/start` 与 `llm/retry-started` 打开真实 attempt;最终消息用量替换该 attempt 的流式样本;终止失败、重试与步骤边界关闭它。缺少生命周期证据、计数不安全或精确总量矛盾时一律 fail-closed。展示消费方只选择完整 Turn 窗口并渲染结果,不再定义第二套记账状态机。
|
||||
|
||||
`contextPressure` 携带可选的 `pressureTokens`(提供方报告的最新提示词规模,为未缓存输入加缓存读取与写入之和)、可选的 `projectedTokens`,以及来自最新一条 `request/context` 记录的可选 `contextWindow`。提供方报告用量前两个数字都保持缺失;路由适配器未公布容量时容量也保持缺失。输出不计入其中,因此轮次流式输出期间 `pressureTokens` 保持不动,等到下一个请求报告用量时才前进。
|
||||
|
||||
|
||||
@@ -40,6 +40,7 @@
|
||||
"@deepseek-ai/dsh-compaction": "workspace:^",
|
||||
"@deepseek-ai/dsh-invariants": "workspace:^",
|
||||
"@deepseek-ai/dsh-llm": "workspace:^",
|
||||
"@deepseek-ai/dsh-llm-retry": "workspace:^",
|
||||
"@deepseek-ai/dsh-session": "workspace:^",
|
||||
"@deepseek-ai/dsh-session-projection": "workspace:^",
|
||||
"@deepseek-ai/cordis": "workspace:^"
|
||||
@@ -52,6 +53,7 @@
|
||||
"@deepseek-ai/dsh-compaction": "workspace:^",
|
||||
"@deepseek-ai/dsh-invariants": "workspace:^",
|
||||
"@deepseek-ai/dsh-llm": "workspace:^",
|
||||
"@deepseek-ai/dsh-llm-retry": "workspace:^",
|
||||
"@deepseek-ai/dsh-session": "workspace:^",
|
||||
"@deepseek-ai/dsh-session-projection": "workspace:^",
|
||||
"@deepseek-ai/cordis": "workspace:^"
|
||||
|
||||
@@ -1,7 +1,9 @@
|
||||
/**
|
||||
* Client-namespace projection of token-meter's browser-safe types.
|
||||
* Client-namespace projection of token-meter's browser-safe contracts and folds.
|
||||
*
|
||||
* @module @deepseek-ai/dsh-token-meter/client
|
||||
*/
|
||||
|
||||
export type * from './projection.ts'
|
||||
export { deriveTurnTokenUsage } from './turn-usage.ts'
|
||||
export type { TurnTokenUsage, TurnTokenUsageRoute } from './turn-usage.ts'
|
||||
|
||||
@@ -18,8 +18,8 @@ export const inject = ['invariants']
|
||||
* No runtime invariant: token estimates are per-call outputs and the private
|
||||
* session cache is invalidated at its event mutation boundary. The package's
|
||||
* three projections do expose observation streams, but their schemas fix the
|
||||
* JSON payloads; the usage folds replace same-step samples, so totals need not
|
||||
* be monotone when a final sample corrects an earlier chunk, and the
|
||||
* JSON payloads; the usage folds replace same-attempt samples, so totals need
|
||||
* not be monotone when a final sample corrects an earlier chunk, and the
|
||||
* composition fold prices through the same `estimate.ts` heuristic as the
|
||||
* measurement service and subtracts producer-logged shadow prices derived
|
||||
* from that service's own nodes, which makes its message figure equal
|
||||
|
||||
@@ -0,0 +1,271 @@
|
||||
import type { AssistantMessage, TokenUsage } from '@deepseek-ai/dsh-llm/types'
|
||||
import type {} from '@deepseek-ai/dsh-llm-retry/types'
|
||||
import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
|
||||
|
||||
/** One provider/model route that contributed a billed request attempt. */
|
||||
export interface TurnTokenUsageRoute {
|
||||
readonly provider: string
|
||||
readonly model: string
|
||||
}
|
||||
|
||||
/** Exact provider-reported token accounting for every attempt in one completed Turn. */
|
||||
export interface TurnTokenUsage {
|
||||
/** Sum of uncached prompt input across all attempts. */
|
||||
readonly uncachedInputTokens: number
|
||||
readonly outputTokens: number
|
||||
/** Exact aggregate prompt plus output total across all attempts. */
|
||||
readonly totalTokens: number
|
||||
/** Present only when every attempt reported the bucket. */
|
||||
readonly cacheReadTokens?: number
|
||||
/** Present only when every attempt reported the bucket. */
|
||||
readonly cacheWriteTokens?: number
|
||||
/** Output subset, present only when every attempt reported it. */
|
||||
readonly reasoningTokens?: number
|
||||
/** Present only when every billed attempt has provider/model attribution. */
|
||||
readonly routes?: readonly TurnTokenUsageRoute[]
|
||||
}
|
||||
|
||||
interface NormalizedAttempt {
|
||||
readonly inputTokens: number
|
||||
readonly outputTokens: number
|
||||
readonly totalTokens: number
|
||||
readonly cacheReadTokens?: number
|
||||
readonly cacheWriteTokens?: number
|
||||
readonly reasoningTokens?: number
|
||||
readonly route?: TurnTokenUsageRoute
|
||||
}
|
||||
|
||||
type AttemptState =
|
||||
| { readonly kind: 'idle' }
|
||||
| {
|
||||
readonly kind: 'open'
|
||||
readonly turn: number
|
||||
readonly step: number
|
||||
readonly sample?: TokenUsage
|
||||
}
|
||||
| {
|
||||
readonly kind: 'finishClosed'
|
||||
readonly turn: number
|
||||
readonly step: number
|
||||
}
|
||||
| {
|
||||
readonly kind: 'settled'
|
||||
readonly turn: number
|
||||
readonly step: number
|
||||
readonly by: 'message' | 'retry'
|
||||
}
|
||||
|
||||
function isCount(value: unknown): value is number {
|
||||
return typeof value === 'number' && Number.isSafeInteger(value) && value >= 0
|
||||
}
|
||||
|
||||
function safeSum(values: readonly number[]): number | undefined {
|
||||
let total = 0
|
||||
for (const value of values) {
|
||||
total += value
|
||||
if (!Number.isSafeInteger(total)) return undefined
|
||||
}
|
||||
return total
|
||||
}
|
||||
|
||||
function messageRoute(message: AssistantMessage): TurnTokenUsageRoute | undefined {
|
||||
const { provider, model } = message.source
|
||||
return provider.length > 0 && model.length > 0 ? { provider, model } : undefined
|
||||
}
|
||||
|
||||
function normalizeUsage(usage: TokenUsage, route?: TurnTokenUsageRoute): NormalizedAttempt | undefined {
|
||||
const {
|
||||
inputTokens, outputTokens, cacheReadTokens, cacheWriteTokens, reasoningTokens, totalTokens,
|
||||
} = usage
|
||||
if (!isCount(inputTokens) || !isCount(outputTokens)) return undefined
|
||||
if (cacheReadTokens !== undefined && !isCount(cacheReadTokens)) return undefined
|
||||
if (cacheWriteTokens !== undefined && !isCount(cacheWriteTokens)) return undefined
|
||||
if (reasoningTokens !== undefined && (!isCount(reasoningTokens) || reasoningTokens > outputTokens)) {
|
||||
return undefined
|
||||
}
|
||||
|
||||
const knownPrompt = safeSum([
|
||||
inputTokens,
|
||||
...cacheReadTokens === undefined ? [] : [cacheReadTokens],
|
||||
...cacheWriteTokens === undefined ? [] : [cacheWriteTokens],
|
||||
])
|
||||
if (knownPrompt === undefined) return undefined
|
||||
|
||||
let exactTotal: number
|
||||
if (totalTokens !== undefined) {
|
||||
if (!isCount(totalTokens)) return undefined
|
||||
const exactPrompt = totalTokens - outputTokens
|
||||
if (!isCount(exactPrompt) || exactPrompt < knownPrompt) return undefined
|
||||
if (cacheReadTokens !== undefined && cacheWriteTokens !== undefined && exactPrompt !== knownPrompt) {
|
||||
return undefined
|
||||
}
|
||||
exactTotal = totalTokens
|
||||
} else {
|
||||
if (cacheReadTokens === undefined || cacheWriteTokens === undefined) return undefined
|
||||
const derivedTotal = safeSum([knownPrompt, outputTokens])
|
||||
if (derivedTotal === undefined) return undefined
|
||||
exactTotal = derivedTotal
|
||||
}
|
||||
|
||||
return {
|
||||
inputTokens,
|
||||
outputTokens,
|
||||
totalTokens: exactTotal,
|
||||
...cacheReadTokens === undefined ? {} : { cacheReadTokens },
|
||||
...cacheWriteTokens === undefined ? {} : { cacheWriteTokens },
|
||||
...reasoningTokens === undefined ? {} : { reasoningTokens },
|
||||
...route === undefined ? {} : { route },
|
||||
}
|
||||
}
|
||||
|
||||
function aggregateAttempts(attempts: readonly NormalizedAttempt[]): TurnTokenUsage | undefined {
|
||||
if (attempts.length === 0) return undefined
|
||||
const inputTokens = safeSum(attempts.map(attempt => attempt.inputTokens))
|
||||
const outputTokens = safeSum(attempts.map(attempt => attempt.outputTokens))
|
||||
const totalTokens = safeSum(attempts.map(attempt => attempt.totalTokens))
|
||||
if (inputTokens === undefined || outputTokens === undefined || totalTokens === undefined) return undefined
|
||||
|
||||
const cacheRead = attempts.map(attempt => attempt.cacheReadTokens)
|
||||
const cacheWrite = attempts.map(attempt => attempt.cacheWriteTokens)
|
||||
const reasoning = attempts.map(attempt => attempt.reasoningTokens)
|
||||
const cacheReadTokens = cacheRead.every(isCount) ? safeSum(cacheRead) : undefined
|
||||
const cacheWriteTokens = cacheWrite.every(isCount) ? safeSum(cacheWrite) : undefined
|
||||
const reasoningTokens = reasoning.every(isCount) ? safeSum(reasoning) : undefined
|
||||
// A present cache bucket is bounded by exact prompt, and reasoning is bounded
|
||||
// by output. Safe required aggregates therefore imply safe optional sums.
|
||||
|
||||
let routes: readonly TurnTokenUsageRoute[] | undefined
|
||||
const attributed = attempts.map(attempt => attempt.route)
|
||||
if (attributed.every((route): route is TurnTokenUsageRoute => route !== undefined)) {
|
||||
const unique = new Map<string, TurnTokenUsageRoute>()
|
||||
for (const route of attributed) unique.set(`${route.provider}\0${route.model}`, route)
|
||||
routes = [...unique.values()]
|
||||
}
|
||||
|
||||
return {
|
||||
uncachedInputTokens: inputTokens,
|
||||
outputTokens,
|
||||
totalTokens,
|
||||
...cacheReadTokens === undefined ? {} : { cacheReadTokens },
|
||||
...cacheWriteTokens === undefined ? {} : { cacheWriteTokens },
|
||||
...reasoningTokens === undefined ? {} : { reasoningTokens },
|
||||
...routes === undefined ? {} : { routes },
|
||||
}
|
||||
}
|
||||
|
||||
function sameAttempt(
|
||||
state: Exclude<AttemptState, { kind: 'idle' }>,
|
||||
turn: number,
|
||||
step: number,
|
||||
): boolean {
|
||||
return state.turn === turn && state.step === step
|
||||
}
|
||||
|
||||
/**
|
||||
* Fold one complete Turn's durable attempt lifecycle into exact token accounting.
|
||||
*
|
||||
* No attempt is inferred from a usage sample. Any missing lifecycle boundary,
|
||||
* incomplete attempt usage, unsafe count, or contradictory exact total makes
|
||||
* the whole disclosure unavailable.
|
||||
* @param events - Turn-local durable events from `turn/start` through `turn/end`.
|
||||
* @returns exact aggregate usage, or undefined when it cannot be proven.
|
||||
*/
|
||||
export function deriveTurnTokenUsage(events: readonly SessionEvent[]): TurnTokenUsage | undefined {
|
||||
let state: AttemptState = { kind: 'idle' }
|
||||
const attempts: NormalizedAttempt[] = []
|
||||
let turn: number | undefined
|
||||
let sawEnd = false
|
||||
let invalid = false
|
||||
|
||||
const closeOpen = (route?: TurnTokenUsageRoute): boolean => {
|
||||
if (state.kind !== 'open' || state.sample === undefined) return false
|
||||
const normalized = normalizeUsage(state.sample, route)
|
||||
if (normalized === undefined) return false
|
||||
attempts.push(normalized)
|
||||
return true
|
||||
}
|
||||
|
||||
for (const event of events) {
|
||||
if (invalid) break
|
||||
if (event.type === 'turn/start') {
|
||||
if (turn !== undefined || state.kind !== 'idle') invalid = true
|
||||
else turn = event.data.turn
|
||||
continue
|
||||
}
|
||||
if (turn === undefined) {
|
||||
invalid = true
|
||||
break
|
||||
}
|
||||
if (event.type === 'turn/end') {
|
||||
if (event.data.turn !== turn || state.kind !== 'idle' || sawEnd) invalid = true
|
||||
else sawEnd = true
|
||||
continue
|
||||
}
|
||||
if (sawEnd) {
|
||||
invalid = true
|
||||
break
|
||||
}
|
||||
if (event.type === 'step/start') {
|
||||
if (event.data.turn !== turn || state.kind !== 'idle') invalid = true
|
||||
else state = { kind: 'open', turn, step: event.data.step }
|
||||
continue
|
||||
}
|
||||
if (event.type === 'llm/retry-started') {
|
||||
if (event.data.turn !== turn
|
||||
|| state.kind !== 'settled'
|
||||
|| state.by !== 'retry'
|
||||
|| !sameAttempt(state, event.data.turn, event.data.step)) invalid = true
|
||||
else state = { kind: 'open', turn, step: event.data.step }
|
||||
continue
|
||||
}
|
||||
if (event.type === 'assistant/chunk') {
|
||||
if (event.data.turn !== turn
|
||||
|| state.kind !== 'open'
|
||||
|| !sameAttempt(state, event.data.turn, event.data.step)) {
|
||||
invalid = true
|
||||
continue
|
||||
}
|
||||
if (event.data.chunk.type === 'usage') {
|
||||
state = { ...state, sample: event.data.chunk.usage }
|
||||
} else if (event.data.chunk.type === 'finish'
|
||||
&& (event.data.chunk.reason.kind === 'error' || event.data.chunk.reason.kind === 'aborted')) {
|
||||
if (!closeOpen()) invalid = true
|
||||
else state = { kind: 'finishClosed', turn, step: event.data.step }
|
||||
}
|
||||
continue
|
||||
}
|
||||
if (event.type === 'assistant/message') {
|
||||
if (event.data.turn !== turn
|
||||
|| state.kind !== 'open'
|
||||
|| !sameAttempt(state, event.data.turn, event.data.step)) {
|
||||
invalid = true
|
||||
continue
|
||||
}
|
||||
if (event.data.usage !== undefined) state = { ...state, sample: event.data.usage }
|
||||
if (!closeOpen(messageRoute(event.data.message))) invalid = true
|
||||
else state = { kind: 'settled', turn, step: event.data.step, by: 'message' }
|
||||
continue
|
||||
}
|
||||
if (event.type === 'llm/retry') {
|
||||
if (event.data.turn !== turn || state.kind === 'idle'
|
||||
|| !sameAttempt(state, event.data.turn, event.data.step)) {
|
||||
invalid = true
|
||||
continue
|
||||
}
|
||||
if (state.kind === 'settled' || (state.kind === 'open' && !closeOpen())) invalid = true
|
||||
if (!invalid) state = { kind: 'settled', turn, step: event.data.step, by: 'retry' }
|
||||
continue
|
||||
}
|
||||
if (event.type === 'step/end') {
|
||||
if (event.data.turn !== turn || state.kind === 'idle'
|
||||
|| !sameAttempt(state, event.data.turn, event.data.step)) {
|
||||
invalid = true
|
||||
continue
|
||||
}
|
||||
if (state.kind === 'open' && !closeOpen()) invalid = true
|
||||
if (!invalid) state = { kind: 'idle' }
|
||||
}
|
||||
}
|
||||
|
||||
return invalid || !sawEnd || state.kind !== 'idle' ? undefined : aggregateAttempts(attempts)
|
||||
}
|
||||
@@ -4,6 +4,7 @@
|
||||
|
||||
import { z } from 'zod'
|
||||
import type { TokenUsage } from '@deepseek-ai/dsh-llm'
|
||||
import type {} from '@deepseek-ai/dsh-llm-retry/types'
|
||||
import type { SessionEvent } from '@deepseek-ai/dsh-session'
|
||||
import type { ProjectionDefinition } from '@deepseek-ai/dsh-session-projection'
|
||||
import type { ContextPressureProjection, TokenUsageProjection } from './projection.ts'
|
||||
@@ -110,18 +111,23 @@ type ContextPressureState = z.infer<typeof contextPressureStateSchema>
|
||||
* Token-meter's session projection unit.
|
||||
*
|
||||
* Usage chunks provide an early sample that survives a later request failure;
|
||||
* an assistant message provides the final sample for the same turn/step. A
|
||||
* repeated sample replaces that step's earlier value instead of double
|
||||
* counting it. The single `last` slot relies on the session-log invariant
|
||||
* that usage reports for one turn/step are adjacent: once a later step begins,
|
||||
* a legal log never reports usage for an earlier step again.
|
||||
* an assistant message provides the final sample for the same attempt. A
|
||||
* repeated sample replaces that attempt's earlier value instead of double
|
||||
* counting it, while `llm/retry-started` closes the replacement slot so the
|
||||
* retried attempt adds to the total. The single `last` slot relies on the
|
||||
* session-log invariant that usage reports for one attempt are adjacent.
|
||||
*/
|
||||
export const tokenUsageProjectionDefinition = {
|
||||
key: 'tokenUsage',
|
||||
stateVersion: 1,
|
||||
stateVersion: 2,
|
||||
stateSchema: tokenUsageStateSchema,
|
||||
init: () => ({ totals: zeroBuckets(), last: null }),
|
||||
apply: (state, event) => {
|
||||
if (event.type === 'llm/retry-started') {
|
||||
return state.last?.turn === event.data.turn && state.last.step === event.data.step
|
||||
? { ...state, last: null }
|
||||
: state
|
||||
}
|
||||
let turn: number
|
||||
let step: number
|
||||
let usage: TokenUsage
|
||||
|
||||
@@ -7,6 +7,7 @@ import type { Session } from '@deepseek-ai/dsh-session'
|
||||
import SessionProjectionRegistry from '@deepseek-ai/dsh-session-projection'
|
||||
import TokenMeter from '@deepseek-ai/dsh-token-meter'
|
||||
import type { ContextPressureProjection, TokenUsageProjection } from '@deepseek-ai/dsh-token-meter/client'
|
||||
import { RetryId } from '@deepseek-ai/dsh-llm-retry'
|
||||
import { CompactionId } from '@deepseek-ai/dsh-compaction'
|
||||
import type {} from '../src/usage-projection.ts'
|
||||
|
||||
@@ -94,9 +95,16 @@ function appendSummaryMeter(ctx: Context, session: Session, start: number, end:
|
||||
}
|
||||
|
||||
describe('tokenUsage session projection', () => {
|
||||
it('serves zero buckets for an empty log', async () => {
|
||||
it('serves zero buckets without usage samples', async () => {
|
||||
const { ctx, session } = await harness()
|
||||
expect(projected(ctx, session)).toEqual(ZERO)
|
||||
session.append('llm/retry-started', {
|
||||
retryId: RetryId('token-meter-no-usage-retry'),
|
||||
turn: 1,
|
||||
step: 1,
|
||||
retry: 1,
|
||||
})
|
||||
expect(projected(ctx, session)).toEqual(ZERO)
|
||||
})
|
||||
|
||||
it('does not count a usage chunk and identical final usage twice', async () => {
|
||||
@@ -148,6 +156,58 @@ describe('tokenUsage session projection', () => {
|
||||
})
|
||||
})
|
||||
|
||||
it('accumulates retried attempts while replacing samples within each attempt', async () => {
|
||||
const { ctx, session } = await harness()
|
||||
const retryId = RetryId('token-meter-retry')
|
||||
session.append('turn/start', { turn: 1 })
|
||||
startStep(session, 1, 1)
|
||||
usageChunk(session, {
|
||||
inputTokens: 10,
|
||||
outputTokens: 2,
|
||||
cacheReadTokens: 3,
|
||||
}, 1, 1)
|
||||
session.append('assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: {
|
||||
type: 'finish',
|
||||
reason: { kind: 'error', failure: { code: 'RATE_LIMIT', message: 'busy', status: 429 } },
|
||||
},
|
||||
})
|
||||
session.append('llm/retry', {
|
||||
retryId,
|
||||
turn: 1,
|
||||
step: 1,
|
||||
provider: 'mock',
|
||||
mode: 'normal',
|
||||
policyKey: 'test',
|
||||
retry: 1,
|
||||
maxRetries: 1,
|
||||
delayMs: 0,
|
||||
failure: { code: 'RATE_LIMIT', message: 'busy', status: 429 },
|
||||
})
|
||||
session.append('llm/retry-started', { retryId, turn: 1, step: 1, retry: 1 })
|
||||
const second = usageChunk(session, {
|
||||
inputTokens: 12,
|
||||
outputTokens: 4,
|
||||
cacheReadTokens: 6,
|
||||
}, 1, 1)
|
||||
finalUsage(session, {
|
||||
inputTokens: 14,
|
||||
outputTokens: 5,
|
||||
cacheReadTokens: 8,
|
||||
cacheWriteTokens: 1,
|
||||
}, 1, 1, [second])
|
||||
session.append('turn/end', { turn: 1, reason: { kind: 'completed' } })
|
||||
|
||||
expect(projected(ctx, session)).toEqual({
|
||||
uncachedInputTokens: 24,
|
||||
outputTokens: 7,
|
||||
cacheReadTokens: 11,
|
||||
cacheWriteTokens: 1,
|
||||
})
|
||||
})
|
||||
|
||||
it('accumulates disjoint buckets across steps without adding reasoning twice', async () => {
|
||||
const { ctx, session } = await harness()
|
||||
startStep(session, 1, 1)
|
||||
|
||||
@@ -0,0 +1,397 @@
|
||||
import { describe, expect, it } from 'vitest'
|
||||
import type { TokenUsage } from '@deepseek-ai/dsh-llm'
|
||||
import type { SessionEvent } from '@deepseek-ai/dsh-session'
|
||||
import { deriveTurnTokenUsage } from '../src/turn-usage.ts'
|
||||
|
||||
function event(seq: number, type: string, data: unknown): SessionEvent {
|
||||
return { seq, time: seq, type, data } as unknown as SessionEvent
|
||||
}
|
||||
|
||||
type UsageOverrides = { [Key in keyof TokenUsage]?: TokenUsage[Key] | undefined }
|
||||
|
||||
function usage(overrides: UsageOverrides = {}): TokenUsage {
|
||||
const value = {
|
||||
inputTokens: 100,
|
||||
outputTokens: 20,
|
||||
totalTokens: 170,
|
||||
cacheReadTokens: 50,
|
||||
...overrides,
|
||||
}
|
||||
return Object.fromEntries(Object.entries(value).filter(([, entry]) => entry !== undefined)) as unknown as TokenUsage
|
||||
}
|
||||
|
||||
function message(
|
||||
seq: number,
|
||||
tokenUsage?: TokenUsage,
|
||||
provider = 'deepseek',
|
||||
model = 'deepseek-chat',
|
||||
step = 1,
|
||||
) {
|
||||
return event(seq, 'assistant/message', {
|
||||
turn: 1,
|
||||
step,
|
||||
message: {
|
||||
id: `message-${seq}`,
|
||||
role: 'assistant',
|
||||
content: [{ type: 'text', text: 'done' }],
|
||||
source: { kind: 'model', provider, model },
|
||||
},
|
||||
...tokenUsage === undefined ? {} : { usage: tokenUsage },
|
||||
})
|
||||
}
|
||||
|
||||
function completeAttempt(...middle: readonly SessionEvent[]): SessionEvent[] {
|
||||
return [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
...middle,
|
||||
event(90, 'step/end', { turn: 1, step: 1 }),
|
||||
event(91, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]
|
||||
}
|
||||
|
||||
describe('deriveTurnTokenUsage', () => {
|
||||
it('preserves authoritative totals and explicit optional buckets', () => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, usage({
|
||||
cacheWriteTokens: 0,
|
||||
reasoningTokens: 8,
|
||||
}))))).toEqual({
|
||||
uncachedInputTokens: 100,
|
||||
outputTokens: 20,
|
||||
totalTokens: 170,
|
||||
cacheReadTokens: 50,
|
||||
cacheWriteTokens: 0,
|
||||
reasoningTokens: 8,
|
||||
routes: [{ provider: 'deepseek', model: 'deepseek-chat' }],
|
||||
})
|
||||
})
|
||||
|
||||
it('derives an exact total only when both cache buckets are present', () => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, usage({
|
||||
totalTokens: undefined,
|
||||
inputTokens: 10,
|
||||
outputTokens: 4,
|
||||
cacheReadTokens: 2,
|
||||
cacheWriteTokens: 1,
|
||||
}))))?.totalTokens).toBe(17)
|
||||
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, usage({
|
||||
totalTokens: undefined,
|
||||
cacheWriteTokens: undefined,
|
||||
}))))).toBeUndefined()
|
||||
})
|
||||
|
||||
it('lets final message usage replace the latest streaming sample', () => {
|
||||
const result = deriveTurnTokenUsage(completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
message(4, usage({ inputTokens: 30, outputTokens: 5, totalTokens: 45, cacheReadTokens: 10 })),
|
||||
))
|
||||
expect(result).toMatchObject({ uncachedInputTokens: 30, outputTokens: 5, totalTokens: 45 })
|
||||
})
|
||||
|
||||
it('keeps the latest streaming sample when the final message omits usage', () => {
|
||||
const result = deriveTurnTokenUsage(completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
message(4),
|
||||
))
|
||||
expect(result).toMatchObject({ uncachedInputTokens: 100, outputTokens: 20, totalTokens: 170 })
|
||||
})
|
||||
|
||||
it('counts an error-finished attempt once across its retry boundary', () => {
|
||||
const events = completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'finish', reason: { kind: 'error', failure: { code: 'HTTP', message: 'failed' } } },
|
||||
}),
|
||||
event(5, 'llm/retry', { turn: 1, step: 1 }),
|
||||
event(6, 'llm/retry-started', { turn: 1, step: 1, retry: 1 }),
|
||||
message(7, usage({ inputTokens: 40, outputTokens: 10, totalTokens: 70, cacheReadTokens: 20 })),
|
||||
)
|
||||
expect(deriveTurnTokenUsage(events)).toEqual({
|
||||
uncachedInputTokens: 140,
|
||||
outputTokens: 30,
|
||||
totalTokens: 240,
|
||||
cacheReadTokens: 70,
|
||||
})
|
||||
})
|
||||
|
||||
it('does not invent an attempt for a scheduled retry that never started', () => {
|
||||
const result = deriveTurnTokenUsage(completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'llm/retry', { turn: 1, step: 1 }),
|
||||
))
|
||||
expect(result).toMatchObject({ totalTokens: 170 })
|
||||
})
|
||||
|
||||
it('fails closed for missing lifecycle or missing attempt usage', () => {
|
||||
expect(deriveTurnTokenUsage([
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
message(2, usage()),
|
||||
event(3, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
])).toBeUndefined()
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3)))).toBeUndefined()
|
||||
})
|
||||
|
||||
it.each([
|
||||
['negative', usage({ inputTokens: -1 })],
|
||||
['fractional', usage({ outputTokens: 1.5 })],
|
||||
['unsafe', usage({ totalTokens: Number.MAX_SAFE_INTEGER + 1 })],
|
||||
['invalid cache read', usage({ cacheReadTokens: -1 })],
|
||||
['invalid cache write', usage({ cacheWriteTokens: 1.5 })],
|
||||
['negative exact prompt', usage({ outputTokens: 20, totalTokens: 10, cacheReadTokens: undefined })],
|
||||
['total below known prompt', usage({ totalTokens: 160 })],
|
||||
['contradictory complete buckets', usage({ totalTokens: 171, cacheWriteTokens: 0 })],
|
||||
['reasoning exceeds output', usage({ reasoningTokens: 21 })],
|
||||
['prompt bucket overflow', usage({
|
||||
inputTokens: Number.MAX_SAFE_INTEGER,
|
||||
outputTokens: 0,
|
||||
totalTokens: Number.MAX_SAFE_INTEGER,
|
||||
cacheReadTokens: 1,
|
||||
})],
|
||||
['derived total overflow', usage({
|
||||
inputTokens: Number.MAX_SAFE_INTEGER,
|
||||
outputTokens: 1,
|
||||
totalTokens: undefined,
|
||||
cacheReadTokens: 0,
|
||||
cacheWriteTokens: 0,
|
||||
})],
|
||||
])('fails closed for %s usage', (_label, invalidUsage) => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, invalidUsage)))).toBeUndefined()
|
||||
})
|
||||
|
||||
it('omits optional aggregates and routes unless every attempt reports them', () => {
|
||||
const events = [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, usage({ totalTokens: 175, cacheWriteTokens: 5, reasoningTokens: 2 })),
|
||||
event(4, 'step/end', { turn: 1, step: 1 }),
|
||||
event(5, 'step/start', { turn: 1, step: 2 }),
|
||||
event(6, 'assistant/message', {
|
||||
turn: 1,
|
||||
step: 2,
|
||||
message: {
|
||||
id: 'message-6', role: 'assistant', content: [],
|
||||
source: { kind: 'model', provider: '', model: '' },
|
||||
},
|
||||
usage: usage({ cacheReadTokens: undefined, cacheWriteTokens: undefined, reasoningTokens: undefined }),
|
||||
}),
|
||||
event(7, 'step/end', { turn: 1, step: 2 }),
|
||||
event(8, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]
|
||||
expect(deriveTurnTokenUsage(events)).toEqual({ uncachedInputTokens: 200, outputTokens: 40, totalTokens: 345 })
|
||||
})
|
||||
|
||||
it('sums multiple steps and preserves distinct attributed routes', () => {
|
||||
const events = [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, usage()),
|
||||
event(4, 'step/end', { turn: 1, step: 1 }),
|
||||
event(5, 'step/start', { turn: 1, step: 2 }),
|
||||
message(6, usage(), 'openai', 'gpt-5', 2),
|
||||
event(7, 'step/end', { turn: 1, step: 2 }),
|
||||
event(8, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]
|
||||
expect(deriveTurnTokenUsage(events)).toEqual({
|
||||
uncachedInputTokens: 200,
|
||||
outputTokens: 40,
|
||||
totalTokens: 340,
|
||||
cacheReadTokens: 100,
|
||||
routes: [
|
||||
{ provider: 'deepseek', model: 'deepseek-chat' },
|
||||
{ provider: 'openai', model: 'gpt-5' },
|
||||
],
|
||||
})
|
||||
})
|
||||
|
||||
it('fails closed when aggregation overflows a safe integer', () => {
|
||||
const half = Math.floor(Number.MAX_SAFE_INTEGER / 2) + 1
|
||||
const attempt = usage({ inputTokens: 0, outputTokens: 0, cacheReadTokens: undefined, totalTokens: half })
|
||||
const events = [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, attempt),
|
||||
event(4, 'step/end', { turn: 1, step: 1 }),
|
||||
event(5, 'step/start', { turn: 1, step: 2 }),
|
||||
event(6, 'assistant/message', {
|
||||
turn: 1,
|
||||
step: 2,
|
||||
message: {
|
||||
id: 'message-6', role: 'assistant', content: [],
|
||||
source: { kind: 'model', provider: 'deepseek', model: 'deepseek-chat' },
|
||||
},
|
||||
usage: attempt,
|
||||
}),
|
||||
event(7, 'step/end', { turn: 1, step: 2 }),
|
||||
event(8, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]
|
||||
expect(deriveTurnTokenUsage(events)).toBeUndefined()
|
||||
})
|
||||
|
||||
it.each([
|
||||
['uncached input', usage({
|
||||
inputTokens: Math.floor(Number.MAX_SAFE_INTEGER / 2) + 1,
|
||||
outputTokens: 0,
|
||||
cacheReadTokens: undefined,
|
||||
totalTokens: Math.floor(Number.MAX_SAFE_INTEGER / 2) + 1,
|
||||
})],
|
||||
['output', usage({
|
||||
inputTokens: 0,
|
||||
outputTokens: Math.floor(Number.MAX_SAFE_INTEGER / 2) + 1,
|
||||
cacheReadTokens: undefined,
|
||||
totalTokens: Math.floor(Number.MAX_SAFE_INTEGER / 2) + 1,
|
||||
})],
|
||||
])('fails closed when aggregate %s overflows', (_label, attempt) => {
|
||||
expect(deriveTurnTokenUsage([
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, attempt),
|
||||
event(4, 'step/end', { turn: 1, step: 1 }),
|
||||
event(5, 'step/start', { turn: 1, step: 1 }),
|
||||
message(6, attempt),
|
||||
event(7, 'step/end', { turn: 1, step: 1 }),
|
||||
event(8, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
])).toBeUndefined()
|
||||
})
|
||||
|
||||
it('closes a sampled attempt at step/end', () => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'finish', reason: { kind: 'stop' } },
|
||||
}),
|
||||
event(5, 'tool/call', { turn: 1, step: 1 }),
|
||||
))).toMatchObject({ totalTokens: 170 })
|
||||
})
|
||||
|
||||
it('accepts an aborted finish after observing usage', () => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'finish', reason: { kind: 'aborted' } },
|
||||
}),
|
||||
))).toMatchObject({ totalTokens: 170 })
|
||||
})
|
||||
|
||||
it.each([
|
||||
['empty turn', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]],
|
||||
['duplicate turn start', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'turn/start', { turn: 1 }),
|
||||
]],
|
||||
['wrong turn end', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'turn/end', { turn: 2, reason: { kind: 'completed' } }),
|
||||
]],
|
||||
['turn end during an open attempt', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]],
|
||||
['duplicate turn end', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
event(3, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
]],
|
||||
['event after turn end', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'turn/end', { turn: 1, reason: { kind: 'completed' } }),
|
||||
event(3, 'step/start', { turn: 1, step: 1 }),
|
||||
]],
|
||||
['wrong-turn step start', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 2, step: 1 }),
|
||||
]],
|
||||
['nested step start', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'step/start', { turn: 1, step: 2 }),
|
||||
]],
|
||||
['retry start without a scheduled retry', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'llm/retry-started', { turn: 1, step: 1, retry: 1 }),
|
||||
]],
|
||||
['retry start after a final message', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, usage()),
|
||||
event(4, 'llm/retry-started', { turn: 1, step: 1, retry: 1 }),
|
||||
]],
|
||||
['retry start for the wrong step', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'llm/retry', { turn: 1, step: 1 }),
|
||||
event(5, 'llm/retry-started', { turn: 1, step: 2, retry: 1 }),
|
||||
]],
|
||||
['usage chunk outside an attempt', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
]],
|
||||
['usage chunk for the wrong step', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 2, chunk: { type: 'usage', usage: usage() } }),
|
||||
]],
|
||||
['error finish without usage', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'assistant/chunk', {
|
||||
turn: 1,
|
||||
step: 1,
|
||||
chunk: { type: 'finish', reason: { kind: 'error', failure: { code: 'HTTP', message: 'failed' } } },
|
||||
}),
|
||||
]],
|
||||
['retry outside an attempt', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'llm/retry', { turn: 1, step: 1 }),
|
||||
]],
|
||||
['retry for the wrong step', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'assistant/chunk', { turn: 1, step: 1, chunk: { type: 'usage', usage: usage() } }),
|
||||
event(4, 'llm/retry', { turn: 1, step: 2 }),
|
||||
]],
|
||||
['retry after a final message', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
message(3, usage()),
|
||||
event(4, 'llm/retry', { turn: 1, step: 1 }),
|
||||
]],
|
||||
['retry before any usage', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'llm/retry', { turn: 1, step: 1 }),
|
||||
]],
|
||||
['step end outside an attempt', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/end', { turn: 1, step: 1 }),
|
||||
]],
|
||||
['step end for the wrong step', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'step/end', { turn: 1, step: 2 }),
|
||||
]],
|
||||
['step end before any usage', [
|
||||
event(1, 'turn/start', { turn: 1 }),
|
||||
event(2, 'step/start', { turn: 1, step: 1 }),
|
||||
event(3, 'step/end', { turn: 1, step: 1 }),
|
||||
]],
|
||||
])('fails closed for invalid lifecycle: %s', (_label, events) => {
|
||||
expect(deriveTurnTokenUsage(events)).toBeUndefined()
|
||||
})
|
||||
|
||||
it('requires the complete turn window', () => {
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, usage())).slice(1))).toBeUndefined()
|
||||
expect(deriveTurnTokenUsage(completeAttempt(message(3, usage())).slice(0, -1))).toBeUndefined()
|
||||
})
|
||||
})
|
||||
@@ -20,6 +20,9 @@
|
||||
{
|
||||
"path": "../../llm/llm"
|
||||
},
|
||||
{
|
||||
"path": "../../llm/llm-retry"
|
||||
},
|
||||
{
|
||||
"path": "../../core/session"
|
||||
},
|
||||
|
||||
@@ -35,10 +35,14 @@ class CliMockAdapter extends LlmAdapter {
|
||||
}
|
||||
const toolResult = options.messages.at(-1)?.content.find(block => block.type === 'tool-result')
|
||||
if (toolResult === undefined) {
|
||||
const reasoning = 'Inspecting the task before the tool call.'
|
||||
const args = JSON.stringify({ command: 'printf CLI_TOOL_ROUND_TRIP', description: 'Prove the CLI tool round trip.' })
|
||||
yield { type: 'block-start', index: 0, blockType: 'tool-call' }
|
||||
yield { type: 'tool-call-delta', index: 0, id: CallId('cli-smoke-call'), name: 'bash', argumentsDelta: args }
|
||||
yield { type: 'block-end', index: 0, block: { type: 'tool-call', id: CallId('cli-smoke-call'), name: 'bash', arguments: args } }
|
||||
yield { type: 'block-start', index: 0, blockType: 'reasoning' }
|
||||
yield { type: 'reasoning-delta', index: 0, text: reasoning }
|
||||
yield { type: 'block-end', index: 0, block: { type: 'reasoning', text: reasoning } }
|
||||
yield { type: 'block-start', index: 1, blockType: 'tool-call' }
|
||||
yield { type: 'tool-call-delta', index: 1, id: CallId('cli-smoke-call'), name: 'bash', argumentsDelta: args }
|
||||
yield { type: 'block-end', index: 1, block: { type: 'tool-call', id: CallId('cli-smoke-call'), name: 'bash', arguments: args } }
|
||||
yield { type: 'usage', usage: { inputTokens: 11, outputTokens: 3, cacheReadTokens: 2 } }
|
||||
yield { type: 'finish', reason: { kind: 'tool-calls' } }
|
||||
return
|
||||
|
||||
Generated
+3
@@ -6281,6 +6281,9 @@ importers:
|
||||
'@deepseek-ai/dsh-llm':
|
||||
specifier: workspace:^
|
||||
version: link:../llm
|
||||
'@deepseek-ai/dsh-llm-retry':
|
||||
specifier: workspace:^
|
||||
version: link:../llm-retry
|
||||
'@deepseek-ai/dsh-session':
|
||||
specifier: workspace:^
|
||||
version: link:../../core/session
|
||||
|
||||
@@ -95,6 +95,9 @@ describe('client bundle purity gate', () => {
|
||||
expect(resolveId('@deepseek-ai/dsh-host-apiproxy/api')).toBeNull()
|
||||
expect(resolveId('@deepseek-ai/dsh-session/surface')).toBeNull()
|
||||
expect(resolveId('@deepseek-ai/dsh-brand')).toBeNull()
|
||||
expect(resolveId('@deepseek-ai/dsh-token-meter/client')).toBeNull()
|
||||
expect(() => resolveId('@deepseek-ai/dsh-token-meter')).toThrow(/purity/)
|
||||
expect(() => resolveId('@deepseek-ai/dsh-token-meter/client/internal')).toThrow(/purity/)
|
||||
})
|
||||
|
||||
it('lets exact generated Remote contributions inline without admitting their package implementation', () => {
|
||||
|
||||
@@ -233,7 +233,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -279,7 +280,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -428,7 +430,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -474,7 +477,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -661,7 +665,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -707,7 +712,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -887,7 +893,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -933,7 +940,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -1078,7 +1086,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1124,7 +1133,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -1309,7 +1319,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1355,7 +1366,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -1532,7 +1544,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1576,7 +1589,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -1930,7 +1944,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1988,7 +2003,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -2191,7 +2207,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2249,7 +2266,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -2496,7 +2514,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2554,7 +2573,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -2800,7 +2820,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2858,7 +2879,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -3248,7 +3270,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -3304,7 +3327,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -3541,7 +3565,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -3599,7 +3624,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -4021,7 +4047,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -4077,7 +4104,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -4345,7 +4373,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -4403,7 +4432,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
@@ -4640,7 +4670,8 @@
|
||||
"type": "usage",
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -4696,7 +4727,8 @@
|
||||
},
|
||||
"usage": {
|
||||
"inputTokens": 3,
|
||||
"outputTokens": 3
|
||||
"outputTokens": 3,
|
||||
"totalTokens": 6
|
||||
}
|
||||
},
|
||||
"sourceEventSeqs": [
|
||||
|
||||
@@ -16,8 +16,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"text-delta","index":0,"text":"DIRECT_CHILD_OK"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"DIRECT_CHILD_OK"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"DIRECT_CHILD_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[14,15,16,17,18],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"DIRECT_CHILD_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[14,15,16,17,18],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -16,8 +16,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"text-delta","index":0,"text":"WORKFLOW_CHILD_OK"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"WORKFLOW_CHILD_OK"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"WORKFLOW_CHILD_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[14,15,16,17,18],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"WORKFLOW_CHILD_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[14,15,16,17,18],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -15,9 +15,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-define","name":"cordis_define","argumentsDelta":"{\"plugin\": {\"kind\": \"new\", \"idPrefix\": \"snap\"}, \"name\": \"Snapshot Double\", \"purpose\": \"Expose a deterministic doubling tool for executable snapshot verification.\", \"code\": {\"host\": \"return (ctx) => {\\n harness.registerTool(ctx, harness.defineTool({\\n name: 'snapshot_double',\\n description: 'Double a number for executable snapshot verification.',\\n parameters: { value: { type: 'number', required: true } },\\n output: {\\n schema: { type: 'number' },\\n render(_args, value) {\\n return [{ type: 'text', text: String(value) }]\\n }\\n },\\n async execute(args) {\\n return args.value * 2\\n }\\n }))\\n}\\n\"}}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-define","name":"cordis_define","arguments":"{\"plugin\": {\"kind\": \"new\", \"idPrefix\": \"snap\"}, \"name\": \"Snapshot Double\", \"purpose\": \"Expose a deterministic doubling tool for executable snapshot verification.\", \"code\": {\"host\": \"return (ctx) => {\\n harness.registerTool(ctx, harness.defineTool({\\n name: 'snapshot_double',\\n description: 'Double a number for executable snapshot verification.',\\n parameters: { value: { type: 'number', required: true } },\\n output: {\\n schema: { type: 'number' },\\n render(_args, value) {\\n return [{ type: 'text', text: String(value) }]\\n }\\n },\\n async execute(args) {\\n return args.value * 2\\n }\\n }))\\n}\\n\"}}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-define","name":"cordis_define","arguments":"{\"plugin\": {\"kind\": \"new\", \"idPrefix\": \"snap\"}, \"name\": \"Snapshot Double\", \"purpose\": \"Expose a deterministic doubling tool for executable snapshot verification.\", \"code\": {\"host\": \"return (ctx) => {\\n harness.registerTool(ctx, harness.defineTool({\\n name: 'snapshot_double',\\n description: 'Double a number for executable snapshot verification.',\\n parameters: { value: { type: 'number', required: true } },\\n output: {\\n schema: { type: 'number' },\\n render(_args, value) {\\n return [{ type: 'text', text: String(value) }]\\n }\\n },\\n async execute(args) {\\n return args.value * 2\\n }\\n }))\\n}\\n\"}}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-define","name":"cordis_define","arguments":"{\"plugin\": {\"kind\": \"new\", \"idPrefix\": \"snap\"}, \"name\": \"Snapshot Double\", \"purpose\": \"Expose a deterministic doubling tool for executable snapshot verification.\", \"code\": {\"host\": \"return (ctx) => {\\n harness.registerTool(ctx, harness.defineTool({\\n name: 'snapshot_double',\\n description: 'Double a number for executable snapshot verification.',\\n parameters: { value: { type: 'number', required: true } },\\n output: {\\n schema: { type: 'number' },\\n render(_args, value) {\\n return [{ type: 'text', text: String(value) }]\\n }\\n },\\n async execute(args) {\\n return args.value * 2\\n }\\n }))\\n}\\n\"}}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":1,"callId":"advanced-define","name":"cordis_define","arguments":"{\"plugin\": {\"kind\": \"new\", \"idPrefix\": \"snap\"}, \"name\": \"Snapshot Double\", \"purpose\": \"Expose a deterministic doubling tool for executable snapshot verification.\", \"code\": {\"host\": \"return (ctx) => {\\n harness.registerTool(ctx, harness.defineTool({\\n name: 'snapshot_double',\\n description: 'Double a number for executable snapshot verification.',\\n parameters: { value: { type: 'number', required: true } },\\n output: {\\n schema: { type: 'number' },\\n render(_args, value) {\\n return [{ type: 'text', text: String(value) }]\\n }\\n },\\n async execute(args) {\\n return args.value * 2\\n }\\n }))\\n}\\n\"}}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"advanced-define"},"content":[{"type":"tool-result","toolCallId":"advanced-define","content":[{"type":"text","text":"Defined snap-1/pkg-1 (Snapshot Double); it is not running yet. Use cordis_run to activate this Package."}],"isError":false}],"role":"user","id":"{{messageId}}"},"meta":{"pluginId":"snap-1","packageId":"pkg-1"}},"sourceEventSeqs":[19],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
@@ -26,9 +26,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-run","name":"cordis_run","argumentsDelta":"{\"pluginId\": \"snap-1\", \"packageId\": \"pkg-1\", \"mode\": \"run\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-run","name":"cordis_run","arguments":"{\"pluginId\": \"snap-1\", \"packageId\": \"pkg-1\", \"mode\": \"run\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-run","name":"cordis_run","arguments":"{\"pluginId\": \"snap-1\", \"packageId\": \"pkg-1\", \"mode\": \"run\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[24,25,26,27,28],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-run","name":"cordis_run","arguments":"{\"pluginId\": \"snap-1\", \"packageId\": \"pkg-1\", \"mode\": \"run\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[24,25,26,27,28],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":2,"callId":"advanced-run","name":"cordis_run","arguments":"{\"pluginId\": \"snap-1\", \"packageId\": \"pkg-1\", \"mode\": \"run\"}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":2,"message":{"source":{"kind":"tool","callId":"advanced-run"},"content":[{"type":"tool-result","toolCallId":"advanced-run","content":[{"type":"text","text":"snap-1/pkg-1 is running (run-1)."}],"isError":false}],"role":"user","id":"{{messageId}}"},"meta":{"pluginId":"snap-1","packageId":"pkg-1","pluginRunId":"run-1"}},"sourceEventSeqs":[30],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":2}}
|
||||
@@ -38,9 +38,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-code","name":"run_code","argumentsDelta":"{\"code\": \"return await tools.snapshot_double({ value: 21 })\", \"description\": \"Run the temporary Plugin tool\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-code","name":"run_code","arguments":"{\"code\": \"return await tools.snapshot_double({ value: 21 })\", \"description\": \"Run the temporary Plugin tool\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":3,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":3,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-code","name":"run_code","arguments":"{\"code\": \"return await tools.snapshot_double({ value: 21 })\", \"description\": \"Run the temporary Plugin tool\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[36,37,38,39,40],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":3,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-code","name":"run_code","arguments":"{\"code\": \"return await tools.snapshot_double({ value: 21 })\", \"description\": \"Run the temporary Plugin tool\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[36,37,38,39,40],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":3,"callId":"advanced-code","name":"run_code","arguments":"{\"code\": \"return await tools.snapshot_double({ value: 21 })\", \"description\": \"Run the temporary Plugin tool\"}"}}
|
||||
{"type":"tool/code-dispatch-start","data":{"rootCallId":"advanced-code","parentCallId":"advanced-code","subCallId":"advanced-code:code:1","name":"snapshot_double","arguments":{"value":21}}}
|
||||
{"type":"tool/code-dispatch","data":{"rootCallId":"advanced-code","parentCallId":"advanced-code","subCallId":"advanced-code:code:1","name":"snapshot_double","arguments":{"value":21},"isError":false,"content":[{"type":"text","text":"42"}]}}
|
||||
@@ -51,9 +51,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-direct-child","name":"subagent","argumentsDelta":"{\"description\": \"Check direct child\", \"prompt\": \"Reply with exactly DIRECT_CHILD_OK and nothing else.\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-direct-child","name":"subagent","arguments":"{\"description\": \"Check direct child\", \"prompt\": \"Reply with exactly DIRECT_CHILD_OK and nothing else.\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":4,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":4,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-direct-child","name":"subagent","arguments":"{\"description\": \"Check direct child\", \"prompt\": \"Reply with exactly DIRECT_CHILD_OK and nothing else.\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[49,50,51,52,53],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":4,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-direct-child","name":"subagent","arguments":"{\"description\": \"Check direct child\", \"prompt\": \"Reply with exactly DIRECT_CHILD_OK and nothing else.\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[49,50,51,52,53],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":4,"callId":"advanced-direct-child","name":"subagent","arguments":"{\"description\": \"Check direct child\", \"prompt\": \"Reply with exactly DIRECT_CHILD_OK and nothing else.\"}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":4,"message":{"source":{"kind":"tool","callId":"advanced-direct-child"},"content":[{"type":"tool-result","toolCallId":"advanced-direct-child","content":[{"type":"text","text":"DIRECT_CHILD_OK"}],"isError":false}],"role":"user","id":"{{messageId}}"}},"sourceEventSeqs":[55],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":4}}
|
||||
@@ -62,9 +62,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-workflow","name":"workflow","argumentsDelta":"{\"script\": \"phase('Delegate')\\nconst reply = await agent('Reply with exactly WORKFLOW_CHILD_OK and nothing else.', { label: 'workflow-child' })\\nreturn { reply }\", \"meta\": {\"name\": \"advanced-exe-snapshot\", \"description\": \"exercise one packaged workflow child\"}}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-workflow","name":"workflow","arguments":"{\"script\": \"phase('Delegate')\\nconst reply = await agent('Reply with exactly WORKFLOW_CHILD_OK and nothing else.', { label: 'workflow-child' })\\nreturn { reply }\", \"meta\": {\"name\": \"advanced-exe-snapshot\", \"description\": \"exercise one packaged workflow child\"}}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":5,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":5,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-workflow","name":"workflow","arguments":"{\"script\": \"phase('Delegate')\\nconst reply = await agent('Reply with exactly WORKFLOW_CHILD_OK and nothing else.', { label: 'workflow-child' })\\nreturn { reply }\", \"meta\": {\"name\": \"advanced-exe-snapshot\", \"description\": \"exercise one packaged workflow child\"}}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[60,61,62,63,64],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":5,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-workflow","name":"workflow","arguments":"{\"script\": \"phase('Delegate')\\nconst reply = await agent('Reply with exactly WORKFLOW_CHILD_OK and nothing else.', { label: 'workflow-child' })\\nreturn { reply }\", \"meta\": {\"name\": \"advanced-exe-snapshot\", \"description\": \"exercise one packaged workflow child\"}}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[60,61,62,63,64],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":5,"callId":"advanced-workflow","name":"workflow","arguments":"{\"script\": \"phase('Delegate')\\nconst reply = await agent('Reply with exactly WORKFLOW_CHILD_OK and nothing else.', { label: 'workflow-child' })\\nreturn { reply }\", \"meta\": {\"name\": \"advanced-exe-snapshot\", \"description\": \"exercise one packaged workflow child\"}}"}}
|
||||
{"type":"tool-workflow/run-start","data":{"runId":"{{workflow-run}}","name":"advanced-exe-snapshot"}}
|
||||
{"type":"tool-workflow/agent-start","data":{"runId":"{{workflow-run}}","seq":1,"label":"workflow-child","phase":"Delegate","childId":"{{child-2}}"}}
|
||||
@@ -77,9 +77,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"tool-call-delta","index":0,"id":"advanced-undefine","name":"cordis_undefine","argumentsDelta":"{\"pluginId\": \"snap-1\"}"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"advanced-undefine","name":"cordis_undefine","arguments":"{\"pluginId\": \"snap-1\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":6,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":6,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-undefine","name":"cordis_undefine","arguments":"{\"pluginId\": \"snap-1\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[75,76,77,78,79],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":6,"message":{"role":"assistant","content":[{"type":"tool-call","id":"advanced-undefine","name":"cordis_undefine","arguments":"{\"pluginId\": \"snap-1\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[75,76,77,78,79],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":6,"callId":"advanced-undefine","name":"cordis_undefine","arguments":"{\"pluginId\": \"snap-1\"}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":6,"message":{"source":{"kind":"tool","callId":"advanced-undefine"},"content":[{"type":"tool-result","toolCallId":"advanced-undefine","content":[{"type":"text","text":"Removed dynamic Plugin snap-1 and all of its Packages."}],"isError":false}],"role":"user","id":"{{messageId}}"}},"sourceEventSeqs":[81],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":6}}
|
||||
@@ -89,8 +89,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"text-delta","index":0,"text":"ADVANCED_EXECUTABLE_OK"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"ADVANCED_EXECUTABLE_OK"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":7,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":7,"message":{"role":"assistant","content":[{"type":"text","text":"ADVANCED_EXECUTABLE_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[87,88,89,90,91],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":7,"message":{"role":"assistant","content":[{"type":"text","text":"ADVANCED_EXECUTABLE_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[87,88,89,90,91],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":7}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -15,8 +15,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"text-delta","index":0,"text":"PROCESS_ONE_OK"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"PROCESS_ONE_OK"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"PROCESS_ONE_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"PROCESS_ONE_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -15,8 +15,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"text-delta","index":0,"text":"PROCESS_TWO_OK"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"PROCESS_TWO_OK"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"PROCESS_TWO_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"text","text":"PROCESS_TWO_OK"}],"source":{"kind":"model","provider":"deepseek-official","model":"smoke-model"},"id":"{{messageId}}"},"usage":{"inputTokens":3,"outputTokens":3,"totalTokens":6}},"sourceEventSeqs":[13,14,15,16,17],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -258,13 +258,76 @@ function turnReasonFromSession(log: string): JsonObject | undefined {
|
||||
}
|
||||
|
||||
function stderrFromSession(log: string): string {
|
||||
let output = ''
|
||||
let started = false
|
||||
let open = false
|
||||
let endsWithNewline = true
|
||||
const appendReasoning = (text: string): void => {
|
||||
if (text === '') return
|
||||
if (!open) {
|
||||
output += 'dsh: reasoning:\n'
|
||||
open = true
|
||||
}
|
||||
output += text
|
||||
endsWithNewline = text.endsWith('\n')
|
||||
}
|
||||
const close = (): void => {
|
||||
if (!open) return
|
||||
if (!endsWithNewline) output += '\n'
|
||||
open = false
|
||||
endsWithNewline = true
|
||||
}
|
||||
for (const record of records(log)) {
|
||||
if (record.type === 'turn/start') {
|
||||
close()
|
||||
started = true
|
||||
continue
|
||||
}
|
||||
if (!started) continue
|
||||
const data = record.data as JsonObject | undefined
|
||||
if (record.type === 'reasoning-chunks') {
|
||||
if (!Array.isArray(data?.texts) || data.texts.some(text => typeof text !== 'string')) {
|
||||
throw new Error('headless snapshot reasoning chunks have invalid text')
|
||||
}
|
||||
for (const text of data.texts as string[]) appendReasoning(text)
|
||||
continue
|
||||
}
|
||||
if (record.type === 'text-chunks' || record.type === 'tool-call-chunks') {
|
||||
close()
|
||||
continue
|
||||
}
|
||||
if (record.type !== 'assistant/chunk') continue
|
||||
const chunk = data?.chunk as JsonObject | undefined
|
||||
switch (chunk?.type) {
|
||||
case 'reasoning-delta':
|
||||
if (typeof chunk.text !== 'string') throw new Error('headless snapshot reasoning delta has invalid text')
|
||||
appendReasoning(chunk.text)
|
||||
break
|
||||
case 'block-start':
|
||||
if (chunk.blockType !== 'reasoning') close()
|
||||
break
|
||||
case 'block-end': {
|
||||
const block = chunk.block as JsonObject | undefined
|
||||
if (block?.type !== 'reasoning') close()
|
||||
break
|
||||
}
|
||||
case 'usage':
|
||||
break
|
||||
case 'text-delta':
|
||||
case 'tool-call-delta':
|
||||
case 'finish':
|
||||
close()
|
||||
break
|
||||
}
|
||||
}
|
||||
close()
|
||||
const reason = turnReasonFromSession(log)
|
||||
if (reason?.kind !== 'error') return ''
|
||||
if (reason?.kind !== 'error') return output
|
||||
const error = reason.error as JsonObject | undefined
|
||||
if (typeof error?.code !== 'string' || typeof error.message !== 'string') {
|
||||
throw new Error('headless snapshot error reason has no code and message')
|
||||
}
|
||||
return `dsh: ${error.code}: ${error.message}\n`
|
||||
return `${output}dsh: ${error.code}: ${error.message}\n`
|
||||
}
|
||||
|
||||
function modelFromSession(log: string): { provider: string; model: string } {
|
||||
@@ -488,6 +551,28 @@ describe('headless recorded-session snapshots', () => {
|
||||
expect(logical(packed)).toStrictEqual(logical(source))
|
||||
})
|
||||
|
||||
it('reconstructs reasoning stderr across packed output boundaries', () => {
|
||||
const log = [
|
||||
{ type: 'turn/start', data: { turn: 1 } },
|
||||
{ type: 'reasoning-chunks', data: { texts: ['first', ''] } },
|
||||
{ type: 'text-chunks', data: { texts: ['text'] } },
|
||||
{ type: 'reasoning-chunks', data: { texts: ['second'] } },
|
||||
{ type: 'tool-call-chunks', data: { args: ['{}'] } },
|
||||
{ type: 'reasoning-chunks', data: { texts: ['third\n'] } },
|
||||
{ type: 'turn/end', data: { turn: 1, reason: { kind: 'completed' } } },
|
||||
].map(record => JSON.stringify(record)).join('\n')
|
||||
|
||||
expect(stderrFromSession(log)).toBe([
|
||||
'dsh: reasoning:',
|
||||
'first',
|
||||
'dsh: reasoning:',
|
||||
'second',
|
||||
'dsh: reasoning:',
|
||||
'third',
|
||||
'',
|
||||
].join('\n'))
|
||||
})
|
||||
|
||||
for (const scenario of scenarios) {
|
||||
const skipped = scenario.manifest.platform === 'posix' && process.platform === 'win32'
|
||||
|| scenario.manifest.platform === 'pwsh' && !hasPwsh
|
||||
@@ -587,12 +672,16 @@ describe('headless recorded-session snapshots', () => {
|
||||
await rm(spillRoot, { recursive: true, force: true })
|
||||
}
|
||||
|
||||
const stderrLog = mode === 'replay' ? primaryFixture : actualLogs[0]?.content
|
||||
if (stderrLog === undefined) throw new Error(`${scenario.name}: stderr projection has no primary session`)
|
||||
const expectedStderr = stderrFromSession(stderrLog)
|
||||
|
||||
if (mode !== 'replay') {
|
||||
fixtures = await writeSessionFixtures(scenario, actualLogs, fixtures, contextOf(actualLogs.map(log => log.content)))
|
||||
}
|
||||
|
||||
expect(result.stdout).toBe(`${finalTextFromSession(fixtures[0] as string)}\n`)
|
||||
expect(result.stderr).toBe(stderrFromSession(fixtures[0] as string))
|
||||
expect(result.stderr).toBe(expectedStderr)
|
||||
expect(actualLogs, `${scenario.name}: persisted session count`).toHaveLength(fixtures.length)
|
||||
const actualContext = contextOf(actualLogs.map(log => log.content))
|
||||
const fixtureContext = contextOf(fixtures)
|
||||
|
||||
@@ -20,9 +20,9 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"reasoning","text":"The user wants me to begin with \"Reading the workspace now.\" and call bash with \"echo alpha\" in the same message. Then after the tool result, reply with the single word DONE and stop."}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":1,"block":{"type":"text","text":"Reading the workspace now."}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":2,"block":{"type":"tool-call","id":"call_00_1yZGg4XTqe0N5r1rnDLx5082","name":"bash","arguments":"{\"command\": \"echo alpha\", \"description\": \"Print alpha to stdout\"}"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":7788,"outputTokens":109,"cacheReadTokens":0,"reasoningTokens":42}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":7788,"outputTokens":109,"totalTokens":7897,"cacheReadTokens":0,"reasoningTokens":42}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"reasoning","text":"The user wants me to begin with \"Reading the workspace now.\" and call bash with \"echo alpha\" in the same message. Then after the tool result, reply with the single word DONE and stop."},{"type":"text","text":"Reading the workspace now."},{"type":"tool-call","id":"call_00_1yZGg4XTqe0N5r1rnDLx5082","name":"bash","arguments":"{\"command\": \"echo alpha\", \"description\": \"Print alpha to stdout\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-flash"},"id":"{{message:3}}"},"usage":{"inputTokens":7788,"outputTokens":109,"cacheReadTokens":0,"reasoningTokens":42}},"sourceEventSeqs":[12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"reasoning","text":"The user wants me to begin with \"Reading the workspace now.\" and call bash with \"echo alpha\" in the same message. Then after the tool result, reply with the single word DONE and stop."},{"type":"text","text":"Reading the workspace now."},{"type":"tool-call","id":"call_00_1yZGg4XTqe0N5r1rnDLx5082","name":"bash","arguments":"{\"command\": \"echo alpha\", \"description\": \"Print alpha to stdout\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-flash"},"id":"{{message:3}}"},"usage":{"inputTokens":7788,"outputTokens":109,"totalTokens":7897,"cacheReadTokens":0,"reasoningTokens":42}},"sourceEventSeqs":[12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88],"surfaceOp":"append"}
|
||||
{"type":"tool/call","data":{"turn":1,"step":1,"callId":"call_00_1yZGg4XTqe0N5r1rnDLx5082","name":"bash","arguments":"{\"command\": \"echo alpha\", \"description\": \"Print alpha to stdout\"}"}}
|
||||
{"type":"tool/result","data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"call_00_1yZGg4XTqe0N5r1rnDLx5082"},"content":[{"type":"tool-result","toolCallId":"call_00_1yZGg4XTqe0N5r1rnDLx5082","content":[{"type":"text","text":"alpha\n"}],"isError":false}],"role":"user","id":"{{message:4}}"}},"sourceEventSeqs":[90],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":1}}
|
||||
@@ -31,8 +31,8 @@
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"text-delta","index":0,"text":"D"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"text-delta","index":0,"text":"ONE"}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"block-end","index":0,"block":{"type":"text","text":"DONE"}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":103,"outputTokens":3,"cacheReadTokens":7808,"reasoningTokens":0}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":103,"outputTokens":3,"totalTokens":7914,"cacheReadTokens":7808,"reasoningTokens":0}}}}
|
||||
{"type":"assistant/chunk","data":{"turn":1,"step":2,"chunk":{"type":"finish","reason":{"kind":"stop"}}}}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"text","text":"DONE"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-flash"},"id":"{{message:5}}"},"usage":{"inputTokens":103,"outputTokens":3,"cacheReadTokens":7808,"reasoningTokens":0}},"sourceEventSeqs":[94,95,96,97,98,99],"surfaceOp":"append"}
|
||||
{"type":"assistant/message","data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"text","text":"DONE"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-flash"},"id":"{{message:5}}"},"usage":{"inputTokens":103,"outputTokens":3,"totalTokens":7914,"cacheReadTokens":7808,"reasoningTokens":0}},"sourceEventSeqs":[94,95,96,97,98,99],"surfaceOp":"append"}
|
||||
{"type":"step/end","data":{"turn":1,"step":2}}
|
||||
{"type":"turn/end","data":{"turn":1,"reason":{"kind":"completed"}}}
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
- banner:
|
||||
- navigation "Session hierarchy":
|
||||
- button "Begin your reply with the" [disabled]
|
||||
- img
|
||||
- text: Standard mode
|
||||
- button "Session log":
|
||||
- text: Session log
|
||||
- img
|
||||
- tablist:
|
||||
- tab "Chat" [selected]
|
||||
- tab "Trajectory"
|
||||
- button "System prompt":
|
||||
- img
|
||||
- img
|
||||
- text: System prompt
|
||||
- text: Begin your reply with the plain sentence "Reading the workspace now." as text, and in that same message call the bash tool with the command "echo alpha". After the tool result, reply with the single word DONE and stop. {{clock}}
|
||||
- button "Copy":
|
||||
- img
|
||||
- button "Context injection @deepseek-ai/dsh-system-prompt":
|
||||
- img
|
||||
- img
|
||||
- text: Context injection @deepseek-ai/dsh-system-prompt
|
||||
- button "Think The user wants me to begin with \"Reading the workspace now.\" and call bash with \"echo alpha\" in the same message. Then after the tool result, reply with the single word DONE and stop.":
|
||||
- img
|
||||
- img
|
||||
- text: Think The user wants me to begin with "Reading the workspace now." and call bash with "echo alpha" in the same message. Then after the tool result, reply with the single word DONE and stop.
|
||||
- paragraph: Reading the workspace now.
|
||||
- button "Bash Print alpha to stdout":
|
||||
- img
|
||||
- img
|
||||
- text: Bash Print alpha to stdout
|
||||
- paragraph: DONE
|
||||
- button "Turn usage 15.8K tok · Cache hit 49.7%" [expanded]:
|
||||
- img
|
||||
- text: Turn usage 15.8K tok · Cache hit 49.7%
|
||||
- term: Provider / model
|
||||
- definition: deepseek-official/deepseek-v4-flash
|
||||
- term: Uncached input
|
||||
- definition: 7,891 tok
|
||||
- term: Cached input
|
||||
- definition: 7,808 tok
|
||||
- term: Output
|
||||
- definition: 112 tok (42 tok reasoning)
|
||||
- term: Total
|
||||
- definition: 15,811 tok
|
||||
- button "Copy":
|
||||
- img
|
||||
- button "Good response":
|
||||
- img
|
||||
- button "Bad response":
|
||||
- img
|
||||
- button "Branch into a new conversation":
|
||||
- img
|
||||
- text: {{clock}} Ran for {{duration}} TTFT {{duration}} {{throughput}} tok/s
|
||||
- textbox "Message the agent"
|
||||
- button "Commands":
|
||||
- img
|
||||
- 'button "Access mode, current: Workspace Write"': Workspace Write
|
||||
- button "Select model, current DeepSeek-V4-Flash":
|
||||
- text: DeepSeek-V4-Flash
|
||||
- img
|
||||
- button "6% of context used"
|
||||
- button "Send message" [disabled]
|
||||
- text: 1 turns · 2 steps LLM {{duration}} · Tool call {{duration}} TTFT avg {{duration}} · {{throughput}} tok/s Cache hit 50% Input 15.7K tok · Output 112 tok
|
||||
@@ -75,6 +75,7 @@
|
||||
"@deepseek-ai/dsh-util-workspace-path": ["./packages/util/workspace-path/src/index.ts"],
|
||||
"@deepseek-ai/dsh-session-stats/types": ["./packages/session/session-stats/src/types.ts"],
|
||||
"@deepseek-ai/dsh-session-stats/client": ["./packages/session/session-stats/src/client.ts"],
|
||||
"@deepseek-ai/dsh-token-meter/client": ["./packages/llm/token-meter/src/client.ts"],
|
||||
"@deepseek-ai/dsh-plan-mode/types": ["./packages/plan/plan-mode/src/types.ts"],
|
||||
"@deepseek-ai/dsh-plan-mode/client": ["./packages/plan/plan-mode/src/client.ts"],
|
||||
"@deepseek-ai/dsh-agent-presets/types": ["./packages/preset/agent-presets/src/types.ts"],
|
||||
|
||||
Reference in New Issue
Block a user