Full With · EXPERIMENT DETAIL

Full With · 完整 Superpowers 复合流程

当设计、任务拆分、多代理、测试、review 和最终验证全部联动时,增加的质量是否值得资源?

01 / 结果快照
盲评分99.00
tool calls / 条244.0
dedup token / 条15.87M
墙钟 / 条40:24

证据:三条 run 平均 244 次 tool calls、15.87M dedup token、99.00 分;Full − Without 的新增 token 最大阶段代理是 coordinate(37.9%)。

边界:run-02 在 token cap 截止但得分 100,说明产品分、流程完成状态和 hidden acceptance proxy 不能合并成一个结论。

02 / 作业路径

这个条件具体要求候选做什么?

  1. 01brainstorming 与行为设计 / spec
  2. 02writing-plans 与任务拆分
  3. 03parent 派发 child agents 并等待结果
  4. 04实现、TDD / focused tests、review / 修复
  5. 05verification-before-completion 与完成 gate
03 / 阶段与 actor

时间 / token 的主要落点

协调36.8%
实现27.0%
Review14.6%
测试 / 调试9.0%
计划4.9%
需求 / 设计3.2%

阶段是可见动作分类器的派生标签。Full 页面中的 coordinate 具体包括 spawn_agentwait_agentfollowup_tasksend_message、worktree / approval / guardian 等过程动作,不等同于“模型在思考”。

04 / 本组 run

每条轨迹的真实状态

run-02100.0
墙钟
47:50
Token
18.47M
tool calls
266
Operator
11
需求审批
Review / 修复
0 / 0
状态
token_cap
首次 mutation
+7:27
在总时间线中定位 →
run-0399.0
墙钟
32:08
Token
13.32M
tool calls
186
Operator
7
需求审批
Review / 修复
0 / 0
状态
completed
首次 mutation
+5:18
在总时间线中定位 →
run-0698.0
墙钟
41:14
Token
15.84M
tool calls
280
Operator
3
需求审批
Review / 修复
0 / 0
状态
completed
首次 mutation
+3:40
在总时间线中定位 →
05 / 回到问题

把本组放回五组比较

这张子页解释“这条方法怎么走”;首页的三个研究问题再回答质量—成本、需求 / review 阶梯和 token 归因。

返回首页三问 →