Grill Me · 23/24 tests
平均 5.85M candidate token、18.50 分钟;在当前面板中分数与精确测试通过数最高。
这是 Workflow Arena 的 Luna 五工作流、八任务横评面板:比较产品分、精确测试、workflow 完成率和候选侧资源,把五种方法放在同一组任务尺度上观察。
5 workflows × 8 tasks × 3
每条 run 两个 score
compact 表中的通过标志
completed / protocol-failed / token-limit
宏均分是八个 task cell 等权平均;每格只有三条。
平均 5.85M candidate token、18.50 分钟;在当前面板中分数与精确测试通过数最高。
Superpowers Full 为 89.64、21/24 tests,但平均 32.99M token 和 48.38 分钟,约为 Grill Me 的 5.64× token、2.62× 时间。
MatrixSpec L0 得分 86.71、19/24 tests,但 workflow 只完成 11/24;协议失败必须单列。
该面板不能回答“相比原生 Luna 提升多少”,也不能与本站 Terra 单任务五组直接相减。
Score、tests 和 workflow completion 是不同终点;token、时间和 credits 只描述候选侧计算负担。
24/24 workflow · 23/24 tests · 11.58 operator decisions/run
24/24 workflow · 21/24 tests · 8.92 operator decisions/run
11/24 workflow · 19/24 tests · 10.33 operator decisions/run
24/24 workflow · 16/24 tests · 0.08 operator decisions/run
24/24 workflow · 14/24 tests · 0.00 operator decisions/run
每个单元格是三条 run 的 mean score,再按八个任务等权求宏均分。
| Workflow | CLI fields | ESLint | pytest plugin | pytest addini | bat sanitize | SQLModel | Axum | Prometheus | Macro |
|---|---|---|---|---|---|---|---|---|---|
| Grill Me | 100 | 99.67 | 98.67 | 86.83 | 97.33 | 86 | 97.83 | 70.83 | 92.14 |
| Superpowers Full | 99.17 | 99 | 99.50 | 89.50 | 75.33 | 84.33 | 97.50 | 72.83 | 89.64 |
| MatrixSpec L0 | 94.50 | 94.50 | 95.17 | 87.17 | 87 | 84.83 | 87.17 | 63.33 | 86.71 |
| OpenSpec core | 83.33 | 48 | 94.83 | 79.33 | 89.83 | 81.67 | 65 | 34 | 72.00 |
| Ponytail | 78.33 | 47.67 | 94.50 | 85 | 80 | 73 | 61.67 | 41.33 | 70.19 |
稳定高分区:CLI、pytest plugin 与 Axum 上,多数 workflow 都能达到较高分;流程差异更容易体现在资源和完成状态。
分化区:ESLint 和 Prometheus 上差距最大。Grill Me 的 ESLint 为 99.67,但 Prometheus 仍只有 70.83;单一宏均分会隐藏任务敏感性。
/grill-me wrapper显式入口,原始内容只是启动一个 /grilling session;它禁止模型自动调用。
grilling primitive沿决策树一次问一个问题,每题给推荐答案;可查事实自己探索,产品决策交给用户。
强制 operator marker、多轮 GT 回答和实施前共享理解审批;批准后才允许原生实现与验证。
它没有 Full Superpowers 的正式 spec、writing-plans、TDD、任务拆分、子代理实施或独立 review loop。当前结果测量的是“两个窄 Skill + harness 强制的 GT 决策闭环”,不能归因于一行 prompt。
9 条是独立 reviewer 修改产品文件,另有 review 尝试耗尽、无效 stage verdict、以及 review 前创建 full baseline 文档各 1 条。报告保留这些失败,没有用高盲评分覆盖流程终态。