← Whetstone overview← 返回 Whetstone 总览
Feature · Use-time check功能 · 用时纠错

Check the note when you use it, not when you file it.经验对不对,用的时候才判得准。

When a note is filed there is only text to judge. When it is used, the code, the log and the board are right there. So that is when whetstone checks it: a note that disagrees with reality gets sorted, reported with its evidence, and left exactly as it is until you say yes. 经验入库的时候,手里只有文本可判;用的时候,代码、日志、板子都在眼前。所以 whetstone 在用的时候核它:和实际对不上的经验,先分类,带着证据报给你,在你说好之前一个字都不改。

5kinds of disagreement; only one means the note was wrong类冲突,只有一类是经验本身错了
0changes to a note without your yes处改动,没有你点头就发生
22agreeing real records, at least, before you are asked more briefly条真实记录全部判对,至少这么多,才简化问法
The premise前提

A note can only be graded against the thing it describes.一条经验,只能拿它描述的那个东西来判。

Every note in the library went through review once, on the way in. At that moment the reviewer, human or model, had the session transcript and the text of the note, and nothing else. That is thin ground: an unguided model judge picks the better of two skills 46.4% of the time (SkillLens, arXiv 2605.23899), which is a coin toss. Then the code moves on, the platform changes, and the note keeps its confident label.库里的每一条经验,入库时都审过一次。那一刻,审的人(或模型)手里只有会话记录和经验的文本,别的什么都没有。这个依据很薄:无引导的模型评委在两个 skill 里挑更好的那个,准确率 46.4%(SkillLens,arXiv 2605.23899),等于抛硬币。之后代码往前走、平台换了,经验还顶着原来那个自信的标签。

The moment a note is used is different. The agent is about to act on it, with the current source tree, the git history and often the hardware in reach. Whether the note still holds is no longer a matter of taste; it can be looked up. The same moment is also where the most common failure hides: the agent notices the mismatch, works around it, finishes the task and says nothing, so the note misleads the next session too.用的那一刻不一样。agent 马上要照着它动手,当前代码、git 历史,常常连硬件都够得着。这条经验还成不成立,不再是品味问题,是查得到的事实。最常见的坏法也藏在这一刻:agent 发现对不上,自己绕过去把活干完,一个字不提,于是这条经验接着误导下一次会话。

“The next time a note is used is the real exam. Everything before that is a rehearsal.”「下次真用这条经验的时候,才是真正的考试。在那之前的,都是彩排。」

From the design review, 2026-09-29设计评审,2026-09-29
The path路径

From “that doesn’t match” to a changed note.从「对不上」到经验被改,中间每一步。

② checks the note twice: at the commit it was written from, and in the code as it is now. ③ sorts the mismatch into one of five kinds. ④ puts the AI’s call on file before anyone is asked, and ⑤ is the question itself, carrying the fingerprint that ④ printed. Only after ⑥ does anything change, and then only by appending.② 把经验核两次:在它的来源 commit 上核一次,在现在的代码上核一次。③ 把对不上的情况分到五类之一。④ 在问任何人之前先把 AI 的判断写进记录,⑤ 才是提问,报告里带着 ④ 打印出的那行指纹。要等 ⑥ 你做了决定才会改,而且只追加。

Five kinds五类冲突

A mismatch is not yet a wrong note.对不上,还不等于经验错了。

The sorting is mechanical: check the claim at the note’s own source commit, then at HEAD, then ask whether the note is a warning.分类是机械的:在经验自己的来源 commit 上核一次说法,在当前代码上再核一次,再看这条经验是不是一条警告。

stalestale · 过期

True then, not now当时对,现在不对

It held at its source commit; a later commit changed the code it describes.在来源 commit 上成立;后来有提交改过它描述的代码。

Update it, keep the old value, record the version where it stopped holding.更新,旧值保留,写明从哪个版本起不成立。
scopescope · 范围不符

Right, somewhere else对,但不是这里

It still holds where it was learned. This is a different platform or branch.在学到它的地方仍然成立,只是现在换了平台或分支。

Add a platform variant. The general rule is not touched.加一个平台变体,通用规律不动。
wrongwrong · 原本就错

Never held从来没成立过

Checked at its own source commit, the claim was already false.在它自己的来源 commit 上核,说法就已经不成立。

Supersede it, and keep why it was misjudged as a lesson of its own.替代它;当初为什么判错,本身记成一条教训。
code-regresscode-regress · 代码又犯了

The note is right. The code is not.经验没错,代码错了

The note is a warning, and the code now does the very thing it warns against. A pitfall exists because the obvious reasoning fails there, so reasoning cannot overrule it.这条经验是一条警告,而现在的代码正好是它警告的那种写法。坑之所以是坑,正是因为按直觉推会错,所以推理推翻不了它。

Leave the note alone and report the code. Count it as a confirmation.经验不动,报代码问题;这次算经验被印证了一次。
clashclash · 经验互相矛盾

Two notes disagree两条经验说法不同

Two skills say different things about the same step.两个 skill 对同一步给了不同的做法。

Both go in front of you. The one you set aside is marked, not deleted.两条都摆到你面前;落选那条标失效,不删。
not a conflict不算冲突

A new case is only a supplement新情况只是补充

Something the note never covered is not a contradiction. Untrained memory managers delete the old entry and add a new one here, and fragment the memory (Memory-R1, arXiv 2508.19828).经验没覆盖到的新情况不是矛盾。未经训练的记忆管理器在这里会删旧加新,把记忆打碎(Memory-R1,arXiv 2508.19828)。

Goes through normal distillation as a new entry or variant.走正常蒸馏,作为新条目或变体。
Evidence证据

What the AI saw decides what it may say.看到了什么,决定 AI 能说到哪一步。

One rule sits above the table: a note that came from a measurement can only be overturned by a new measurement, never by reasoning. Many notes record exactly the case where the documentation or the first principle was wrong.表上面压着一条规矩:实测来的经验,只能被新的实测推翻,不能被推理推翻。很多经验记的恰恰就是「文档或原理说 X、测出来是 Y」。

Grade档Counts as this grade算这一档的The AI may sayAI 能说到
seenA file:line in the current code, the note’s own check re-run, a register read back. You could repeat it.当前代码的 file:line、重跑了经验自带的检查、读回了寄存器。你照着能重做一遍。“This is wrong; change it to X; here is the evidence.”「这里不对,建议改成 X,证据是……」
inferredWorked out from principle, analogy or documentation.从原理推、从别处类比、从文档读出来的。“I suspect this is wrong, because… please judge.”「我怀疑这条不对,理由是……请你判断」
unseenNo board, no platform, or the check could not be run.板子不在、平台不在手边,或者核不了。Only that there is a conflict. No verdict.只报冲突,不给结论。
Only the seen grade can ever earn a shorter way of asking. What the AI inferred, however often it turned out right, says nothing about what it would see.只有 seen 这一档有资格简化问法。AI 推出来的,判对的次数再多,也说明不了它看到的时候会怎样。
The record记录

The AI writes its call down before it hears yours.AI 在听到你的回答之前,先写下自己的判断。

The party being scored is the one keeping the record. Written in one line after your answer, its “own” call could simply be copied from yours, and agreement would read 100% with nothing on screen to say so. So the record comes in two steps.写记录的,正是被打分的一方。等你回答完再一次写一行,它「自己的判断」完全可以照着你的答案填,一致率就成了 100%,屏幕上还没有任何迹象。所以记录分两步。

① The call goes on file and ② the id and fingerprint come back. ③ The card cannot exist before that, because it has to carry the fingerprint. ④ Your answer arrives and ⑤ is appended as a separate line. ⑥ If an entry was reported twice in one session, the first call is the one scored.① 判断先写进记录,② 拿回编号和指纹。③ 报告在这之前出不来,因为它必须带着这行指纹。④ 你回答后,⑤ 另起一行追加。⑥ 同一条经验在一次会话里报了两次的,只给第一次的判断打分。

Four commands, one file.四个命令,一个文件。

  • report — the AI’s call, before asking: kind, evidence grade, proposed action, how sure it is, and the model that made it.report —— 问之前的 AI 判断:哪一类、哪档证据、建议怎么办、有几成把握,以及是哪个模型判的。
  • resolve — your answer to a report already on file. No report, no answer.resolve —— 你对某条已有报告的回答。没有报告,就收不了回答。
  • miss — a conflict the AI met and never reported, found by you. It carries no AI call, because there was none.miss —— AI 碰到了却没报、被你发现的冲突。它不带 AI 判断,因为 AI 根本没判过。
  • stats — scores each group and says how it may be asked next time.stats —— 给每一组打分,并说明下次该怎么问。

Everything lands in journal/review-decisions.jsonl, which is git-ignored and stays on your machine.全部写进 journal/review-decisions.jsonl,它在 gitignore 里,只留在你本机。

$ whetstone decision report --entry boot-x/pitfall-7 \
    --ai-type stale --ai-evidence seen --ai-action update \
    --ai-certainty high --model model-a …
reported C-20260929-01 · fingerprint b19b2a

$ whetstone decision stats
use-time conflicts — 25 reported by the AI, 0 missed
  model: model-a (all 25 reports)

  group                    judged  basis  lower-bound  ask as
  code-regress × inferred       1    1/1        0.100  full (nothing seen)
  stale × seen                 22  22/22        0.901  summary
  wrong × seen                  2    0/2        0.000  full
Real output from a generic demo log, trimmed for width: two columns and the tail of two lines are cut.用通用演示数据跑出的真实输出,为宽度做了截短:去掉了两列和两行的末尾。
The score打分

A score changes how you are asked. Never whether.分数只改怎么问你,不改问不问。

Agreement is judged by an exact one-sided lower bound (Clopper-Pearson, δ = 0.1), not by the share that happened to agree. Setting thresholds from the sample rate met its target only 47.5% of the time in Trust or Escalate (arXiv 2407.18370).判对率用精确单侧下界(Clopper-Pearson,δ = 0.1),不用碰巧判对的比例。Trust or Escalate(arXiv 2407.18370)里,直接用样本一致率定阈值,只有 47.5% 的时候达到了目标。

Eighteen right out of twenty reads as 90%, but it only proves 75.5%. The least that proves 90% is twenty-two out of twenty-two.20 条对 18 条,看上去是 90%,能证明的只有 75.5%。要证明 90%,最少也要 22 条全对。

  • A missed conflict sends its whole kind back to full confirmation; the count starts again after the miss.出现一次漏报,这一类立刻退回完整确认,之后只算漏报之后的记录。
  • A new model starts every group from zero. Agreement one model earned says nothing about the next.换了模型,每一组都从 0 开始。一个模型攒下的判对记录,说明不了下一个。
  • If more than 30% of the calls rest on inference or on nothing seen, no group is relaxed at all.推断和看不到的判断超过三成,任何一组都不放宽。
  • Safety-relevant notes are always asked in full and re-measured before any change.安全相关的经验永远完整确认,改之前先重测一次。
Where the ideas came from思路从哪来

Borrowed where it held up, left behind where it did not.能用的借过来,用不上的不借。

The design came out of reading twelve open-source memory systems at the code level, a few dozen papers, and three self-improvement projects. None of them lets a model decide alone whether an old note should change. The ones that work share one shape: mechanical checks underneath, the model only classifying, a person when it is unsure.这套设计来自一次调研:读了十二个开源记忆系统的代码、几十篇论文、三个自我进化项目。没有一个能让模型单独决定旧经验该不该改。做得好的都是同一个结构:机械检查打底,模型只做分类,没把握就交给人。

Took借了

karpathy/autoresearch

Its loop can run overnight because the agent cannot edit the evaluation. Here the AI may not touch the decision table, the scoring or the statistics code, and a check it writes must fail on the old value before it counts.它能无人值守跑一夜,前提是 agent 改不了评测。这里 AI 不许改判定表、打分算法和统计代码;它新写的检查,拿旧值跑必须失败才算数。

Took借了

alchaincyf/darwin-skill

Its warning when more than 30% of runs are dry runs became the rule here that nothing is relaxed once guesses pass 30%. Its deliberately degraded versions, which every blind judge had to catch, are the same idea as this repo’s mutation tests. Its weighted score deciding what to keep was left behind.它「空跑超过三成就警告」变成了这里的「推断占比超过三成就不放宽」;它故意造劣化版本、要求盲评裁判全部识别,和本仓的故意改坏测试是同一个思路。用加权总分决定留不留,没借。

Took借了

alchaincyf/nuwa-skill

Its edge test expects “probably, but not certain” rather than a confident answer. That is the unseen grade: no evidence, no verdict. Its keyword-count quality check is on the list of things not to do.它的边缘测试要的是「可能……但不确定」,不是斩钉截铁。这就是 unseen 那一档:没证据就不下结论。它数关键词的质量检查,进了禁止清单。

Read读了

TbusOS/ai-doc

The paper notes behind the whole framework: memory systems and self-improving agents. The repeated lesson was that self-correction without a reliable signal tends to fail, which is why the signal here comes from the code in front of the agent.整个框架背后的论文笔记:记忆系统、自我改进的 agent。反复出现的教训是:没有可靠信号的自我纠错往往失败 —— 所以这里的信号,取自 agent 眼前的代码。

Measured实测依据

Scoring an AI judge against people拿人校准 AI 裁判

Trust or Escalate (arXiv 2407.18370) for the exact lower bound and for escalating when unsure. SkillLens (arXiv 2605.23899) for the 46.4% coin toss. On anchoring: experts shown an AI suggestion abandoned a correct call 7% of the time (arXiv 2603.11821).Trust or Escalate(arXiv 2407.18370):精确下界,以及没把握就上交给人。SkillLens(arXiv 2605.23899):46.4% 的抛硬币。关于被带偏:专家看到 AI 建议后,7% 放弃了原本正确的判断(arXiv 2603.11821)。

Measured实测依据

Updating old memory without breaking it改旧记忆而不改坏

SEMV (arXiv 2609.27175) stores a verified contradiction as counter-evidence; negative transfer fell from 5.7% to 0.2%. GLOVE (arXiv 2601.19249) re-runs before rewriting. StateAuditor (arXiv 2608.01619) verifies provenance and order, not meaning. misevolution (arXiv 2509.26354) shows safety constraints being selected away.SEMV(arXiv 2609.27175):核实过的矛盾存为反证,负迁移从 5.7% 降到 0.2%。GLOVE(arXiv 2601.19249):改记忆前先重跑。StateAuditor(arXiv 2608.01619):只验出处和先后,不验语义。misevolution(arXiv 2509.26354):按表现筛选时,安全约束会被一条条淘汰。

Honest caveat坦诚的边界

Four things the record cannot do for you.记录替你做不了的四件事。

It cannot make the AI report. A rule can only ask; a miss is known only when you notice it and say so, so the true number of misses is always at least what the log shows.它没法强制 AI 上报。规则只能要求;漏报只有你发现并说一声才会被记下,所以真实的漏报数只会比记录里多。

It cannot see a report written after your answer. The file would still show the report first. The one tell is the fingerprint line on the card: a card without it counts as not reported, and that check is yours.它看不出「等你回答完才写报告」。文件里照样是报告在前。能看出来的只有报告上那行指纹:没有这一行的报告当作没报,这一关靠你看一眼。

With a single reviewer, agreeing with you is not the same as being right; there is no second person to measure against. And only notes that get used get checked. A note nobody opens stays exactly as confident as the day it was filed.只有一个审核人时,和你判得一样不等于判对了,没有第二个人可以比。另外,只有被用到的经验才会被核;没人打开的那条,永远和入库那天一样自信。

Technical details技术细节

The parts you can check.能核对的那些部分。

Commands命令

whetstone decision report · resolve · miss · stats [--model X]

Fingerprint指纹

The first six hex digits of SHA-256 over the id, entry, kind, evidence, action, certainty, time and model. Recomputed on every read; a line that no longer matches is not scored.对编号、条目、类型、证据档、建议、把握、时间和模型取 SHA-256 的前 6 位。每次读取都重算,对不上的那行不计分。

Thresholds门槛

Summary at a lower bound of 0.90, batch at 0.95, no relaxing above a 30% guess share. Every one of these numbers is a guess meant to be recalibrated from the log itself.摘要确认要下界 ≥ 0.90,批量确认要 ≥ 0.95,推断占比超过三成就不放宽。这几个数都是拍的,留着让记录本身来校准。

Tests测试

119 selftest checks and 31 mutations. Each mutation removes one guard, such as the model filter, the fingerprint or the lower bound, and the suite must go red.119 项自检、31 条故意改坏。每一条拿掉一道保护(按模型过滤、指纹、精确下界……),整套测试必须报红。

Spec规范

spec/use-time-conflicts.md — decision table, report format, scoring, and the list of things never to do.—— 判定表、报告格式、打分规则和禁止清单。

whetstone is MIT, standard library only, and works with any runtime that reads the Agent Skills layout. The use-time check ships in the same command line as the rest.whetstone 是 MIT 协议,只用标准库,任何读 Agent Skills 布局的 runtime 都能用。用时纠错和其他功能在同一个命令行里。