Field note · 2026-09-02实践记录 · 2026-09-02

I marked 61 entries “high confidence.” Not one of them named a test. 我给 61 条经验标了「高置信度」,没有一条说得出验证方式。

I keep a skill library — a folder of markdown files my coding agent loads when a task matches: secure boot, SDK migration, partition layouts. After I finish something I write down what I learned, so the next project doesn't start from zero. 我维护一个 skill 库,就是一个 markdown 文件夹,任务匹配上了编程 agent 就加载:安全启动、SDK 升级、分区布局。每做完一件事我把学到的写下来,下个项目就不用从零开始。

Months ago I put a rule at the top of the framework document that governs those files: 几个月前,我在管着这些文件的框架文档开头写下一条规矩:

No entry may be marked high confidence without an executable check that actually passed. 没有一条实测通过的可执行检查,不许标 high。

Last week I wrote the checker that enforces it, and pointed it at the two packages the framework had produced. 上周我把执行这条规矩的检查器写了出来,拿它对准这套框架自己产出的两个包。

61entries marked high条标着 high
0of them named a test条说得出验证方式
54graded above what my own table permits条高过我自己那张表的上限
180errors in two files I had been using个错误,出在两个我一直在用的文件里
claims more than the evidence allows标的比证据允许的高marked high, within the cap on other grounds标了 high,按其他依据未超上限one dot = one entry一个点 = 一条经验

Nobody had broken the rule. Nothing had ever enforced it. 没有人违反过这条规矩。只是从来没有任何东西执行过它。

Why written rules stop working写下来的规矩为什么会失效

I don't think this is specific to me. You put a discipline in a document. The model reads it at the start of a session and follows it. Forty messages later the document has been summarised, the context has rolled over, and the entries coming out still have the right fields and the same confident label. The format is what carries across a summarisation. The reason the format exists does not, and nothing in the output tells you which of the two you are looking at. 我不认为这只发生在我身上。你把一条纪律写进文档,模型在会话开头读到它、照做。四十条消息之后,文档被摘要压缩了,上下文已经滚过一轮,产出的条目仍然带着对的字段和一样自信的标签。能扛过一次摘要压缩的是格式,格式存在的那个理由扛不过去,而输出里没有任何东西告诉你眼前这条是哪一种。

You don't notice, because nothing breaks. A wrong high just sits there looking authoritative. Six months later someone acts on advice that was true once, on one board, and isn't any more. 你不会发现,因为什么都没坏。一个标错的 high 就那么摆着,一副权威的样子。半年后有人照着一条曾经在某一块板上成立、现在已经不成立的建议动手。

That is the argument for making rules executable, and it is why I had stopped trusting the document version of mine. 这就是把规矩做成能跑的东西的理由,也是我后来不再相信自己那份文档版规矩的原因。

What I could and couldn't turn into a check哪些能变成检查,哪些不能

Some of the discipline needs judgement. Is this principle really exception-free? Was that contradiction preserved, or averaged away? Mechanising those would just produce false positives, and a checker that fires on good input gets ignored, which leaves you worse off than before. 有些纪律需要判断力。这条原理是不是真的举不出反例?那个矛盾是被保留了,还是被和成了平均值?把这些硬做成机械判断只会制造误报,而一个在好输入上乱报的检查器会被人忽略,那比原来还糟。

Plenty of it is decidable, though. Does the entry have a source, a date, a reproduction record, a verification method? Is the date absolute, or does it say “recently”? Given the evidence actually present, is the declared confidence higher than the table permits? Is the reproduction record a list of platforms, or did someone just type “3 times”? Was the pitfall split into a transferable lesson and a platform-specific fact, or is the lesson still welded to the one chip where I learned it? 但相当一部分是可判的。这条有没有来源、日期、复现记录、验证方式?日期是绝对日期,还是写着「最近」?按实际拿得出的证据,标的置信度有没有高过表允许的上限?复现记录是一张平台列表,还是有人随手打了个「3 次」?这个坑拆成了可迁移的教训和平台专属的事实,还是教训仍然焊死在我学到它的那一颗芯片上?

That came to 37 checks. It prints findings and exits non-zero. 数下来 37 条检查。它打印问题,退出码非零。

Update: it is 38 now — V25, which asks whether a package says what must never be done, was added two days after this note. 后记:现在是 38 条 —— 两天后加了 V25,查一个包有没有写明这个领域里绝对不能做的事。

whetstone verify examples/demo-skill --brief
$ whetstone verify examples/demo-skill --brief

ERROR (5)
  ✗ SKILL.md:31  [V16] [溯源] 复现记录 has two lines for the same platform/project: 'soc-x'
  ✗ SKILL.md:31  [V12] [溯源] L2 claims 置信度 high, the table allows only low — 复现记录 has 1 line(s), needs 2 (single-platform L1/L2 is capped at low — the §7 constraint outranks the table)
  ✗ pitfalls.md:5  [V14] [从上层状态反推底层状态会看错] 复现记录 is a bare count 「2 次」
  ✗ pitfalls.md:5  [V12] [从上层状态反推底层状态会看错] L2 claims 置信度 high, the table allows only low — 复现记录 has 0 line(s), needs 2 (single-platform L1/L2 is capped at low — the §7 constraint outranks the table)
  ✗ params/soc-x.md:6  [V10] [存储起始位置] L3 claims 置信度 high, the table allows only low — 验证方式 never passed a test; 复现记录 has 1 line(s), needs 2

WARN (3)
  ! SKILL.md:18  [V19] concrete value in the body (hex address/value): '0x1A40'
  ! SKILL.md  [V25] no section says what must NEVER be done in this domain
  ! params/soc-x.md  [V22] no 替代记录 section

summary: 5 error(s), 3 warning(s), 0 info

Today's real output, unedited — --brief only drops the “→ how to fix” hint under each line. examples/demo-skill ships in the repo, deliberately flawed — run it and compare. 这是今天的真实输出,没有改动 —— --brief 只是去掉了每行下面那句「→ 怎么改」的提示。examples/demo-skill 就在仓库里,故意写坏的,跑一遍就能对照。

The grading is asymmetric on purpose: a reviewer can downgrade any entry but cannot upgrade past the conditions, so going above the cap is an error while going below it is only a note. I have never once caught myself being too harsh on my own notes, and the checker is built around that. 这个评级是故意做成不对称的:人审可以把任何一条下调,但不能越过条件上调,所以高过上限报错,低于上限只是提示。我从来没有一次抓到自己把笔记标得太严,检查器就是照着这一点造的。

Level等级What the evidence has to show证据必须能拿出什么
highAn executable check that passed, plus reproduction on at least two distinct platforms or projects.一条可执行检查实测通过,加上在至少两个不同平台或项目上复现过。
medReproduction on two or more. Or verified once — and that second route is open to platform facts only.在两个以上平台复现过。或者实测通过一次 —— 而后一条路只对平台事实开放。
lowEverything else, including a method that passed its check on the single platform where you first met it.其余全部,包括那种在你第一次遇到它的那个平台上验证通过了的做法。
The whole table. Verifying a method once doesn't make it general, which is why a single-platform lesson is capped at low even when its check passes. Try it interactively →整张表就这么大。一个做法只验过一次不等于它通用,所以单平台的经验即使验证通过也封顶 low。在首页亲手试一下 →

Four ways past it四条绕过去的路

I asked three reviews to attack the checker, each in a separate context so none of them started from my assumptions. The third came back blocking, with four ways to get a single-platform lesson graded high without a single warning. 我请三份评审去攻击这个检查器,各自在独立的上下文里跑,谁都不从我的预设出发。第三份回来判定阻塞,给出四条能让单平台的经验拿到 high、而且一个告警都不出的路。

  1. 1

    Delete a separator删掉一个分隔符

    Reproduction lines read platform · date · pointer and are keyed on the platform field. Take out the · and the whole line becomes the key, so two records from the same board count as two platforms.复现记录的格式是 平台 · 日期 · 指针,按平台字段做键。把 · 去掉,整行就成了键,于是同一块板的两条记录算成两个平台。

    FIXED · MUTATION-TESTED
  2. 2

    Put “L3” in a heading在标题里写「L3」

    Layer was inferred from the nearest heading containing L1 to L4. A section called “Don't put L3 values in the body” — an instruction against doing the thing — reclassified everything beneath it and raised the ceiling.层级是从最近的、含 L1 到 L4 的标题推断的。一个叫「不要把 L3 的值写进正文」的小节 —— 一句反对这么做的告诫 —— 把它底下所有条目重新分了类,还把上限抬高了。

    FIXED · MUTATION-TESTED
  3. 3

    Add a suffix加个后缀

    soc-x and soc-x(second run) normalised to two different keys, which got the same board counted twice and cleared the gate.soc-x 和 soc-x(第二次) 归一化成两个不同的键,同一块板因此被算成两个,就过了升级条件。

    FIXED · MUTATION-TESTED
  4. 4

    Drop the leading pipe不写前导竖线

    GFM tables don't require | at the start of a line. My parser did. An entire platform-parameter file could be invisible to every check while the run reported no problems — including the check whose only job was to notice a table with no confidence column, because it looked for the leading pipe the same way.GFM 表格不要求行首有 |,我的解析器要求。于是一整份平台参数文件可以对所有检查隐形,而那次运行报告零问题 —— 连那条唯一职责就是发现「这张表没有置信度列」的检查也一样,因为它判前导竖线的方式跟解析器一模一样。

    FIXED · MUTATION-TESTED

All four are fixed. Each one is now a mutation-test entry: remove the fix and the suite has to fail. 四条都修了。每一条现在都是一个变异测试条目:把修法拿掉,整套测试必须失败。

Two of my tests were fake我有两条测试是假的

Making the tests prove the fixes turned up something worse. One test was watching the wrong error code, so it passed because a different rule fired. Another declared high in its own fixture data, which meant it errored whether or not the bug existed. It had been green the entire time and proving nothing. 让测试去证明那些修法的时候,翻出了更难看的东西。一条测试盯的是错误的检查码,它能通过是因为另一条规则响了。另一条在自己的测试数据里就写了 high,于是不管 bug 在不在都会报错。它一直是绿的,什么也没证明。

The coverage check I had added specifically to catch this had its own problem. Its failures ran inside a pipeline, so the counter incremented in a subshell and disappeared. It printed FAIL and reported 0 failed. 我专门为了抓这种事加的覆盖率检查,自己也有毛病。它的失败跑在管道里,计数在子 shell 里加,加完就没了。它打印 FAIL,然后报告 0 failed。

I found all three the same way: by deleting checking logic and requiring the suite to go red. A test that passes and a test that would fail are different things, and breaking the code on purpose is the only way I know to tell them apart. This is the 61 entries again with one more layer on it — I had written both the rule and the test that was supposed to hold me to it, and neither was doing any work. 这三个我都是用同一个办法找出来的:删掉检查逻辑,要求整套测试变成失败。「测试通过了」和「这条测试真的会失败」是两回事,而故意把代码弄坏是我知道的唯一区分办法。这就是那 61 条又来了一遍,只是上面多压了一层:规矩是我写的,本该管住我的那条测试也是我写的,两个都没在干活。

If you keep this kind of thing如果你也维护这类东西

If you keep a skill library, a rules file, a CLAUDE.md, or any corpus you curate by hand, the question worth asking is which of your rules would still be obeyed after you stopped believing in them. A rule is either run by a command, or checked by a person on some schedule, or it is decoration — and decoration is what you get when nobody chooses. All three look the same from outside the file, which is why mine sat there for months without anyone noticing, me included. 如果你也维护 skill 库、规则文件、CLAUDE.md,或者任何手工整理的语料,值得问的问题是:你那些规矩里,哪几条在你不再相信它们之后还会被遵守。一条规矩要么有命令在执行它,要么有人按某个周期检查它,要么就是装饰 —— 而没人去选的时候,得到的就是装饰。从文件外面看,这三种长得一模一样,所以我那条在那儿摆了几个月都没人发现,包括我自己。

The tool is whetstone. MIT, standard library only, works with any runtime that reads the Agent Skills layout. There's a deliberately broken example package in the repo so the checker has something to catch, and --explain prints what it checks, where each check is narrower than the rule behind it, and what it will not judge at all. 工具叫 whetstone。MIT,只用标准库,任何读 Agent Skills 布局的 runtime 都能装。仓库里有一个故意写坏的示例包,好让检查器有东西可抓;--explain 会打印它查什么、每条检查比背后那条规则窄在哪、以及它压根不判什么。