Counting one project twice同一个项目算两次
Reproduction lines are keyed by platform/project. Ten runs on one project is still one line — otherwise the write-back becomes a back door around the gate.复现记录按平台/项目做唯一键。同一个项目跑十次还是一行,否则「回写证据」就成了绕过升级条件的后门。
Whetstone turns a finished dev session — transcript, git diff, the pitfalls you hit — into a portable Agent Skill. Then it holds that skill to an evidence standard you can actually run. Whetstone 把一次做完的开发 —— transcript、git diff、踩过的坑 —— 蒸馏成可移植的 Agent Skill,再用一条能跑的命令守住证据纪律。
Claude Code · Codex · Cursor · any skills-compatible runtime任何支持 skills 的 runtime
verifyverify 里的机械检查The design and pitfalls live in the skill body. Addresses, key lengths, tool names live in one params/<platform>.md. A new chip is a new table, nothing else.设计和坑留在 skill 正文;地址、密钥长度、工具名放进一张 params/<platform>.md。换芯片就新加一张表,别的不动。
The output is a plain-markdown folder with no runtime binding. Hand it to someone on Cursor or Codex; they drop it in and have your know-how.产物是一个纯 markdown 目录,不绑定任何 runtime。递给用 Cursor 或 Codex 的同事,放进去就能用。
Each new session is reconciled against the library: duplicates merge, changed facts are superseded, evidence for old entries accumulates. The library grows instead of resetting.每次新蒸馏都和已有库比对:重复的合并,变了的事实标注替代,旧条目的证据继续累加。库在长,而不是每次归零。
A skill library doesn't fail loudly. It fills up with plausible, once-true advice that still wears a confident label. So every rule a machine can decide became a check with an exit code. Below is a real run on examples/demo-skill, a deliberately broken package that ships with the repo.skill 库不会大声地坏掉,它会慢慢塞满听着有理、只在当时成立、却还顶着高置信标签的条目。所以凡是机器能判的规则,都做成了带退出码的检查。下面是对仓库自带的、故意写坏的 examples/demo-skill 的一次真实运行。
Each of the 38 codes is tested in both directions — it fires on a bad package and stays quiet on a good one — and a mutation test deletes checks one at a time to prove the suite notices: 14 of 14 caught.38 个检查码每一个都正反两个方向测过:坏包上必须报、好包上必须不报。另有变异测试逐条删掉检查,确认测试套件会发现:14 条全部抓到。
Try it. This is the exact rule verify enforces (extraction-framework §7). A reviewer may always lower a level; nobody may raise it past the conditions.自己试试。这就是 verify 执行的那条规则(extraction-framework §7)。人审永远可以往下调,但谁都不能越过条件往上调。
Reproduction lines are keyed by platform/project. Ten runs on one project is still one line — otherwise the write-back becomes a back door around the gate.复现记录按平台/项目做唯一键。同一个项目跑十次还是一行,否则「回写证据」就成了绕过升级条件的后门。
A changed value is appended and marked superseded; the old one stays readable. Being quietly wrong is the expensive kind of wrong, most of all for irreversible values.变了的值只能追加并标注替代,旧值仍然可读。错得没人发现是代价最大的错法,不可逆的值尤其如此。
Review runs in a separate context. Self-review bias is something this design assumes will happen, not something it hopes to avoid.质检在独立的上下文里跑。自评偏差是这套设计预设一定会发生的事,不是指望能躲开的事。
Whether an L1 truly has no counter-example stays with the reviewer, and --explain prints that boundary in full. A mechanised non-mechanical judgement only makes false positives.一条 L1 是不是真的举不出反例,留给人审,--explain 把这条边界原样打印出来。硬把不机械的判断做成机械的,只会制造误报。
What happened when I ran it on my own library: 拿它跑我自己的库,结果是:61 entries marked “high confidence”. Not one could name a test. →61 条标着「高置信度」的经验,没有一条说得出验证方式 →
And after an entry is in the library, it is checked again each time it is used: 入库之后,每次用到一条经验还会再核一次:check the note when you use it, not when you file it →经验对不对,用的时候才判得准 →
verify 全绿,不等于 skill 好用。Every check asks whether the knowledge is true and traceable. None asks whether installing it makes its user better — and the two don't predict each other. That is why the human gates stay. Source: SkillLens (Microsoft).每条检查问的都是「这条知识真不真、能不能追溯」,没有一条问「装上它,使用者是不是变强了」,而这两件事互相预测不了。所以人审关一直保留。数据来源:SkillLens(Microsoft)。
of extractor/consumer pairings in the SkillLens corpus transferred negatively — the skill made things worse.SkillLens 语料里,25% 的「抽取方 × 使用方」组合出现负迁移 —— 装上 skill 反而变差。
is how often an unguided model judge picks the better of two skills. A coin does 50%.无引导的模型评委在两个 skill 里挑对更好那个的比例。掷硬币是 50%。
Think of it as strata. The deeper a piece of knowledge sits, the more platforms it survives — and the closer it lives to the reusable skill body. Click a layer.把它想成岩层:越往下的知识,换越多平台仍然成立,也越靠近可复用的 skill 正文。点任意一层看看。
Two human gates bracket an automatic middle. Nothing reaches your library without your sign-off.两道人审关夹着中间的自动流程。没有你点头,什么都进不了正式库。
Runtime-specific code is fenced into one thin capture adapter. What you share is plain markdown. Everything below the dashed line is optional — installed, it helps; absent, nothing breaks.跟 runtime 绑定的代码只放在薄薄一层采集适配器里。对外分享的是纯 markdown。虚线以下全是可选:装了能增强,不装也不影响使用。
Keys, OTP, the signing chain — and a buffer overrun you only found by dumping bytes.密钥、OTP、签名链,还有一个靠 dump 字节才揪出来的缓冲区越界。
It separates the transferable design and pitfall (L1/L2) from this chip's addresses and key length (L3). You approve at two gates.它把可迁移的设计和坑(L1/L2)跟这颗芯片的地址、密钥长度(L3)分开。你在两道关点头。
verified-boot/产出 verified-boot/Plain markdown, self-contained, and verify is clean.纯 markdown,自包含,verify 通过。
They now have the design, the method, and the byte-dump pitfall. No onboarding call.设计、做法、那个 dump 字节的坑,他马上就有了。不用专门开会讲。
params/soc-b.md. The design and pitfalls carry over untouched.params/soc-b.md。设计和坑原样复用。
A runtime keeps every skill's name + description in context to choose between them. Bodies load only after a pick. So the noise lives in the menu — and that's where lint looks.runtime 把每个 skill 的 name + description 常驻在上下文里用来挑选,正文要等被选中才加载。所以乱就乱在菜单上,lint 查的也正是这里。
Finds missing triggers, overlapping trigger sets and near-identical names — tells deliberate companions (-audit) from real clashes — and checks whether the whole menu still fits.找出缺触发词、触发词重叠、名字几乎一样的 skill,能分清故意成对的(-audit)和真撞车,还会查整张菜单是不是还放得下。
Groups skills into families and writes a browsable INDEX.md.按族归组,生成一份能浏览的 INDEX.md。
Capability line, concrete triggers, and a boundary that pushes siblings apart — baked into the template.能力一句话 + 具体触发词 + 跟同族划清边界,写进模板,新 skill 一生成就符合。
And once the library outgrows what one menu can hold, the right entry still has to be found: 库长到一张菜单装不下时,还得让对的那条经验被找到:a growing library needs a finder, not a longer menu →经验越攒越多,靠的是会找,不是菜单更长 →
Whetstone ships as an Agent Skill plus plain scripts. No package manager, no service, no dependencies.Whetstone 就是一个 Agent Skill 加几个普通脚本。不用包管理器,不起服务,零依赖。
git clone https://github.com/TbusOS/whetstone.git~/.claude/skills/whetstone/The tool is public. Your distilled library — with real platform values — lives in a private repo. Keeping them apart means the public one never carries your internal data.工具是公开的;你蒸出来、带真实平台值的 skill 库放在私有仓。两者分开,公开仓就永远不会带出你的内部数据。
whetstone (the public tool) and your private library.whetstone(公开工具)和你的私有库。deploy.sh --link <library>Symlinks every skill into the skills folder — one source, no drift.把每个 skill 软链进 skills 目录,只有一份源,不会两边不一致。deploy.sh --gen-claudemd <library>Regenerates your always-on rules; --add-import wires them per project.重新生成常驻规则;每个项目再用 --add-import 接上。Capture the reusable part once. Prove it. Carry it to every platform and every teammate.把能复用的那部分捕获一次、证明一次,带到每个平台、每个同事那里。