Self-contained · zero runtime dependency自包含 · 零运行时依赖

Distill the session.
Refuse what it
can't prove.
把会话蒸馏成 skill,
再把证不出来的
挡在门外。

Whetstone turns a finished dev session — transcript, git diff, the pitfalls you hit — into a portable Agent Skill. Then it holds that skill to an evidence standard you can actually run. Whetstone 把一次做完的开发 —— transcript、git diff、踩过的坑 —— 蒸馏成可移植的 Agent Skill,再用一条能跑的命令守住证据纪律。

Claude Code · Codex · Cursor · any skills-compatible runtime任何支持 skills 的 runtime

4layers, by how far knowledge carries层,按知识能走多远来分
38mechanical checks in verifyverify 里的机械检查
2human review gates道人审关
0runtime dependencies运行时依赖
01 · Why为什么

Every session ends. What you learned usually ends with it.会话一结束,学到的东西大多跟着没了。

reusable knowledge you still have 手里还剩多少能复用的经验
━ with whetstone: the floor rises用 whetstone:底线一次比一次高   ━ without: back to zero不用:每次归零 illustration, not measured data示意图,不是测量数据
  1. 1

    New platform, same explanation — again.换个平台,又从头讲一遍。

    The design and pitfalls live in the skill body. Addresses, key lengths, tool names live in one params/<platform>.md. A new chip is a new table, nothing else.设计和坑留在 skill 正文;地址、密钥长度、工具名放进一张 params/<platform>.md。换芯片就新加一张表,别的不动。

  2. 2

    Your teammate can't get it at all.同事根本拿不到。

    The output is a plain-markdown folder with no runtime binding. Hand it to someone on Cursor or Codex; they drop it in and have your know-how.产物是一个纯 markdown 目录,不绑定任何 runtime。递给用 Cursor 或 Codex 的同事,放进去就能用。

  3. 3

    A better fix on B never reaches A.在 B 上学到的,A 永远用不上。

    Each new session is reconciled against the library: duplicates merge, changed facts are superseded, evidence for old entries accumulates. The library grows instead of resetting.每次新蒸馏都和已有库比对:重复的合并,变了的事实标注替代,旧条目的证据继续累加。库在长,而不是每次归零。

02 · Evidence discipline证据纪律

A rule only the model obeys is one context window from gone.只靠模型自觉遵守的规则,离「不存在」只差一个上下文窗口。

A skill library doesn't fail loudly. It fills up with plausible, once-true advice that still wears a confident label. So every rule a machine can decide became a check with an exit code. Below is a real run on examples/demo-skill, a deliberately broken package that ships with the repo.skill 库不会大声地坏掉,它会慢慢塞满听着有理、只在当时成立、却还顶着高置信标签的条目。所以凡是机器能判的规则,都做成了带退出码的检查。下面是对仓库自带的、故意写坏的 examples/demo-skill 的一次真实运行。

~/whetstone — zsh

      
38checks, grouped by what they guard条检查,按守的是什么分组
error · exit 1 warn · exit 1 with --strictwarn · 加 --strict 才 exit 1 info

Each of the 38 codes is tested in both directions — it fires on a bad package and stays quiet on a good one — and a mutation test deletes checks one at a time to prove the suite notices: 14 of 14 caught.38 个检查码每一个都正反两个方向测过:坏包上必须报、好包上必须不报。另有变异测试逐条删掉检查,确认测试套件会发现:14 条全部抓到。

Confidence is a table, not a feeling.置信度是一张表,不是一种感觉。

Try it. This is the exact rule verify enforces (extraction-framework §7). A reviewer may always lower a level; nobody may raise it past the conditions.自己试试。这就是 verify 执行的那条规则(extraction-framework §7)。人审永远可以往下调,但谁都不能越过条件往上调。

What kind of entry?这是哪一层的条目?
Executable check (验证方式)可执行检查(验证方式)
Seen on how many distinct platforms / projects?在几个不同的平台 / 项目上复现过?
Highest level allowed最高允许标到
LOW
MED
HIGH

V16 · §7REFUSED

Counting one project twice同一个项目算两次

Reproduction lines are keyed by platform/project. Ten runs on one project is still one line — otherwise the write-back becomes a back door around the gate.复现记录按平台/项目做唯一键。同一个项目跑十次还是一行,否则「回写证据」就成了绕过升级条件的后门。

V22 · §11REFUSED

Overwriting a fact in silence悄悄改掉一个事实

A changed value is appended and marked superseded; the old one stays readable. Being quietly wrong is the expensive kind of wrong, most of all for irreversible values.变了的值只能追加并标注替代,旧值仍然可读。错得没人发现是代价最大的错法,不可逆的值尤其如此。

PHASE 4REFUSED

Grading its own homework自己给自己判卷

Review runs in a separate context. Self-review bias is something this design assumes will happen, not something it hopes to avoid.质检在独立的上下文里跑。自评偏差是这套设计预设一定会发生的事,不是指望能躲开的事。

--explainWON'T FAKE

Pretending it can judge everything假装什么都能判

Whether an L1 truly has no counter-example stays with the reviewer, and --explain prints that boundary in full. A mechanised non-mechanical judgement only makes false positives.一条 L1 是不是真的举不出反例,留给人审,--explain 把这条边界原样打印出来。硬把不机械的判断做成机械的,只会制造误报。

What happened when I ran it on my own library: 拿它跑我自己的库,结果是:61 entries marked “high confidence”. Not one could name a test. →61 条标着「高置信度」的经验,没有一条说得出验证方式 →

And after an entry is in the library, it is checked again each time it is used: 入库之后,每次用到一条经验还会再核一次:check the note when you use it, not when you file it →经验对不对,用的时候才判得准 →

03 · The honest limit它管不到的地方

A green run is not a good skill.verify 全绿,不等于 skill 好用。

Every check asks whether the knowledge is true and traceable. None asks whether installing it makes its user better — and the two don't predict each other. That is why the human gates stay. Source: SkillLens (Microsoft).每条检查问的都是「这条知识真不真、能不能追溯」,没有一条问「装上它,使用者是不是变强了」,而这两件事互相预测不了。所以人审关一直保留。数据来源:SkillLens(Microsoft)。

25%

of extractor/consumer pairings in the SkillLens corpus transferred negatively — the skill made things worse.SkillLens 语料里,25% 的「抽取方 × 使用方」组合出现负迁移 —— 装上 skill 反而变差。

46.4%

is how often an unguided model judge picks the better of two skills. A coin does 50%.无引导的模型评委在两个 skill 里挑对更好那个的比例。掷硬币是 50%。

04 · The ruler一把尺子

One question sorts everything: how far does it carry?只问一个问题:这条知识能走多远?

Think of it as strata. The deeper a piece of knowledge sits, the more platforms it survives — and the closer it lives to the reusable skill body. Click a layer.把它想成岩层:越往下的知识,换越多平台仍然成立,也越靠近可复用的 skill 正文。点任意一层看看。

this session只在这次every platform任何平台
A pitfall is always split in two. The lesson (“don't infer hardware state from a software flag”) goes to L2; the value it hinged on (“the flag lives at OTP[ADDR]”) goes to L3. Unsplit pitfalls are refused (V20). 一个坑一定拆成两半。教训(「别从软件 flag 推硬件状态」)进 L2;它依赖的那个值(「flag 在 OTP[ADDR]」)进 L3。不拆的坑会被拒收(V20)。
05 · How it works怎么工作

Mine. Layer. Reconcile. You approve.挖掘,分层,比对,你点头。

Two human gates bracket an automatic middle. Nothing reaches your library without your sign-off.两道人审关夹着中间的自动流程。没有你点头,什么都进不了正式库。

06 · Architecture架构

A distiller, a portable package, optional extras.一个蒸馏器,一个可移植的包,几个可选的下游。

Runtime-specific code is fenced into one thin capture adapter. What you share is plain markdown. Everything below the dashed line is optional — installed, it helps; absent, nothing breaks.跟 runtime 绑定的代码只放在薄薄一层采集适配器里。对外分享的是纯 markdown。虚线以下全是可选:装了能增强,不装也不影响使用。

Distiller蒸馏器runtime-neutral method跟 runtime 无关的方法
extraction-frameworkthe L1–L4 rulesL1–L4 分层规则
SKILL.mdPhase 0–5 distillerPhase 0–5 蒸馏流程
verify · lintchecks that exit non-zero会返回非零的检查
find · routefinds the right skill once the menu is full菜单装不下之后,帮着找到对的 skill
capture/<runtime>the only runtime-specific part唯一跟 runtime 绑定的部分
produces产出
Portable skill package可移植 skill 包the unit you share对外分享的就是它
SKILL.mdL1 principles + L2 methodsL1 原理 + L2 做法
params/<platform>.mdL3 values — swap per platformL3 平台值,换平台只换它
pitfalls.mdL2 pitfall libraryL2 坑库
optional sync可选同步
Optional extras可选下游install-to-use装了才用
engramlocal memory store: recall, supersede本地记忆库:召回、替代
llm-wikia wiki page for humans to read发成给人读的 wiki 页
darwin-skillscore and polish one skill (3rd-party)给单个 skill 打分打磨(第三方)
07 · One real scenario一个真实场景

Secure boot, done once — owned by the whole team.安全启动做一次,整个团队都会。

  1. 1

    You finish bring-up on one SoC你在一颗 SoC 上做完安全启动

    Keys, OTP, the signing chain — and a buffer overrun you only found by dumping bytes.密钥、OTP、签名链,还有一个靠 dump 字节才揪出来的缓冲区越界。

  2. 2

    Run the distiller跑一次蒸馏

    It separates the transferable design and pitfall (L1/L2) from this chip's addresses and key length (L3). You approve at two gates.它把可迁移的设计和坑(L1/L2)跟这颗芯片的地址、密钥长度(L3)分开。你在两道关点头。

  3. 3

    Out comes verified-boot/产出 verified-boot/

    Plain markdown, self-contained, and verify is clean.纯 markdown,自包含,verify 通过。

  4. 4

    A teammate on Cursor drops it in用 Cursor 的同事把它放进去

    They now have the design, the method, and the byte-dump pitfall. No onboarding call.设计、做法、那个 dump 字节的坑,他马上就有了。不用专门开会讲。

  5. 5

    Next SoC? Add one table.下一颗 SoC?加一张表。

    params/soc-b.md. The design and pitfalls carry over untouched.params/soc-b.md。设计和坑原样复用。

08 · Index hygiene索引卫生

More skills shouldn't mean a noisier menu.skill 越多,菜单不该越乱。

A runtime keeps every skill's name + description in context to choose between them. Bodies load only after a pick. So the noise lives in the menu — and that's where lint looks.runtime 把每个 skill 的 name + description 常驻在上下文里用来挑选,正文要等被选中才加载。所以乱就乱在菜单上,lint 查的也正是这里。

SKILL BODYskill 正文loaded only when picked选中才加载
whetstone lint

Finds missing triggers, overlapping trigger sets and near-identical names — tells deliberate companions (-audit) from real clashes — and checks whether the whole menu still fits.找出缺触发词、触发词重叠、名字几乎一样的 skill,能分清故意成对的(-audit)和真撞车,还会查整张菜单是不是还放得下。

whetstone index

Groups skills into families and writes a browsable INDEX.md.按族归组,生成一份能浏览的 INDEX.md。

description contractdescription 写法约定

Capability line, concrete triggers, and a boundary that pushes siblings apart — baked into the template.能力一句话 + 具体触发词 + 跟同族划清边界,写进模板,新 skill 一生成就符合。

And once the library outgrows what one menu can hold, the right entry still has to be found: 库长到一张菜单装不下时,还得让对的那条经验被找到:a growing library needs a finder, not a longer menu →经验越攒越多,靠的是会找,不是菜单更长 →

09 · Install安装

Drop the folder into any runtime.把目录放进任何 runtime。

Whetstone ships as an Agent Skill plus plain scripts. No package manager, no service, no dependencies.Whetstone 就是一个 Agent Skill 加几个普通脚本。不用包管理器,不起服务,零依赖。

git clone https://github.com/TbusOS/whetstone.git
Claude Code~/.claude/skills/whetstone/
Codex · Cursorthe matching skills / rules folder对应的 skills / rules 目录
10 · Move machines换机迁移

New laptop? Clone two repos and you're back.换电脑?克隆两个仓,能力全回来。

The tool is public. Your distilled library — with real platform values — lives in a private repo. Keeping them apart means the public one never carries your internal data.工具是公开的;你蒸出来、带真实平台值的 skill 库放在私有仓。两者分开,公开仓就永远不会带出你的内部数据。

whetstone public · tool my-skills private · data new machine deploy.sh --link skills/ ✓ rules ✓
  1. Clone both repos两个仓都 clone 下来whetstone (the public tool) and your private library.whetstone(公开工具)和你的私有库。
  2. deploy.sh --link <library>Symlinks every skill into the skills folder — one source, no drift.把每个 skill 软链进 skills 目录,只有一份源,不会两边不一致。
  3. deploy.sh --gen-claudemd <library>Regenerates your always-on rules; --add-import wires them per project.重新生成常驻规则;每个项目再用 --add-import 接上。

Read the CLAUDE.md layering guide →读 CLAUDE.md 分层设计指南 →

Your next feature下一个功能

Stop letting it evaporate.别再让经验蒸发了。

Capture the reusable part once. Prove it. Carry it to every platform and every teammate.把能复用的那部分捕获一次、证明一次,带到每个平台、每个同事那里。