The budget is fixed. The library is not.预算是固定的,库是一直长的。
Claude Code gives the skill menu roughly one percent of the context window. When the library outgrows it, no skill is removed; the least-used ones lose their descriptions and stay as a bare name. That is where a skill stops being picked on topic: on one machine over 30 days, a trigger word in a message led to the skill 26% of the time when its description was on the menu, and 1% of the time when only its name was (98 and 99 matching messages).Claude Code 给 skill 菜单大约上下文的百分之一。库超出这个预算时,它不删 skill,而是把用得最少的那些描述去掉,只留名字。skill 从这一步开始就很难按话题被挑中:一台机器上 30 天的记录里,消息里出现触发词时,菜单上有描述的 skill 被用上 26%,只剩名字的 1%(分别 98 条、99 条命中消息)。
Shortening descriptions only delays the outcome. Split a fixed budget over more skills and every share shrinks, until a description is too short to match anything. The entry point cannot grow with the library. It has to be a fixed-size entry plus a search, or a signal that brings a skill out without going through the menu at all: the folder you are working in, the file you just opened.缩短描述只是推迟这个结局。固定的预算分给越来越多的 skill,每条的份额越来越小,直到短得什么都匹配不上。入口不能跟着库一起变长,只能是「固定大小的入口 + 查找」,或者干脆不经过菜单,由确定的信号直接带出来:你在哪个目录干活,你刚打开了哪个文件。
Shares are the library’s part of the budget divided evenly, after about 20 characters of name and punctuation per entry. The budget follows the model: a smaller context gets a smaller menu, so the same library runs out sooner.份额 = 留给库的预算平均分,每条先扣掉约 20 字符的名字和格式。预算跟着模型走:上下文小,菜单就小,同一个库更早装不下。
“I don’t want them called by hand. There are so many skills that nobody knows which ones exist.”「不希望手动调用呀,因为 skills 非常多,根本不知道有哪些 skills 的。」
The library’s owner, on why manual-only skills were ruled out, 2026-09-29库的主人,解释为什么不用「只许手动调用」,2026-09-29
Keep the menu short, ask when unsure, learn afterwards.菜单保持精简,拿不准就查,事后再学。
Certain signals come first: the folder and the file are facts, “the model read a description and thought it fit” is a guess. On one machine, a rule tied to the working folder led to the right skill in 8 of 10 sessions; trigger words in descriptions, in 5 of 38.确定的信号优先:目录、文件是事实,「模型读了描述觉得相关」是推断。一台机器上,按工作目录写的规则 10 个会话里 8 个用上了对应 skill;按描述里的触发词,38 组里 5 组。
Routing skill路由 skill
one short entry; the whole catalog sits behind it菜单里只占一条,整个库的目录在它后面
Yield by project按项目让位
used elsewhere, never here: name only, still callable用在别处、本项目没用过:只留名字,照样能调用
from usage来自使用Appear by file按文件出现
joins the menu when a matching file is read读到匹配的文件,才加进菜单
measured已实测Project-only项目专属
lives in that project’s own skills folder放在那个项目自己的 skill 目录里
whetstone find "<what you are about to do><要做的事>"
name ×3 · description ×2 · body ×1 · learned messages ×2名字 ×3 · 描述 ×2 · 正文 ×1 · 学到的原话 ×2
learned学来的Below the threshold够不上门槛
says “no skill fits”直说「没有合适的」
Any agent, a shell有 shell 就能跑
the one entry all can run各家 agent 都能用的入口
Session logs会话记录
each runtime’s format各家自己的格式
Adapter适配器
one per runtime每个运行时一个
Neutral events统一事件
6 kinds, one per line6 种,一行一条
learn · usage学原话 · 使用表
what led to each skill什么话引出了哪条
replay回放
kept only if it wins赢了才保留改动
Gold cards are the runtime’s own switches, so they cost nothing to run; the orange card is the part whetstone adds; green tags mark what the learning layer feeds. The two “always on” switches in the middle were checked in real sessions, not only read from the program.金色卡片是运行时自带的开关,本身不花任何运行成本;橙色卡片是 whetstone 加的部分;绿色小标签标出学习层喂给谁。中间两个常驻开关都在真实会话里验过,不只是读程序得出的。
What each part does, and why it is built that way.每个部分做什么,为什么这样做。
One entry for the whole library整个库只占一条
A generated routing skill: a description under 200 characters, a catalog grouped by family behind it. A new skill has no usage, so the runtime cuts its description first; the same sentence also goes into the always-loaded rules file, which the menu budget does not touch.自动生成的路由 skill:描述不到 200 字符,后面是按族分组的目录。新 skill 没有调用记录,超预算时最先被砍描述,所以同一句话也写进常驻规则文件,那里不受菜单预算影响。
whetstone route catalogYield, don’t hide让位,不藏起来
A skill whose use sits in another project (at least 3 sessions, two thirds of them elsewhere, none here) is set to name-only for this project. It stays callable by name. The rule was fixed before looking at the data; on one machine it currently suggests nothing.用得集中在别的项目(至少 3 个会话、三分之二在别处、本项目一次没用)的 skill,在本项目里设成只留名字,照样能按名字调用。规则是看数据之前定的;一台机器上目前一条建议都没有。
whetstone route plan · prints unless --write不加 --write 只打印Appear when the file does碰到文件才出现
A skill tied to a kind of file gets a paths glob. It is absent from the menu at the start, and joins it the moment a matching file is read, so it costs nothing until then. The trade: a conversation that never touches such a file will not see it.和某类文件绑定的 skill 加上 paths。开局不在菜单里,读到匹配的文件那一刻才加进去,在那之前不占预算。代价:从头到尾没碰这类文件的对话,看不到它。
Search the body, not just the name正文也要搜
BM25F over four fields. The body counts because the menu text alone is not enough: dropping skill bodies took top-1 accuracy from 56.0% to 18.7% in SkillRouter. Chinese is split into pairs of characters, because it has no spaces and single characters match almost anything.四个字段的 BM25F。正文要算进去,因为只看菜单那点文字不够:SkillRouter 去掉正文后首位命中从 56.0% 掉到 18.7%。中文按相邻两个字切,因为中文没有空格,单字几乎什么都能配上。
whetstone findLearn how you actually ask学你平时怎么说
The message you typed right before a skill was first used in a session is added to that skill’s index. How people phrase a request and how a tool describes itself share few words (0.06 overlap in ToolRet), so the words that led to a skill before are the best hint for next time.每个会话里某条 skill 第一次被用之前你打的那句话,并进这条 skill 的索引。人的说法和工具的自我描述很少共用词(ToolRet 测的词面重合 0.06),所以以前引出过它的那句话,是下次最好的线索。
whetstone route learnA change must win the replay改动要在回放里赢
Past sessions are replayed in time order, each scored only with what was learned before it. A change is kept at 5 wins and 0 losses, or 7 and 1, with wins from two projects or more. Otherwise the verdict is “not enough evidence”, which is not a failure.历史会话按时间顺序回放,每个会话只用它之前学到的东西打分。5 胜 0 负或 7 胜 1 负、胜例来自至少 2 个项目,才保留;否则判「证据不足」,这不算失败。
whetstone route replayA skill that waits for its file.一条等文件出现的 skill。
session starts会话开始
the conditional skill is not listed这条条件 skill 不在菜单里
reads notes.txt读 notes.txt
unchanged: the glob does not match不变:路径不匹配
reads board/x.dts读 board/x.dts
added mid-session会话中途加进来
Two headless Claude Code sessions with hooks off, 2026-09-30, skill frontmatter paths: ["**/*.dts"]. A third check found where the per-project switch is read: inside a git repo, from its top level, subfolders included; outside one, only in the folder the session starts in.两个无交互的 Claude Code 测试会话(关掉钩子),2026-09-30,skill 的 frontmatter 写 paths: ["**/*.dts"]。另一组实测查了按项目的开关从哪读:在 git 仓里读仓根,子目录启动也生效;不在 git 仓,只读启动的那个目录。
An answer, or a plain “no”.要么给答案,要么直说没有。
Handing an agent the wrong skill is worse than handing it none: in SkillLens a quarter of the pairings transferred negatively, and in SkillsBench 16 of 84 tasks got worse with a skill added. So find prints at most three, and only those that cover enough of the question.给 agent 一条错的 skill,比什么都不给更糟:SkillLens 里四分之一的组合是负迁移,SkillsBench 84 个任务里有 16 个加了 skill 反而变差。所以 find 最多给三条,而且只给覆盖了足够多问题的那些。
- Cover is the share of the question’s weight a skill matches. A word no skill contains still counts, or one lucky word would read as a full match.覆盖率 = 这条 skill 配上的查询词权重占全部查询词的比例。库里没有的词也要算进去,不然碰巧配上一个词就成了百分之百。
- The threshold, 0.3, is a starting value. The second query on the right is a real miss: “screen” and “tune” appear nowhere in that skill’s text.门槛 0.3 只是起点值。右边第二条是真实的漏判:「屏幕」「调参」在那条 skill 的文字里都没出现。
- From now on the adapter records every query an agent writes to find. Those, not your raw messages, are what the threshold will be set on.从现在起,适配器会记下 agent 每次给 find 写的查询。门槛以后用这些来定,而不是用你的原话。
$ whetstone find "AVB 签名 验签怎么配置"
1. secure-boot (covers 50%: avb 签名 验签) ~/skills/secure-boot/SKILL.md
Secure boot chain: signing, verification, AVB, fastboot flash checks. 触发词:验签, 签名, AVB.
find: 5 skills searched
$ whetstone find "屏幕闪烁 调参"
no skill fits well enough (best: panel-tuning covers 25%, needs 30%); go ahead without one
find: 5 skills searched
$ whetstone find "今天天气怎么样"
no skill fits well enough (nothing matched, needs 30%); go ahead without one
find: 5 skills searched--src pointed at it and learned messages off. Paths shortened.在为本页写的 5 个 skill 演示库上的真实输出(安全启动、分区表、屏幕调参、git 流程、版本归档),--src 指向它,关掉学到的原话。路径做了缩短。One core that never reads a runtime’s format.核心只有一份,不读任何一家的格式。
Six runtimes store sessions six different ways and switch skills with six different settings. What they all have is the same three facts: what the user said, which files were touched, which skill was used. The core learns, searches and scores from those facts alone; each runtime gets an adapter that extracts them and translates the results back into its own switches. Claude Code’s adapter is built and tested. The other five are designed from reading their source, and not written yet.六家运行时存会话的方式各不相同,开关 skill 的设置也各不相同。它们共有的是同样三件事:用户说了什么、碰了哪些文件、用了哪条 skill。核心只拿这三件事来学习、查找和打分;每个运行时配一个适配器,负责把这三件事取出来,再把结果翻译回它自己的开关。Claude Code 的适配器已经做好并测过;另外五家是读了它们的源码之后设计的,还没写。
Runtime · adapter运行时 · 适配器
Neutral events统一事件
Neutral events统一事件
one JSON object per line一行一个 JSON 对象- session · user
- file {path, op}
- load {via, use}
- menu · find
Core核心
bin/route.py
reads only the events只读统一事件- learnlearn 学原话what was said → which skill什么话 → 用了哪条
- usage · elsewherewho uses what, per project各项目用了哪些
- findfind 查找rank, then a threshold打分,再过门槛
- replayreplay 回放keep only what wins赢了才保留
Outputs产出
whetstone find
a shell command, so any of the six can run it一条 shell 命令,六家都能跑Menu switches各家的菜单开关
- Claude Code: overrides · pathsClaude Code:overrides · paths
- opencode: permission.skillopencode:permission.skill
- pi: .pi/settings.json skillspi:.pi/settings.json 的 skills
- Gemini CLI: skills.disabledGemini CLI:skills.disabled
- Codex: no per-project switchCodex:没有项目级开关
Pointer line常驻指针
CLAUDE.md · AGENTS.md · GEMINI.md~/.claude/skills → Claude Code · opencode · Cursor ~/.agents/skills → Codex · opencode · pi · Gemini CLI · Cursor (linking into the second: not done yet)(第二个目录还没接上)The grey notes under each unbuilt adapter are the main obstacle found in that runtime’s source, dated 2026-09-29/30. use is decided by the adapter because runtimes differ: in Claude Code, reading a SKILL.md is usually someone editing it; in Codex or pi, reading it is how a skill gets used.每个未实现适配器下面的灰字,是读那家源码(2026-09-29/30 的版本)时找到的主要难点。use(算不算在用)交给适配器判断,因为各家不同:Claude Code 里读 SKILL.md 多半是在改它,Codex 和 pi 里读 SKILL.md 就是在用。
What the six runtimes offer, side by side六家运行时能用上的东西,并排看
| Claude Code | Codex | opencode | pi | Gemini CLI | Cursor | |
|---|---|---|---|---|---|---|
| Adapter适配器 | built已实现 | not written未写 | not written未写 | not written未写 | not written未写 | not written未写 |
Reads ~/.agents/skills读 ~/.agents/skills | no否 | yes是 | yes是 | yes是 | yes是 | yes (docs)是(文档) |
| Menu limit菜单上限 | ~1% of context; cuts descriptions of the least used约 1% 上下文;砍用得最少的描述 | 2% of context, max 10k tokens; then drops whole entries2% 上下文、最多 1 万 token;再不够就整条删 | none无 | none无 | none无 | not documented文档没写 |
| Per-project switch项目级开关 | skillOverrides | project config ignored项目配置不生效 | permission.skill | .pi/settings.json | skills.disabled | not found未查到 |
| Appear by file按文件出现 | paths | — | — | — | — | paths (docs)(文档) |
| Menu kept in the session log会话记录里有菜单 | yes是 | yes是 | no否 | yes是 | no否 | no否 |
| Always-loaded file常驻指令文件 | CLAUDE.md | AGENTS.md | AGENTS.md | AGENTS.md | GEMINI.md | AGENTS.md |
From each runtime’s source at its 2026-09-29/30 head (Codex also from 14 local sessions); Cursor from its documentation only. The spec lists the file and line behind every cell.来自各家 2026-09-29/30 的源码(Codex 另有本机 14 份会话);Cursor 只有文档。规范里给出了每一格对应的文件和行号。
More hits is not the test. Winning case by case is.不看命中多了几个,看逐例是赢是输。
83 sessions on one machine gave 34 cases: a skill used for the first time in a session, and the message typed just before. Two rankings exist only to check the scorer: a shuffled one must lose, and one allowed to peek at the answer must win.一台机器上 83 个会话,得到 34 个例子:某会话里某条 skill 第一次被用,和那之前打的那句话。有两个排序只用来检查打分脚本本身:打乱的必须输,偷看了答案的必须赢。
The project boost has one more top-3 hit than learned messages alone, and was still dropped: compared case by case it moved the right skill down four times and up three. The learned messages won nine times, lost none, and the wins came from six projects. Each formula was written down before its first replay and tested once.「本项目用过就加分」比单用学到的原话多一个前 3 命中,照样不用:逐例比,它把对的那条往下挤了四次、往上提了三次。学到的原话赢九次、一次没输,胜例来自六个项目。每个公式都是第一次回放前写下的,只测一次。
Borrowed where it held up, left behind where it did not.能用的借过来,用不上的不借。
karpathy/autoresearch
Keep a change only when the frozen score improves. What it lacks was added: splitting by time, a paired case-by-case test, “not enough evidence” as its own verdict, and a scorer that has to beat a shuffle and lose to a peek.只在冻结的评分变好时保留改动。它缺的几样补上了:按时间切分、逐例配对比较、「证据不足」单独算一种结论、打分脚本必须赢过打乱的、输给偷看的。
openai/codex
Its skill selector runs a dozen cheap rankings (BM25F, character n-grams, recently used) in shadow mode each turn, logging where the skill the model really called would have ranked, without changing the prompt. Marked temporary in its source.它的 skill 选择实验每轮在后台跑十几种便宜的排序(BM25F、字符 n-gram、最近用过),只记录模型真正调用的那条会排第几,不改提示词。源码里标着是临时实验。
lomeshdutta/skill-router
A hook on every message was tried there and withdrawn in v0.2 for once per session: good evaluation scores, too noisy in real sessions (paraphrased). Here the hook count is zero unless a replay shows a median of one prompt per session or less.那里试过每条消息都跑的钩子,v0.2 撤回,改成每会话一次:评测分数好,真实会话里太吵(大意)。这里钩子数为零,除非回放证明每个会话提示次数的中位数不超过一次。
Picking the right tool怎么挑准工具
SkillRouter (arXiv 2603.22455): bodies matter, 56.0% → 18.7% top-1 without them. ToolRet: requests and tool docs share little wording, 0.06 overlap. SkillLens (arXiv 2605.23899) and SkillsBench: a wrong skill can do harm. All figures as reported by the authors.SkillRouter(arXiv 2603.22455):正文要紧,去掉后首位命中 56.0% → 18.7%。ToolRet:请求和工具文档很少共用词,重合 0.06。SkillLens(arXiv 2605.23899)与 SkillsBench:错的 skill 会帮倒忙。数字均为作者自报。
What is not done, and what the numbers cannot say.哪些还没做,数字说明不了什么。
Only Claude Code is wired up. The adapters for Codex, opencode, pi, Gemini CLI and Cursor are designed from their source and not written; linking the library into ~/.agents/skills and checking Codex’s menu limit at install time are not done either. whetstone find itself runs anywhere there is a shell and Python.目前只接好了 Claude Code。Codex、opencode、pi、Gemini CLI、Cursor 的适配器是按源码设计的,还没写;把库链接到 ~/.agents/skills、装库时按 Codex 的规则算菜单上限,也都还没做。whetstone find 本身在任何有 shell 和 Python 的地方都能跑。
Every number comes from one machine and one user. The replay uses today’s skill text, and some descriptions gained trigger words only after they were missed, so “text only” reads high. The threshold is not set yet.所有数字来自一台机器、一个使用者。回放用的是 skill 现在的文字,有些描述是在当时没被用上之后才补的触发词,所以「只看文字」那一行偏高。门槛还没定下来。
And find still depends on the agent remembering to run it. The always-loaded line and the paths switch make that likelier; nothing makes it certain, and a replay cannot measure it.另外,find 还是要靠 agent 想起来去跑。常驻那一行和 paths 让它更可能被想起来,但没有办法保证,回放也测不出这一步。
The parts you can check.能核对的那些部分。
whetstone find · route events · route learn · route usage · route replay · route elsewhere · route plan · route catalog
Every line carries v, runtime, session, ts, and optionally src (log file and line). Kinds: session, user, file, load, menu, find. A new runtime needs only an adapter that writes these.每行都带 v、runtime、session、ts,可选 src(会话文件和行号)。种类:session、user、file、load、menu、find。接一个新运行时,只要写一个产出这些行的适配器。
Learned messages are your own words. They stay on your machine (~/.local/share/whetstone/) and never go into a portable skill package.学到的原话是你自己打的字,只存在本机(~/.local/share/whetstone/),不进可移植的 skill 包。
73 selftest checks, each rule in both directions, plus a final check that nothing outside the test folder was written. 57 mutations: each removes one rule and the suite must go red.73 项自检,每条规则正反两个方向都测,最后还查测试目录以外有没有文件被写。57 条故意改坏:每条拿掉一条规则,整套测试必须报红。
spec/routing.md — the three layers, the event format, the six-runtime table with file and line, the replay rules, and what may never be done.—— 三层结构、事件格式、带文件和行号的六家对照、回放规则和禁止清单。
whetstone is MIT, standard library only. Routing ships in the same command line as distilling, verifying and the use-time check.whetstone 是 MIT 协议,只用标准库。经验路由和蒸馏、校验、用时纠错在同一个命令行里。