← Whetstone overview← 返回 Whetstone 总览
Feature · Experience routing功能 · 经验路由

A growing library needs a finder, not a longer menu.经验越攒越多,靠的是会找,不是菜单更长。

An agent picks a skill by reading a menu of names and descriptions, and that menu has a fixed budget. Past it, descriptions are cut and the trigger words go with them. whetstone keeps the menu short with switches the runtime already has, finds the rest with one command any agent can run, and learns from past sessions how you ask. No hooks. Built for Claude Code first, designed for the other runtimes. agent 挑 skill 靠读一张「名字 + 描述」的菜单,这张菜单有固定预算,超了就砍描述,触发词跟着没了。whetstone 用运行时自带的开关让菜单保持精简,其余的用一条任何 agent 都能跑的命令去找,再从历史会话里学你平时怎么说。不装钩子。先在 Claude Code 上做完,其他运行时按同一套设计接入。

35/65skills on one machine’s menu shown as a bare name, description cut个 skill 在一台机器的菜单里只剩名字,描述被砍掉
1%how often a trigger word still led to the skill once its description was cut; 26% while it was there描述被砍后,触发词出现时 skill 还被用上的比例;有描述时是 26%
0hooks installed; everything learned is read from past sessions afterwards个钩子;学的东西全部事后从历史会话里读
The premise前提

The budget is fixed. The library is not.预算是固定的,库是一直长的。

Claude Code gives the skill menu roughly one percent of the context window. When the library outgrows it, no skill is removed; the least-used ones lose their descriptions and stay as a bare name. That is where a skill stops being picked on topic: on one machine over 30 days, a trigger word in a message led to the skill 26% of the time when its description was on the menu, and 1% of the time when only its name was (98 and 99 matching messages).Claude Code 给 skill 菜单大约上下文的百分之一。库超出这个预算时,它不删 skill,而是把用得最少的那些描述去掉,只留名字。skill 从这一步开始就很难按话题被挑中:一台机器上 30 天的记录里,消息里出现触发词时,菜单上有描述的 skill 被用上 26%,只剩名字的 1%(分别 98 条、99 条命中消息)。

Shortening descriptions only delays the outcome. Split a fixed budget over more skills and every share shrinks, until a description is too short to match anything. The entry point cannot grow with the library. It has to be a fixed-size entry plus a search, or a signal that brings a skill out without going through the menu at all: the folder you are working in, the file you just opened.缩短描述只是推迟这个结局。固定的预算分给越来越多的 skill,每条的份额越来越小,直到短得什么都匹配不上。入口不能跟着库一起变长,只能是「固定大小的入口 + 查找」,或者干脆不经过菜单,由确定的信号直接带出来:你在哪个目录干活,你刚打开了哪个文件。

Shares are the library’s part of the budget divided evenly, after about 20 characters of name and punctuation per entry. The budget follows the model: a smaller context gets a smaller menu, so the same library runs out sooner.份额 = 留给库的预算平均分,每条先扣掉约 20 字符的名字和格式。预算跟着模型走:上下文小,菜单就小,同一个库更早装不下。

“I don’t want them called by hand. There are so many skills that nobody knows which ones exist.”「不希望手动调用呀,因为 skills 非常多,根本不知道有哪些 skills 的。」

The library’s owner, on why manual-only skills were ruled out, 2026-09-29库的主人,解释为什么不用「只许手动调用」,2026-09-29
Three layers三层结构

Keep the menu short, ask when unsure, learn afterwards.菜单保持精简,拿不准就查,事后再学。

Certain signals come first: the folder and the file are facts, “the model read a description and thought it fit” is a guess. On one machine, a rule tied to the working folder led to the right skill in 8 of 10 sessions; trigger words in descriptions, in 5 of 38.确定的信号优先:目录、文件是事实,「模型读了描述觉得相关」是推断。一台机器上,按工作目录写的规则 10 个会话里 8 个用上了对应 skill;按描述里的触发词,38 组里 5 组。

Gold cards are the runtime’s own switches, so they cost nothing to run; the orange card is the part whetstone adds; green tags mark what the learning layer feeds. The two “always on” switches in the middle were checked in real sessions, not only read from the program.金色卡片是运行时自带的开关,本身不花任何运行成本;橙色卡片是 whetstone 加的部分;绿色小标签标出学习层喂给谁。中间两个常驻开关都在真实会话里验过,不只是读程序得出的。

Six parts六个部分

What each part does, and why it is built that way.每个部分做什么,为什么这样做。

Always on常驻

One entry for the whole library整个库只占一条

A generated routing skill: a description under 200 characters, a catalog grouped by family behind it. A new skill has no usage, so the runtime cuts its description first; the same sentence also goes into the always-loaded rules file, which the menu budget does not touch.自动生成的路由 skill:描述不到 200 字符,后面是按族分组的目录。新 skill 没有调用记录,超预算时最先被砍描述,所以同一句话也写进常驻规则文件,那里不受菜单预算影响。

whetstone route catalog
Always on常驻

Yield, don’t hide让位,不藏起来

A skill whose use sits in another project (at least 3 sessions, two thirds of them elsewhere, none here) is set to name-only for this project. It stays callable by name. The rule was fixed before looking at the data; on one machine it currently suggests nothing.用得集中在别的项目(至少 3 个会话、三分之二在别处、本项目一次没用)的 skill,在本项目里设成只留名字,照样能按名字调用。规则是看数据之前定的;一台机器上目前一条建议都没有。

whetstone route plan · prints unless --write不加 --write 只打印
Always on常驻

Appear when the file does碰到文件才出现

A skill tied to a kind of file gets a paths glob. It is absent from the menu at the start, and joins it the moment a matching file is read, so it costs nothing until then. The trade: a conversation that never touches such a file will not see it.和某类文件绑定的 skill 加上 paths。开局不在菜单里,读到匹配的文件那一刻才加进去,在那之前不占预算。代价:从头到尾没碰这类文件的对话,看不到它。

Claude Code, Cursor (docs)Claude Code、Cursor(文档)
On demand按需

Search the body, not just the name正文也要搜

BM25F over four fields. The body counts because the menu text alone is not enough: dropping skill bodies took top-1 accuracy from 56.0% to 18.7% in SkillRouter. Chinese is split into pairs of characters, because it has no spaces and single characters match almost anything.四个字段的 BM25F。正文要算进去,因为只看菜单那点文字不够:SkillRouter 去掉正文后首位命中从 56.0% 掉到 18.7%。中文按相邻两个字切,因为中文没有空格,单字几乎什么都能配上。

whetstone find
Learns学习

Learn how you actually ask学你平时怎么说

The message you typed right before a skill was first used in a session is added to that skill’s index. How people phrase a request and how a tool describes itself share few words (0.06 overlap in ToolRet), so the words that led to a skill before are the best hint for next time.每个会话里某条 skill 第一次被用之前你打的那句话,并进这条 skill 的索引。人的说法和工具的自我描述很少共用词(ToolRet 测的词面重合 0.06),所以以前引出过它的那句话,是下次最好的线索。

whetstone route learn
Learns学习

A change must win the replay改动要在回放里赢

Past sessions are replayed in time order, each scored only with what was learned before it. A change is kept at 5 wins and 0 losses, or 7 and 1, with wins from two projects or more. Otherwise the verdict is “not enough evidence”, which is not a failure.历史会话按时间顺序回放,每个会话只用它之前学到的东西打分。5 胜 0 负或 7 胜 1 负、胜例来自至少 2 个项目,才保留;否则判「证据不足」,这不算失败。

whetstone route replay
Measured, not read实测,不是读来的

A skill that waits for its file.一条等文件出现的 skill。

Two headless Claude Code sessions with hooks off, 2026-09-30, skill frontmatter paths: ["**/*.dts"]. A third check found where the per-project switch is read: inside a git repo, from its top level, subfolders included; outside one, only in the folder the session starts in.两个无交互的 Claude Code 测试会话(关掉钩子),2026-09-30,skill 的 frontmatter 写 paths: ["**/*.dts"]。另一组实测查了按项目的开关从哪读:在 git 仓里读仓根,子目录启动也生效;不在 git 仓,只读启动的那个目录。

On demand按需

An answer, or a plain “no”.要么给答案,要么直说没有。

Handing an agent the wrong skill is worse than handing it none: in SkillLens a quarter of the pairings transferred negatively, and in SkillsBench 16 of 84 tasks got worse with a skill added. So find prints at most three, and only those that cover enough of the question.给 agent 一条错的 skill,比什么都不给更糟:SkillLens 里四分之一的组合是负迁移,SkillsBench 84 个任务里有 16 个加了 skill 反而变差。所以 find 最多给三条,而且只给覆盖了足够多问题的那些。

  • Cover is the share of the question’s weight a skill matches. A word no skill contains still counts, or one lucky word would read as a full match.覆盖率 = 这条 skill 配上的查询词权重占全部查询词的比例。库里没有的词也要算进去,不然碰巧配上一个词就成了百分之百。
  • The threshold, 0.3, is a starting value. The second query on the right is a real miss: “screen” and “tune” appear nowhere in that skill’s text.门槛 0.3 只是起点值。右边第二条是真实的漏判:「屏幕」「调参」在那条 skill 的文字里都没出现。
  • From now on the adapter records every query an agent writes to find. Those, not your raw messages, are what the threshold will be set on.从现在起,适配器会记下 agent 每次给 find 写的查询。门槛以后用这些来定,而不是用你的原话。
$ whetstone find "AVB 签名 验签怎么配置"
1. secure-boot  (covers 50%: avb 签名 验签)  ~/skills/secure-boot/SKILL.md
   Secure boot chain: signing, verification, AVB, fastboot flash checks. 触发词:验签, 签名, AVB.
find: 5 skills searched

$ whetstone find "屏幕闪烁 调参"
no skill fits well enough (best: panel-tuning covers 25%, needs 30%); go ahead without one
find: 5 skills searched

$ whetstone find "今天天气怎么样"
no skill fits well enough (nothing matched, needs 30%); go ahead without one
find: 5 skills searched
Real output on a five-skill demo library written for this page (secure boot, partition table, panel tuning, git workflow, release archive), with --src pointed at it and learned messages off. Paths shortened.在为本页写的 5 个 skill 演示库上的真实输出(安全启动、分区表、屏幕调参、git 流程、版本归档),--src 指向它,关掉学到的原话。路径做了缩短。
Across runtimes多种 agent

One core that never reads a runtime’s format.核心只有一份,不读任何一家的格式。

Six runtimes store sessions six different ways and switch skills with six different settings. What they all have is the same three facts: what the user said, which files were touched, which skill was used. The core learns, searches and scores from those facts alone; each runtime gets an adapter that extracts them and translates the results back into its own switches. Claude Code’s adapter is built and tested. The other five are designed from reading their source, and not written yet.六家运行时存会话的方式各不相同,开关 skill 的设置也各不相同。它们共有的是同样三件事:用户说了什么、碰了哪些文件、用了哪条 skill。核心只拿这三件事来学习、查找和打分;每个运行时配一个适配器,负责把这三件事取出来,再把结果翻译回它自己的开关。Claude Code 的适配器已经做好并测过;另外五家是读了它们的源码之后设计的,还没写。

The grey notes under each unbuilt adapter are the main obstacle found in that runtime’s source, dated 2026-09-29/30. use is decided by the adapter because runtimes differ: in Claude Code, reading a SKILL.md is usually someone editing it; in Codex or pi, reading it is how a skill gets used.每个未实现适配器下面的灰字,是读那家源码(2026-09-29/30 的版本)时找到的主要难点。use(算不算在用)交给适配器判断,因为各家不同:Claude Code 里读 SKILL.md 多半是在改它,Codex 和 pi 里读 SKILL.md 就是在用。

What the six runtimes offer, side by side六家运行时能用上的东西,并排看

Claude CodeCodexopencodepiGemini CLICursor
Adapter适配器built已实现not written未写not written未写not written未写not written未写not written未写
Reads ~/.agents/skills读 ~/.agents/skillsno否yes是yes是yes是yes是yes (docs)是(文档)
Menu limit菜单上限~1% of context; cuts descriptions of the least used约 1% 上下文;砍用得最少的描述2% of context, max 10k tokens; then drops whole entries2% 上下文、最多 1 万 token;再不够就整条删none无none无none无not documented文档没写
Per-project switch项目级开关skillOverridesproject config ignored项目配置不生效permission.skill.pi/settings.jsonskills.disablednot found未查到
Appear by file按文件出现paths————paths (docs)(文档)
Menu kept in the session log会话记录里有菜单yes是yes是no否yes是no否no否
Always-loaded file常驻指令文件CLAUDE.mdAGENTS.mdAGENTS.mdAGENTS.mdGEMINI.mdAGENTS.md

From each runtime’s source at its 2026-09-29/30 head (Codex also from 14 local sessions); Cursor from its documentation only. The spec lists the file and line behind every cell.来自各家 2026-09-29/30 的源码(Codex 另有本机 14 份会话);Cursor 只有文档。规范里给出了每一格对应的文件和行号。

The first replay第一次回放

More hits is not the test. Winning case by case is.不看命中多了几个,看逐例是赢是输。

83 sessions on one machine gave 34 cases: a skill used for the first time in a session, and the message typed just before. Two rankings exist only to check the scorer: a shuffled one must lose, and one allowed to peek at the answer must win.一台机器上 83 个会话,得到 34 个例子:某会话里某条 skill 第一次被用,和那之前打的那句话。有两个排序只用来检查打分脚本本身:打乱的必须输,偷看了答案的必须赢。

The project boost has one more top-3 hit than learned messages alone, and was still dropped: compared case by case it moved the right skill down four times and up three. The learned messages won nine times, lost none, and the wins came from six projects. Each formula was written down before its first replay and tested once.「本项目用过就加分」比单用学到的原话多一个前 3 命中,照样不用:逐例比,它把对的那条往下挤了四次、往上提了三次。学到的原话赢九次、一次没输,胜例来自六个项目。每个公式都是第一次回放前写下的,只测一次。

Where the ideas came from思路从哪来

Borrowed where it held up, left behind where it did not.能用的借过来,用不上的不借。

Took借了

karpathy/autoresearch

Keep a change only when the frozen score improves. What it lacks was added: splitting by time, a paired case-by-case test, “not enough evidence” as its own verdict, and a scorer that has to beat a shuffle and lose to a peek.只在冻结的评分变好时保留改动。它缺的几样补上了:按时间切分、逐例配对比较、「证据不足」单独算一种结论、打分脚本必须赢过打乱的、输给偷看的。

Read读了

openai/codex

Its skill selector runs a dozen cheap rankings (BM25F, character n-grams, recently used) in shadow mode each turn, logging where the skill the model really called would have ranked, without changing the prompt. Marked temporary in its source.它的 skill 选择实验每轮在后台跑十几种便宜的排序(BM25F、字符 n-gram、最近用过),只记录模型真正调用的那条会排第几,不改提示词。源码里标着是临时实验。

Left behind没借

lomeshdutta/skill-router

A hook on every message was tried there and withdrawn in v0.2 for once per session: good evaluation scores, too noisy in real sessions (paraphrased). Here the hook count is zero unless a replay shows a median of one prompt per session or less.那里试过每条消息都跑的钩子,v0.2 撤回,改成每会话一次:评测分数好,真实会话里太吵(大意)。这里钩子数为零,除非回放证明每个会话提示次数的中位数不超过一次。

Measured实测依据

Picking the right tool怎么挑准工具

SkillRouter (arXiv 2603.22455): bodies matter, 56.0% → 18.7% top-1 without them. ToolRet: requests and tool docs share little wording, 0.06 overlap. SkillLens (arXiv 2605.23899) and SkillsBench: a wrong skill can do harm. All figures as reported by the authors.SkillRouter(arXiv 2603.22455):正文要紧,去掉后首位命中 56.0% → 18.7%。ToolRet:请求和工具文档很少共用词,重合 0.06。SkillLens(arXiv 2605.23899)与 SkillsBench:错的 skill 会帮倒忙。数字均为作者自报。

Honest caveat坦诚的边界

What is not done, and what the numbers cannot say.哪些还没做,数字说明不了什么。

Only Claude Code is wired up. The adapters for Codex, opencode, pi, Gemini CLI and Cursor are designed from their source and not written; linking the library into ~/.agents/skills and checking Codex’s menu limit at install time are not done either. whetstone find itself runs anywhere there is a shell and Python.目前只接好了 Claude Code。Codex、opencode、pi、Gemini CLI、Cursor 的适配器是按源码设计的,还没写;把库链接到 ~/.agents/skills、装库时按 Codex 的规则算菜单上限,也都还没做。whetstone find 本身在任何有 shell 和 Python 的地方都能跑。

Every number comes from one machine and one user. The replay uses today’s skill text, and some descriptions gained trigger words only after they were missed, so “text only” reads high. The threshold is not set yet.所有数字来自一台机器、一个使用者。回放用的是 skill 现在的文字,有些描述是在当时没被用上之后才补的触发词,所以「只看文字」那一行偏高。门槛还没定下来。

And find still depends on the agent remembering to run it. The always-loaded line and the paths switch make that likelier; nothing makes it certain, and a replay cannot measure it.另外,find 还是要靠 agent 想起来去跑。常驻那一行和 paths 让它更可能被想起来,但没有办法保证,回放也测不出这一步。

Technical details技术细节

The parts you can check.能核对的那些部分。

Commands命令

whetstone find · route events · route learn · route usage · route replay · route elsewhere · route plan · route catalog

Event line事件行

Every line carries v, runtime, session, ts, and optionally src (log file and line). Kinds: session, user, file, load, menu, find. A new runtime needs only an adapter that writes these.每行都带 v、runtime、session、ts,可选 src(会话文件和行号)。种类:session、user、file、load、menu、find。接一个新运行时,只要写一个产出这些行的适配器。

Privacy隐私

Learned messages are your own words. They stay on your machine (~/.local/share/whetstone/) and never go into a portable skill package.学到的原话是你自己打的字,只存在本机(~/.local/share/whetstone/),不进可移植的 skill 包。

Tests测试

73 selftest checks, each rule in both directions, plus a final check that nothing outside the test folder was written. 57 mutations: each removes one rule and the suite must go red.73 项自检,每条规则正反两个方向都测,最后还查测试目录以外有没有文件被写。57 条故意改坏:每条拿掉一条规则,整套测试必须报红。

Spec规范

spec/routing.md — the three layers, the event format, the six-runtime table with file and line, the replay rules, and what may never be done.—— 三层结构、事件格式、带文件和行号的六家对照、回放规则和禁止清单。

whetstone is MIT, standard library only. Routing ships in the same command line as distilling, verifying and the use-time check.whetstone 是 MIT 协议,只用标准库。经验路由和蒸馏、校验、用时纠错在同一个命令行里。