Harness · sky-skills Harness · sky-skills

The Harness.
Nine components. One rule.
架构蓝图。
九个组件。一条规矩。

The full design of the sky-skills harness — nine design generator skills and two evaluators with four review gates. Now all nine are built; eight are done — the canonical library (02) just reached the full 59/59 matrix — and design-evolve (09) has not run an evolution round yet. Every component must defend one question: what specific model weakness does it correspond to? sky-skills 架构的完整蓝图 —— 九个设计生成器 skill 加两个 评审员(4 道审查检查)。现在九个全部建好;八个已完成——canonical 库(02)刚补满 59/59 整张覆盖表——design-evolve(09)还没跑过一轮进化。每个组件都必须能回答一个问题:它挡的是模型的哪一个具体弱点?

See the nine看这九个 What's done today今天做到哪里

Read in another voice用另一种声音读 apple anthropic ember sage glass
23 focused skills — 18 generate, 2 judge, 1 plans, 1 evolves, 1 keeps them current精选 skill — 18 个生成,2 个评审,1 个规划,1 个自进化,1 个保持更新
9 harness components — each defending one model weaknessharness 组件 — 每个挡一个模型弱点
9 design voices — this roadmap is rendered in five of them设计声音 — 这份路线图用其中五种各渲一版

Sources指导来源

Two ideas we're building on. 我们站在的两个肩膀。

GAN (Goodfellow et al., 2014) — a generator and a discriminator train adversarially. The sharper the discriminator, the stronger the generator becomes. Our nine design skills are the generator, design-review is the discriminator.

GAN(Goodfellow et al., 2014) —— generator 和 discriminator 对抗训练。discriminator 越锐利,generator 就越强。我们的 9 个设计 skill 是 generator,design-review 是 discriminator。

Anthropic, Harness design for long-running appsplanner / generator / evaluator three-part split, evaluators must be separate, sprint contracts, 5–15 iteration rounds. The rule we take most seriously: each harness component must correspond to one specific model weakness.

Anthropic 的 Harness design for long-running apps —— planner / generator / evaluator 三段式,evaluator 必须独立,sprint contract,5–15 轮迭代。我们最认真对待的一条:每一个 harness 组件都要对应模型的一个具体弱点

Agents tend to praise their own work confidently — even when quality is obviously mediocre to a human observer. Separating the agent doing the work from the agent evaluating it is a powerful lever. Agent 倾向于自信地称赞自己的作品 —— 哪怕在人类观察者看来质量明显平庸。把做事的 agent 和评判的 agent 分开,是一个真正的杠杆。

— Anthropic, harness-design post—— Anthropic,harness-design 文章

Goal目标

A page that stands next to
anthropic.com · stripe.com · linear.app
一页能和 anthropic.com ·
stripe.com · linear.app 同台站着

Not technically bug-freetaste, originality, rhythm, copy, hand-crafted illustrations, device-agnostic, accessible, shareable. From a one-sentence brief. 不是"技术上没 bug",而是 有品位、有原创、有节奏、copy 过关、图示手工感强、全设备可用、无障碍、可分享。从一句话需求开始。

Nine components九个组件

Every row defends a weakness. 每一行都挡一个弱点。

The only legitimate reason for a harness component to exist is that a specific model weakness makes it necessary. No weakness, no seat. 一个 harness 组件能存在的唯一理由,是有一条具体的模型弱点要求它存在。挡不住弱点,没位置。

Future待做 01

design-planner

skills/design-planner/ Model weakness对应弱点

Faced with a vague brief, the model fills in training-data medians. Scope drifts. A planner expands the brief into a sprint contract before the generator runs. 面对模糊需求,模型会用训练数据的中位数填空,scope 失控。planner 在生成器动手前,把需求展开成 sprint contract

Done · 59/59已完成 · 59/59 02

reference library参考库

skills/<style>-design/references/canonical/ Model weakness对应弱点

Without reference, output regresses to the most average web design. A per-style canonical library of 10 page-types anchors quality at each style's own best — not an industry average. 没有参考,输出就向"最平均的 web 设计"收敛。每个 style 自己的 10 个 page-type canonical 库,把品质锚定在该 style 自己的最好。

Partly done部分完成 03

generator + self-diff生成器 + 自评差异

skills/{apple,anthropic,ember,sage}-design/ Model weakness对应弱点

Models don't naturally articulate their design decisions after writing. Without a required self-diff note, the critic has no concrete target to argue with — just vibes. 模型写完 HTML 以后,不会自然记录它的设计决策。没有强制的 self-diff note,critic 没有具体靶子,只能凭感觉评。

Phase 1 done · expanding期 1 已完成 · 扩展中 04

mechanical-review

skills/design-review/ Model weakness对应弱点

Models don't self-check runtime behavior — multi-viewport, keyboard, color contrast, LCP/CLS, SEO meta. A mechanical gate catches what source-scanning and self-review miss. runtime 行为(多视口、键盘可达、对比度、LCP/CLS、SEO meta)模型自己不会检查。mechanical 检查抓的正是源扫和自评都抓不到的那类 bug。

Done · 2026-04-22已完成 · 2026-04-22 05

multi-critic多位专家评审员

.claude/agents/design-{composition,copy,illustration,brand}-critic.md Model weakness对应弱点

One critic spreads attention across 7 dimensions and underweights any one. Four specialists (weights 25/25/20/30) run in parallel fresh contexts — composition, copy, illustration, brand — then an aggregator merges. First live run: solo 93, multi 88; illustration caught a 6th-hue SVG leak the generalist missed. 一个 critic 要把注意力摊给 7 维,每一维都浅。4 位专家(权重 25/25/20/30)并行 fresh context —— 版式、copy、插画、品牌 —— 再由 aggregator 合成。首次实战:单评 93,多评 88;插画专家抓到了通才漏掉的 SVG 6 色偷渡。

Shipped · first run critic 92已发 · 首跑 critic 92 06

/design-loop

.claude/commands/design-loop.md Model weakness对应弱点

Models default to ship the first-pass output. Without orchestration, actual iteration ends after 1–2 rounds. A loop controller forces up to 5 rounds, or escalates. 模型默认"第一版够用就发"。没有编排,实际只跑 1-2 轮。loop controller 强制最多 5 轮,超过再交还给人。

Done · 2026-04-22已完成 · 2026-04-22 07

learning-loop

.claude/agents/design-learner.md + scripts/learning-loop.mjs Model weakness对应弱点

"The new issue the critic caught today" doesn't auto-archive. design-learner reads critic verdicts and proposes 3 codifications — new known-bugs rows, new visual-audit checks, per-skill dos-and-donts. First run 2026-04-22: 13 raw issues → 5 bug classes → 3 known-bugs rows (1.17/1.18/1.19) + 1 new check + 8 dos-and-donts. Same bug never caught twice. "critic 今天新抓到的问题"不会自动归档。design-learner 读 verdict,提三件事:known-bugs 新行、visual-audit 新 check、每 skill 的 dos-and-donts。首次实战 2026-04-22:13 条 → 5 类 → 3 行 known-bugs + 1 个新 check + 8 条 dos-and-donts。同一个 bug 不再抓第二次。

Future · after 10+ pages待做 · 10+ 真实页后 08

library-grower

.claude/commands/design-distill.md Model weakness对应弱点

Models can't self-organize yesterday's good output into tomorrow's reference. Without this, the canonical library stays fixed and stale over 6 months. 模型没有"把昨天好产物整理成明天参考"的能力。没有它,canonical library 会固定,6 个月后开始过期。

Built · no round yet机制已建 · 未跑 09

design-evolve

skills/design-evolve/SKILL.md Model weakness对应弱点

Models repeat the same taste mistakes and can't tell which of their own rules actually help. A Darwin / autoresearch ratchet: mutate one generator rule, regenerate, score with the frozen evaluator, keep it only if it strictly beats the baseline with no held-out canonical regressing. Selection, not authorship. 模型会重复同样的审美错误,也分不清自己哪条规则真有用。一个 Darwin / autoresearch 式爬坡:改一条生成器规则、重新生成、用冻结的评审器打分,只有严格超过基线且无 held-out canonical 变差才保留。由选择而非作者决定。

Where it stands当前进度

Nine components, by status. 九个组件,按当前状态。

DONE · 7 JUST COMPLETED · 1 BUILT · NO ROUND YET · 1 Done · the pipeline is built 01 plannersprint-contract + --plan 03 generator self-diff§M + verify.py gate 04 mechanical-review51 checks · 94 known-bugs 05 multi-critic4 subagents + solo critic 06 /design-loop ran · critic 92first real end-to-end run 07 learning-loop 08 library-growerdesign-learner · distill 5+ to candidate Just completed 02 canonical library 58 / 58 page-types covered matrix complete — 10 page-types × anthropic / apple / ember / sage + 18 more run --coverage for the live count Built · no round yet 09 design-evolve Darwin / autoresearch ratchet, wired: evolve-ledger.mjsevolve-rules.mjsregression-gate.mjs rules.json empty · 0 rounds run ledger = the loop's first-run deposit when it runs, the evaluator stays frozen
The harness at a glance — eight components done, the canonical library (02) now complete at the full 59/59 matrix, and design-evolve (09) wired end-to-end but yet to run its first evolution round. harness 全貌 —— 八个组件已完成,canonical 库(02)已补满 59/59 整张覆盖表,design-evolve(09)已端到端接线但还没跑过第一轮进化。

Hard constraint硬约束

The harness must not flatten the nine voices. harness 越复杂,越怕抹平 9 种声音。

Universal rules never judge style; style rules live in per-style folders; critic's taste is anchored to each style's own canonical. universal 规则从不评判风格,风格规则住在各自文件夹,critic 的品味锚定在每个 style 自己的 canonical。

Layer Where it lives住在哪 Flattens the voices?会抹平声音吗
Universal quality通用质量 skills/design-review/references/cross-skill-rules.md No不会
Only judges does it work.只管"能不能用"
Cross-skill bugs跨 skill bug 清单 skills/design-review/references/known-bugs.md No不会
Already segmented into per-style sections.按分节组织
Style canonicalsstyle canonical 例子 skills/<style>-design/references/canonical/ Strengthens them反过来强化
Apple pricing and sage pricing look nothing alike. Libraries are never shared.apple pricing 和 sage pricing 长得完全不一样。不互抄
Style personality风格个性 skills/<style>-design/references/dos-and-donts.md Strengthens them反过来强化
Palette, fonts, signature moves — all per-style.色板、字体、签名动作 —— 全是各自的
Critic's tastecritic 的品味 Subagent reads this style's canonical + dos-and-dontssubagent 读 当前 style 的 canonical + dos-and-donts No不会
Critic always asks is this a good <style> page? — not abstract good.critic 永远问"这是不是一张好的 <style>; 页",不是绝对意义上好

Status · 2026-06-11今天 · 2026-06-11

What's real. What's next. 哪些是真的。下一步做什么。

Done 8 (incl. 02) · Built-no-round 1 (09)已完成 8(含 02) · 机制已建未跑 1(09)

What's actually done 实际已完成

51-check visual-audit.mjs + 94 known-bugs catalogued. 59/59 canonical page-types covered — the full matrix, each with a binding .md rubric. multi-critic — 4 specialist subagents in parallel (composition / copy / illustration / brand) plus solo design-critic. learning-loopdesign-learner turns critic misses into new mechanical checks. 23 skills total: 9 design generators, 7 systems / content utilities, 2 workflow generators (gated-dual-clone, doc-review-loop), 2 independent evaluators (design-review, gated-dual-clone-audit), 1 planner (design-planner), 1 self-evolution loop (design-evolve), 1 skills auto-updater (skills-sync). 2026-04-27: anthropic-design v2 (scenario recipes + ux-writing + recipe components in anthropic.css) and bin/design-review --audit batch mode (file / dir / URL). 2026-05-22: anthropic-design v3 — a four-script md rendering pipeline (md-mirror / md-rewrite-links / md-pack / cross-link-pack). Markdown in, styled HTML out. Every link survives cp -r. 2026-06-11: diagram-craft v3 — seven kernel diagram types, constraints in three layers (aesthetics fixed / quality machine-gated / structure free). 5 new visual-audit.mjs gates (known-bugs 1.28–1.31). sprint-contract §1b diagram density. Templates: anthropic 14, apple 14. Two diagram galleries: 23 + 14. Gold corrected #c49464#c9913f. Wave 2: visual-audit adds keyboard-a11y + LCP/CLS perf gates. 38 checks now. verify.py gains a warn tier with SEO meta checks. --audit reports per-kind stats. --distill wired (component 08). design-planner ships as the 13th skill (component 01). Same day: diagram-craft lands in ember-design and sage-design. 8 templates each. 8-figure galleries. The monochrome gate now covers anthropic / ember / sage — apple stays exempt. Grayscale is the identity. 51 项 visual-audit.mjs + 94 条 known-bugs。覆盖 59/59 canonical page-type —— 整张覆盖表,每张配绑定 .mdmulti-critic — 4 位专家 subagent 并行(版式 / 文案 / 插画 / 品牌)+ 单 design-criticlearning-loopdesign-learner 把 critic miss 转成新机器 check。23 个 skill:9 个设计 generator、7 个系统 / 内容工具、2 个 workflow generator(gated-dual-clonedoc-review-loop)、2 个独立 evaluator(design-reviewgated-dual-clone-audit)、1 个 planner(design-planner)、1 个自进化 loop(design-evolve)、1 个 skills 自动更新器(skills-sync)。2026-04-27:anthropic-design v2(scenario recipes + ux-writing + recipe components 入 anthropic.css)、bin/design-review --audit 批量模式(file / dir / URL)。2026-05-22anthropic-design v3 — 4 件套 md 渲染管线(md-mirror / md-rewrite-links / md-pack / cross-link-pack)。md 进,风格化 HTML 出,cp -r 后链接全部有效。2026-06-11diagram-craft v3 — 内核七图型 + 约束三层化(审美不可变 / 质量机器检查 / 结构自由定制)。visual-audit.mjs 新增 5 项检查(known-bugs 1.28–1.31)。sprint-contract §1b 图密度。模板库 anthropic 14 件、apple 14 件。双 diagram gallery:23 图 + 14 图。金主色修正 #c49464#c9913fWave 2visual-audit 加 keyboard 可达性 + LCP/CLS 性能检查。现在 38 项 check。verify.py 加 warn 档,带 SEO meta 检查。--audit 按 kind 出统计。--distill 接线完成(组件 08)。design-planner 作为第 13 个 skill 上线(组件 01)。同日:diagram-craft 落进 ember-designsage-design。各 8 件模板。各 8 图 gallery。monochrome 检查覆盖 anthropic / ember / sage —— apple 豁免。灰阶即身份。

Now · critic validated现在 · critic 已验证

4 real pages reviewed 评审了 4 张真实页

2026-04-22: all 4 HARNESS-ROADMAP variants ran through bin/design-review --critic. anthropic climbed 84 → 93 in three rounds. apple 88, ember 84, sage 78 → pass after fixes. Verdicts cite line numbers, class names, CSS to paste — specific and actionable. Each fix pins an §H / §J / §K regression for the next visual-audit. 2026-04-22:4 个 HARNESS-ROADMAP 都跑了 bin/design-review --critic。anthropic 三轮 84 → 93。apple 88、ember 84、sage 78 → 修完通过。每条建议都引用行号、类名、可粘贴 CSS —— 具体可操作。每次修复都把一个 §H / §J / §K 回归锁进下一次 visual-audit。

Shipped · 2026-06-14/15已发 · 2026-06-14/15

loop · grower · evolve — all built loop · grower · evolve 全建好

(06) /design-loop shipped its first real end-to-end run — an anthropic security page at critic 92, now the corpus's first page. (08) library-grower distills 5+ same-type pages into a candidate (candidate.html + diffs.md). (09) design-evolve has the full ratchet wired but has not run an evolution round yet. And the whole harness now installs globally under ~/.claude/ and runs from any directory (2026-06-15). (06) /design-loop 完成首次端到端实战 —— anthropic security 页 critic 92,已是 corpus 第一页。(08) library-grower 把 5+ 张同类页蒸馏成候选(candidate.html + diffs.md)。(09) design-evolve 棘轮已端到端接线,但还没跑过一轮进化。整套 harness 现在全局装到 ~/.claude/、任意目录可跑(2026-06-15)。

Rules for changing this document修改本文档的规矩

How this doc stays honest. 这份文档怎么保持诚实。

  • Adding a component新增组件 Must state the specific model weakness it defends. No weakness, no seat. 必须写清它挡的是模型的哪一个具体弱点。挡不住弱点,没位置。
  • Removing a component删除组件 Must present evidence the weakness no longer applies. Vibes aren't evidence. 必须给出证据,证明那条弱点已经不成立。感觉不算证据。
  • After every real page每跑完一张真实页 Update the Status today section. If the harness caught a bug class we hadn't catalogued, append to known-bugs.md and add a check. 同步"今天"这一节。如果抓到一个新 bug 类,同时追加到 known-bugs.md + 加一个 check。
  • If the harness drifts如果 harness 跑偏 Flattens the nine voices, produces noise that trains evaluators to ignore warnings, gives the generator a way to game the test — write a new Historical lessons entry. 抹平 9 种声音、产生让评审员习惯跳过的噪音、给 generator 作弊空间 —— 追加一条"历史教训"。

Historical lessons历史教训

Things we were wrong about. 我们曾经搞错的事。

2026-04-20

Simplicity first is a heuristic, not a hard rule. "simplicity first" 是经验法则,不是硬规则。

While designing this roadmap, I (Claude) once argued for cutting 8 components to 5, citing the article's simplicity first principle. The user correctly pushed back. The article's actual hard rule isn't fewer is better — it's every component must correspond to a specific model weakness. The retreat was unwarranted and I restored all 8. Lesson: when tempted to cut a component, the only valid argument is evidence that the weakness it defends doesn't exist. 写这份 roadmap 时,我(Claude)曾经引用文章里 "simplicity first" 把 8 组件砍到 5。用户正确地反问。文章真正的硬规矩不是"越少越好",是"每个组件对应一个具体弱点"。那次退让没有依据,我恢复了 8 个。教训:想裁一个组件时,唯一有效的理由是证据表明它挡的弱点不存在

Think a component is wrong?
Missing? Misranked?
觉得某个组件不对?
漏了?顺序错了?

Argue with evidence. We'd rather restructure the harness than keep a component that isn't really doing anything. 用证据讲话。我们宁愿重构 harness,也不愿让一个其实没挡任何东西的组件继续占位。

Open a discussion ›开一个 discussion ›      File an issue ›提一个 issue ›      Back to sky-skills ›返回 sky-skills ›