00 · Architecture · 2026-04-20架构 · 2026-04-20

The Harness. Nine components. One rule. 架构蓝图。九个组件。一条规矩。

The full design of the sky-skills harness — nine generator skills and two evaluators with four review gates. Now all nine are built; eight are done — the canonical library (02) just reached the full 59/59 matrix — and design-evolve (09) has not run an evolution round yet. Every component must defend one question: what specific model weakness does it correspond to? sky-skills 架构的完整蓝图 —— 九个 生成器 skill 加两个 评审员(4 道审查检查)。现在九个全部建好;八个已完成——canonical 库(02)刚补满 59/59 整张覆盖表——design-evolve(09)还没跑过一轮进化。每个组件都必须能回答一个问题:它挡的是模型的哪一个具体弱点?

See the nine看这九个 What's done today今天做到哪里
Read in another voice用另一种声音读 apple anthropic ember sage glass
23 focused skills — 18 generate, 2 judge, 1 plans, 1 evolves, 1 keeps them current精选 skill —— 18 个生成,2 个评审,1 个规划,1 个自进化,1 个保持更新
9 harness components — each defending one weaknessharness 组件 —— 各挡一个弱点
9 design voices — this roadmap is rendered in five of them设计声音 —— 这份路线图用其中五种各渲一版
01 · Sources指导来源

Two ideas we're building on. 我们站在的两个肩膀。

GAN (Goodfellow et al., 2014) — a generator and a discriminator train adversarially. The sharper the discriminator, the stronger the generator becomes. Our nine design skills are the generator, design-review is the discriminator.

GAN(Goodfellow et al., 2014) —— generator 和 discriminator 对抗训练。discriminator 越锐利,generator 就越强。我们的 9 个设计 skill 是 generator,design-review 是 discriminator。

Anthropic, Harness design for long-running appsplanner / generator / evaluator three-part split, evaluators must be separate, sprint contracts, 5–15 iteration rounds. The rule we take most seriously: each harness component must correspond to one specific model weakness. Assumptions worth stress-testing.

Anthropic 的 Harness design for long-running apps —— planner / generator / evaluator 三段式,evaluator 必须独立,sprint contract,5–15 轮迭代。我们最认真对待的一条:每一个 harness 组件都要对应模型的一个具体弱点。这些假设值得压测。

Agents tend to praise their own work confidently — even when quality is obviously mediocre to a human observer. Separating the agent doing the work from the agent evaluating it is a powerful lever. Agent 倾向于自信地称赞自己的作品 —— 哪怕在人类观察者看来质量明显平庸。把做事的 agent 和评判的 agent 分开,是一个真正的杠杆。

Anthropic · harness-design postAnthropic · harness-design 文章

02 · Goal目标

A page that stands next to
anthropic.com, stripe.com, linear.app.
一页能和 anthropic.com、
stripe.com、linear.app 同台站着。

Not technically bug-freetaste, originality, rhythm, copy, hand-crafted illustrations, device-agnostic, accessible, shareable. From a one-sentence brief. 不是"技术上没 bug",而是 有品位、有原创、有节奏、copy 过关、图示手工感强、全设备可用、无障碍、可分享。从一句话需求开始。

03 · Nine components九个组件

Every row defends a weakness. 每一行都挡一个弱点。

The only legitimate reason for a harness component to exist is that a specific model weakness makes it necessary. No weakness, no seat. 一个 harness 组件能存在的唯一理由,是有一条具体的模型弱点要求它存在。挡不住弱点,没位置。

shipped · 2026-06-11已上线 · 2026-06-11 01

design-planner

skills/design-planner/ Model weakness对应弱点

Faced with a vague brief, the model fills in training-data medians. Scope drifts. A planner expands the brief into a sprint contract before the generator runs. 面对模糊需求,模型用训练数据的中位数填空,scope 失控。planner 在生成器动手前,把需求展开成 sprint contract

Done · 59/59已完成 · 59/59 02

reference library参考库

skills/<style>-design/references/canonical/ Model weakness对应弱点

Without reference, output regresses to the most average web design. A per-style canonical library of 10 page-types anchors quality at each style's own best. 没有参考,输出向"最平均的 web 设计"收敛。每个 style 自己的 10 个 canonical,把品质锚定在该 style 自己的最好。

partly done部分完成 03

generator + self-diff生成器 + 自评差异

skills/{apple,anthropic,ember,sage}-design/ Model weakness对应弱点

Models don't naturally articulate their design decisions after writing. Without a required self-diff note, the critic has no concrete target to argue with. 模型写完 HTML 后不会自然记录它的设计决策。没有强制的 self-diff note,critic 没具体靶子。

phase 1 done期 1 完成 04

mechanical-review

skills/design-review/ Model weakness对应弱点

Models don't self-check runtime behavior — multi-viewport, keyboard, color contrast, LCP/CLS, SEO meta. A mechanical gate catches what source-scanning and self-review miss. runtime 行为(多视口、键盘、对比度、LCP/CLS、SEO meta)模型不自检。mechanical 检查抓源扫和自评都抓不到的。

done · 2026-04-22已完成 · 2026-04-22 05

multi-critic多位专家评审员

.claude/agents/design-{composition,copy,illustration,brand}-critic.md Model weakness对应弱点

One critic spreads attention across 7 dimensions and underweights any one. Four specialists (weights 25/25/20/30) run in parallel fresh contexts — composition, copy, illustration, brand — then an aggregator merges. First live run: solo 93, multi 88; illustration caught a 6th-hue SVG leak the generalist missed. 一个 critic 把注意力摊给 7 维,每一维都浅。4 位专家(权重 25/25/20/30)并行 fresh context —— 版式、copy、插画、品牌 —— 再由 aggregator 合成。首次实战:单评 93,多评 88;插画专家抓到了通才漏掉的 SVG 6 色偷渡。

shipped · first run critic 92已发 · 首跑 critic 92 06

/design-loop

.claude/commands/design-loop.md Model weakness对应弱点

Models default to ship the first-pass output. Without orchestration, actual iteration ends after 1–2 rounds. A loop controller forces up to 5 rounds. 模型默认"第一版够用就发"。没编排,实际只跑 1-2 轮。loop controller 强制最多 5 轮。

done · 2026-04-22已完成 · 2026-04-22 07

learning-loop

.claude/agents/design-learner.md + scripts/learning-loop.mjs Model weakness对应弱点

"The new issue the critic caught today" doesn't auto-archive. design-learner reads critic verdicts and proposes 3 codifications — new known-bugs rows, new visual-audit checks, per-skill dos-and-donts. First run 2026-04-22: 13 raw issues → 5 bug classes → 3 known-bugs rows (1.17/1.18/1.19) + 1 new check + 8 dos-and-donts. Same bug never caught twice. "critic 今天新抓到的问题"不会自动归档。design-learner 读 verdict,提三件事:known-bugs 新行、visual-audit 新 check、每 skill 的 dos-and-donts。首次实战 2026-04-22:13 条 → 5 类 → 3 行 known-bugs + 1 个新 check + 8 条 dos-and-donts。同一个 bug 不再抓第二次。

future · 10+ pages待做 · 10+ 真实页后 08

library-grower

.claude/commands/design-distill.md Model weakness对应弱点

Models can't self-organize yesterday's good output into tomorrow's reference. Without this, the canonical library stays fixed and stale over 6 months. 模型没有"把昨天好产物整理成明天参考"的能力。没有它,canonical library 会固定,6 个月后过期。

Built · no round yet机制已建 · 未跑 09

design-evolve

skills/design-evolve/SKILL.md Model weakness对应弱点

Models repeat the same taste mistakes and can't tell which of their own rules actually help. A Darwin / autoresearch ratchet: mutate one generator rule, regenerate, score with the frozen evaluator, keep it only if it strictly beats the baseline with no held-out canonical regressing. Selection, not authorship. 模型会重复同样的审美错误,也分不清自己哪条规则真有用。一个 Darwin / autoresearch 式爬坡:改一条生成器规则、重新生成、用冻结的评审器打分,只有严格超过基线且无 held-out canonical 变差才保留。由选择而非作者决定。

04 · Where it stands当前进度

Nine components, by status. 九个组件,按当前状态。

DONE · 7 JUST COMPLETED · 1 BUILT · NO ROUND YET · 1 Done — the pipeline is built 01 plannersprint-contract + --plan 03 generator self-diff§M + verify.py gate 04 mechanical-review51 checks · 94 known-bugs 05 multi-critic4 subagents + solo critic 06 /design-loop ran · critic 92first real end-to-end run 07 learning-loop 08 library-growerdesign-learner · distill 5+ to candidate Just completed 02 canonical library 58 / 58 page-types covered matrix complete — 10 page-types × anthropic / apple / ember / sage + 18 more run --coverage for the live count Built · no round yet 09 design-evolve Darwin / autoresearch ratchet, wired: evolve-ledger.mjsevolve-rules.mjsregression-gate.mjs rules.json empty · 0 rounds run ledger = the loop's first-run deposit when it runs, the evaluator stays frozen
The harness at a glance — eight components done, the canonical library (02) now complete at the full 59/59 matrix, and design-evolve (09) wired end-to-end but yet to run its first evolution round. harness 全貌 —— 八个组件已完成,canonical 库(02)已补满 59/59 整张覆盖表,design-evolve(09)已端到端接线但还没跑过第一轮进化。
05 · Hard constraint硬约束

The harness must not flatten the nine voices. harness 越复杂,越怕抹平九种声音。

Universal rules never judge style; style rules live in per-style folders; critic's taste is anchored to each style's own canonical. universal 规则不评判风格,风格规则住在各自文件夹,critic 的品味锚定在每个 style 自己的 canonical。

Layer Where it lives住在哪 Flattens the voices?会抹平声音吗
Universal quality通用质量 skills/design-review/references/cross-skill-rules.md No不会
Only judges does it work.只管"能不能用"
Cross-skill bugs跨 skill bug 清单 skills/design-review/references/known-bugs.md No不会
Segmented into per-style sections.按分节组织
Style canonicalsstyle canonical 例子 skills/<style>-design/references/canonical/ Strengthens them反过来强化
Apple pricing and sage pricing look nothing alike. Libraries are never shared.apple pricing 和 sage pricing 长得完全不一样。不互抄
Style personality风格个性 skills/<style>-design/references/dos-and-donts.md Strengthens them反过来强化
Palette, fonts, signature moves — all per-style.色板、字体、签名动作 —— 全是各自的
Critic's tastecritic 的品味 Subagent reads this style's canonical + dos-and-dontssubagent 读 当前 style 的 canonical + dos-and-donts No不会
Critic always asks is this a good <style> page?critic 永远问"这是不是一张好的 <style> 页"
06 · Status · 2026-06-11今天 · 2026-06-11

What's real. What's next. 哪些是真的。下一步。

done 8 (incl. 02) · built-no-round 1 (09)已完成 8(含 02) · 机制已建未跑 1(09)

What's actually done 实际已完成

51-check visual-audit.mjs + 94 known-bugs catalogued. 59/59 canonical page-types covered — the full matrix, each with a binding .md rubric. multi-critic — 4 specialist subagents in parallel (composition / copy / illustration / brand) plus solo design-critic. learning-loopdesign-learner turns critic misses into new mechanical checks. 23 skills total: 9 design generators, 7 systems / content utilities, 2 workflow generators (gated-dual-clone, doc-review-loop), 2 independent evaluators (design-review, gated-dual-clone-audit), 1 planner (design-planner), 1 self-evolution loop (design-evolve), 1 skills auto-updater (skills-sync). 2026-04-27: anthropic-design v2 (scenario recipes + ux-writing + recipe components in anthropic.css), bin/design-review --audit batch mode (file / dir / URL). 2026-05-22: anthropic-design v3 — a four-script md rendering pipeline (md-mirror / md-rewrite-links / md-pack / cross-link-pack); a styled doc directory that keeps every link intact after cp -r. 2026-06-11: diagram-craft v3 — seven kernel diagram types, and constraints settled into three layers: aesthetics immutable, quality machine-gated, structure free to customize. 5 new visual-audit.mjs gates (known-bugs 1.28–1.31); sprint-contract §1b diagram density; templates anthropic 14 / apple 14; two diagram galleries (23 + 14 diagrams); gold corrected #c49464#c9913f across skills. Wave 2: visual-audit grows keyboard-a11y and LCP/CLS perf gates (38 checks); verify.py gains a warn tier with SEO meta checks; --audit reports per-kind stats; --distill wired (component 08); design-planner ships quietly as the 13th skill (component 01). And quietly, the same day, the craft reaches the other two — ember-design and sage-design each gain diagram-craft, 8 templates, an 8-figure gallery; the monochrome gate now holds for anthropic / ember / sage, and apple, grayscale by nature, stays exempt. 51 项 visual-audit.mjs + 94 条 known-bugs。覆盖 59/59 canonical page-type —— 整张覆盖表,每张配绑定 .mdmulti-critic —— 4 位专家 subagent 并行(版式 / 文案 / 插画 / 品牌)+ 单 design-criticlearning-loop —— design-learner 把 critic miss 转成新机器 check。23 个 skill:9 个设计 generator、7 个系统 / 内容工具、2 个 workflow generator(gated-dual-clonedoc-review-loop)、2 个独立 evaluator(design-reviewgated-dual-clone-audit)、1 个 planner(design-planner)、1 个自进化 loop(design-evolve)、1 个 skills 自动更新器(skills-sync)。2026-04-27:anthropic-design v2(scenario recipes + ux-writing + recipe components 入 anthropic.css)、bin/design-review --audit 批量模式(file / dir / URL)。2026-05-22anthropic-design v3 —— 4 件套 md 渲染管线(md-mirror / md-rewrite-links / md-pack / cross-link-pack),风格化文档目录 cp -r 之后每条链接照常工作。2026-06-11diagram-craft v3 —— 内核七图型,约束安顿成三层:审美不可变、质量机器检查、结构自由定制。visual-audit.mjs 新增 5 项检查(known-bugs 1.28–1.31);sprint-contract §1b 图密度;模板库 anthropic 14 件 / apple 14 件;双 diagram gallery(23 图 + 14 图);金主色修正 #c49464#c9913f,全 skill 同步。Wave 2visual-audit 新增 keyboard 可达性与 LCP/CLS 性能检查(38 项 check);verify.py 添 warn 档,含 SEO meta 检查;--audit 按 kind 给出统计;--distill 接线完成(组件 08);design-planner 作为第 13 个 skill 安静上线(组件 01)。同日,工艺安静地抵达另外两家 —— ember-designsage-design 各得 diagram-craft、8 件模板、8 图画廊;monochrome 检查对 anthropic / ember / sage 生效,apple 以灰阶为本,豁免如常。

now · critic validated现在 · critic 已验证

4 real pages reviewed 评审了 4 张真实页

2026-04-22: all 4 HARNESS-ROADMAP variants ran through bin/design-review --critic. anthropic 84 → 93 across three rounds. apple 88, ember 84, sage 78 → pass after fixes. Verdicts cite line numbers, class names, paste-ready CSS — specific and actionable. Each fix cycle pins an §H / §J / §K regression into the next visual-audit. 2026-04-22:4 张 HARNESS-ROADMAP 都跑了 bin/design-review --critic。anthropic 三轮 84 → 93,apple 88,ember 84,sage 78 → 修完通过。每条 verdict 直接给行号、类名、可粘贴 CSS —— 具体可操作。每一轮修都把 §H / §J / §K 回归锁进下一次 visual-audit。

shipped · 2026-06-14/15已发 · 2026-06-14/15

loop · grower · evolve — all built loop · grower · evolve 全建好

(06) /design-loop shipped its first real end-to-end run — an anthropic security page at critic 92, now the corpus's first page. (08) library-grower distills 5+ same-type pages into a candidate. (09) design-evolve has the full ratchet wired but has not run an evolution round yet. And the whole harness now installs globally under ~/.claude/ and runs from any directory (2026-06-15). (06) /design-loop 完成首次端到端实战 —— anthropic security 页 critic 92,已是 corpus 第一页。(08) library-grower 把 5+ 同类页蒸馏成候选。(09) design-evolve 棘轮已端到端接线,但还没跑过一轮进化。整套 harness 现在全局装到 ~/.claude/、任意目录可跑(2026-06-15)。

07 · Rules for changing this document修改本文档的规矩

How this doc stays honest. 这份文档怎么保持诚实。

  • Adding a component新增组件 Must state the specific model weakness it defends. No weakness, no seat. 必须写清它挡的是模型的哪一个具体弱点。挡不住弱点,没位置。
  • Removing a component删除组件 Must present evidence the weakness no longer applies. Vibes aren't evidence. 必须给出证据,证明那条弱点已经不成立。感觉不算证据。
  • After every real page每跑完一张真实页 Update the Status today section. New bug class → known-bugs.md + a check. 同步"今天"这一节。新 bug 类 → known-bugs.md + 加 check。
  • If the harness drifts如果 harness 跑偏 Flattens the nine voices, produces noise, gives the generator a way to game the test — write a new Historical lessons entry. 抹平 9 种声音、产生噪音、给 generator 作弊空间 —— 追加"历史教训"。
08 · Historical lessons历史教训

Things we were wrong about. 我们曾经搞错的事。

2026-04-20

Simplicity first is a heuristic, not a hard rule. "simplicity first" 是经验法则,不是硬规则。

While designing this roadmap, I (Claude) once argued for cutting 8 components to 5, citing the article's simplicity first principle. The user correctly pushed back. The article's actual hard rule isn't fewer is better — it's every component must correspond to a specific model weakness. The retreat was unwarranted. Lesson: the only valid argument for cutting is evidence that the weakness doesn't exist. 写这份 roadmap 时,我(Claude)曾经引用 "simplicity first" 把 8 组件砍到 5。用户正确地反问。文章真正的硬规矩不是"越少越好",是"每个组件对应一个具体弱点"。那次退让没有依据。教训:裁组件的唯一有效理由是证据表明弱点不存在

09 · Contribute参与

Think a component is wrong?
Missing? Misranked?
觉得某个组件不对?
漏了?顺序错了?

Argue with evidence. We'd rather restructure the harness than keep a component that isn't really doing anything. 用证据讲话。我们宁愿重构 harness,也不愿让一个没挡任何东西的组件继续占位。

Open a discussion开一个 discussion File an issue提一个 issue Back to sky-skills返回 sky-skills