Architecture · 2026-04-20架构 · 2026-04-20

The Harness.
Nine components.
One rule.
架构蓝图。
九个组件。
一条规矩。

The full design of the sky-skills harness — nine generator skills and two evaluators with four review gates. Now all nine are built; eight are done — the canonical library (02) just reached the full 59/59 matrix — and design-evolve (09) has not run an evolution round yet. Every component must defend one question: what specific model weakness does it correspond to? sky-skills 架构的完整蓝图 —— 九个 生成器 skill 加两个 评审员(4 道审查检查)。现在九个全部建好;八个已完成——canonical 库(02)刚补满 59/59 整张覆盖表——design-evolve(09)还没跑过一轮进化。每个组件都必须能回答一个问题:它挡的是模型的哪一个具体弱点?

See the nine看这九个 What's done today今天做到哪里
Read in another voice用另一种声音读 apple anthropic ember sage glass

23 focused skills — 18 generate, 2 judge, 1 plans, 1 evolves, 1 keeps them current精选 skill —— 18 个生成,2 个评审,1 个规划,1 个自进化,1 个保持更新
9 harness components — each defending one weaknessharness 组件 —— 各挡一个弱点
9 design voices — this roadmap is rendered in five of them设计声音 —— 这份路线图用其中五种各渲一版

Sources指导来源

Two ideas we're building on. 我们站在的两个肩膀。

GAN (Goodfellow et al., 2014) — a generator and a discriminator train adversarially. The sharper the discriminator, the stronger the generator becomes. Our nine design skills are the generator, design-review is the discriminator.

GAN(Goodfellow et al., 2014) —— generator 和 discriminator 对抗训练。discriminator 越锐利,generator 就越强。我们的 9 个设计 skill 是 generator,design-review 是 discriminator。


Anthropic, Harness design for long-running appsplanner / generator / evaluator three-part split, evaluators must be separate, sprint contracts, 5–15 iteration rounds. The rule we take most seriously: each harness component must correspond to one specific model weakness. Assumptions worth stress-testing, because they may be incorrect or quickly outdated.

Anthropic 的 Harness design for long-running apps —— planner / generator / evaluator 三段式,evaluator 必须独立,sprint contract,5–15 轮迭代。我们最认真对待的一条:每一个 harness 组件都要对应模型的一个具体弱点。这些假设值得压测,因为它们可能不对,或者很快过时。

Agents tend to praise their own work confidently — even when quality is obviously mediocre to a human observer. Separating the agent doing the work from the agent evaluating it is a powerful lever. Agent 倾向于自信地称赞自己的作品 —— 哪怕在人类观察者看来质量明显平庸。把做事的 agent 和评判的 agent 分开,是一个真正的杠杆。

Anthropic — Harness design postAnthropic —— harness-design 文章

Goal目标

A page that stands next to
anthropic.com, stripe.com, linear.app.
一页能和 anthropic.com、
stripe.com、linear.app 同台站着。

Not technically bug-freetaste, originality, rhythm, copy, hand-crafted illustrations, device-agnostic, accessible, shareable. From a one-sentence brief. 不是"技术上没 bug",而是 有品位、有原创、有节奏、copy 过关、图示手工感强、全设备可用、无障碍、可分享。从一句话需求开始。

Nine components九个组件

Every row defends a weakness. 每一行都挡一个弱点。

The only legitimate reason for a harness component to exist is that a specific model weakness makes it necessary. No weakness, no seat. 一个 harness 组件能存在的唯一理由,是有一条具体的模型弱点要求它存在。挡不住弱点,没位置。

future待做 01

design-planner

skills/design-planner/ Model weakness对应弱点

Faced with a vague brief, the model fills in training-data medians. Scope drifts. A planner expands the brief into a sprint contract before the generator runs. 面对模糊需求,模型用训练数据的中位数填空,scope 失控。planner 在生成器动手前,把需求展开成 sprint contract

Done · 59/59已完成 · 59/59 02

reference library参考库

skills/<style>-design/references/canonical/ Model weakness对应弱点

Without reference, output regresses to the most average web design. A per-style canonical library of 10 page-types anchors quality at each style's own best — not an industry average. 没有参考,输出向"最平均的 web 设计"收敛。每个 style 自己的 10 个 page-type canonical,把品质锚定在该 style 自己的最好。

partly done部分完成 03

generator + self-diff生成器 + 自评差异

skills/{apple,anthropic,ember,sage}-design/ Model weakness对应弱点

Models don't naturally articulate their design decisions after writing. Without a required self-diff note, the critic has no concrete target to argue with — just vibes. 模型写完 HTML 后不会自然记录它的设计决策。没有强制的 self-diff note,critic 没具体靶子。

phase 1 done · expanding期 1 已完成 · 扩展中 04

mechanical-review

skills/design-review/ Model weakness对应弱点

Models don't self-check runtime behavior — multi-viewport, keyboard, color contrast, LCP/CLS, SEO meta. A mechanical gate catches what source-scanning and self-review miss. runtime 行为(多视口、键盘、对比度、LCP/CLS、SEO meta)模型不自检。mechanical 检查抓源扫和自评都抓不到的那类。

done · 2026-04-22已完成 · 2026-04-22 05

multi-critic多位专家评审员

.claude/agents/design-{composition,copy,illustration,brand}-critic.md Model weakness对应弱点

One critic spreads attention across 7 dimensions and underweights any one. Four specialists (weights 25/25/20/30) run in parallel fresh contexts — composition, copy, illustration, brand — then an aggregator merges. First live run: solo 93, multi 88; illustration caught a 6th-hue SVG leak the generalist missed. 一个 critic 把注意力摊给 7 维,每一维都浅。4 位专家(权重 25/25/20/30)并行 fresh context —— 版式、copy、插画、品牌 —— 再由 aggregator 合成。首次实战:单评 93,多评 88;插画专家抓到了通才漏掉的 SVG 6 色偷渡。

shipped · first run critic 92已发 · 首跑 critic 92 06

/design-loop

.claude/commands/design-loop.md Model weakness对应弱点

Models default to ship the first-pass output. Without orchestration, actual iteration ends after 1–2 rounds. A loop controller forces up to 5 rounds, or escalates. 模型默认"第一版够用就发"。没编排,实际只跑 1-2 轮。loop controller 强制最多 5 轮,超过交人。

done · 2026-04-22已完成 · 2026-04-22 07

learning-loop

.claude/agents/design-learner.md + scripts/learning-loop.mjs Model weakness对应弱点

"The new issue the critic caught today" doesn't auto-archive. design-learner reads critic verdicts and proposes 3 codifications — new known-bugs rows, new visual-audit checks, per-skill dos-and-donts. First run 2026-04-22: 13 raw issues → 5 bug classes → 3 known-bugs rows (1.17/1.18/1.19) + 1 new check + 8 dos-and-donts. Same bug never caught twice. "critic 今天新抓到的问题"不会自动归档。design-learner 读 verdict,提三件事:known-bugs 新行、visual-audit 新 check、每 skill 的 dos-and-donts。首次实战 2026-04-22:13 条 → 5 类 → 3 行 known-bugs + 1 个新 check + 8 条 dos-and-donts。同一个 bug 不再抓第二次。

future · after 10+ pages待做 · 10+ 真实页后 08

library-grower

.claude/commands/design-distill.md Model weakness对应弱点

Models can't self-organize yesterday's good output into tomorrow's reference. Without this, the canonical library stays fixed and stale over 6 months. 模型没有"把昨天好产物整理成明天参考"的能力。没有它,canonical library 会固定,6 个月后过期。

Built · no round yet机制已建 · 未跑 09

design-evolve

skills/design-evolve/SKILL.md Model weakness对应弱点

Models repeat the same taste mistakes and can't tell which of their own rules actually help. A Darwin / autoresearch ratchet: mutate one generator rule, regenerate, score with the frozen evaluator, keep it only if it strictly beats the baseline with no held-out canonical regressing. Selection, not authorship. 模型会重复同样的审美错误,也分不清自己哪条规则真有用。一个 Darwin / autoresearch 式爬坡:改一条生成器规则、重新生成、用冻结的评审器打分,只有严格超过基线且无 held-out canonical 变差才保留。由选择而非作者决定。

Where it stands当前进度

Nine components, by status. 九个组件,按当前状态。

DONE · 7 JUST COMPLETED · 1 BUILT · NO ROUND YET · 1 Done — the pipeline is built 01 plannersprint-contract + --plan 03 generator self-diff§M + verify.py gate 04 mechanical-review51 checks · 94 known-bugs 05 multi-critic4 subagents + solo critic 06 /design-loop ran · critic 92first real end-to-end run 07 learning-loop 08 library-growerdesign-learner · distill 5+ to candidate Just completed 02 canonical library 58 / 58 page-types covered matrix complete — 10 page-types × anthropic / apple / ember / sage + 18 more run --coverage for the live count Built · no round yet 09 design-evolve Darwin / autoresearch ratchet, wired: evolve-ledger.mjsevolve-rules.mjsregression-gate.mjs rules.json empty · 0 rounds run ledger = the loop's first-run deposit when it runs, the evaluator stays frozen
The harness at a glance — eight components done, the canonical library (02) now complete at the full 59/59 matrix, and design-evolve (09) wired end-to-end but yet to run its first evolution round. harness 全貌 —— 八个组件已完成,canonical 库(02)已补满 59/59 整张覆盖表,design-evolve(09)已端到端接线但还没跑过第一轮进化。
Hard constraint硬约束

The harness must not flatten the nine voices. harness 越复杂,越怕抹平九种声音。

Universal rules never judge style; style rules live in per-style folders; critic's taste is anchored to each style's own canonical. universal 规则不评判风格,风格规则住在各自文件夹,critic 的品味锚定在每个 style 自己的 canonical。

Layer Where it lives住在哪 Flattens the voices?会抹平声音吗
Universal quality通用质量 skills/design-review/references/cross-skill-rules.md No不会
Only judges does it work.只管"能不能用"
Cross-skill bugs跨 skill bug 清单 skills/design-review/references/known-bugs.md No不会
Segmented into per-style sections.按分节组织
Style canonicalsstyle canonical 例子 skills/<style>-design/references/canonical/ Strengthens them反过来强化
Apple pricing and sage pricing look nothing alike. Libraries are never shared.apple pricing 和 sage pricing 长得完全不一样。不互抄
Style personality风格个性 skills/<style>-design/references/dos-and-donts.md Strengthens them反过来强化
Palette, fonts, signature moves — all per-style.色板、字体、签名动作 —— 全是各自的
Critic's tastecritic 的品味 Subagent reads this style's canonical + dos-and-dontssubagent 读 当前 style 的 canonical + dos-and-donts No不会
Critic always asks is this a good <style> page?critic 永远问"这是不是一张好的 <style>; 页"
Status · 2026-06-11今天 · 2026-06-11

What's real. What's next. 哪些是真的。下一步。

done 8 (incl. 02) · built-no-round 1 (09)已完成 8(含 02) · 机制已建未跑 1(09)

What's actually done 实际已完成

51-check visual-audit.mjs + 94 known-bugs catalogued. 59/59 canonical page-types covered — the full matrix, each with a binding .md rubric. multi-critic — 4 specialist subagents in parallel (composition / copy / illustration / brand) plus solo design-critic. learning-loopdesign-learner turns critic misses into new mechanical checks. 23 skills total: 9 design generators, 7 systems / content utilities, 2 workflow generators (gated-dual-clone, doc-review-loop), 2 independent evaluators (design-review, gated-dual-clone-audit), 1 planner (design-planner), 1 self-evolution loop (design-evolve), 1 skills auto-updater (skills-sync). 2026-04-27: anthropic-design v2 (scenario recipes + ux-writing + recipe components in anthropic.css), bin/design-review --audit batch mode (file / dir / URL). 2026-05-22: anthropic-design v3 — the md rendering four-piece (md-mirror / md-rewrite-links / md-pack / cross-link-pack): point it at a markdown directory, get styled HTML whose links survive cp -r anywhere. 2026-06-11: diagram-craft v3 — seven kernel diagram types, constraints in three layers (aesthetics immutable, quality machine-gated, structure yours to shape); 5 new visual-audit.mjs gates grown from known-bugs 1.28–1.31; sprint-contract learns §1b diagram density; templates at anthropic 14 / apple 14; two diagram galleries (23 + 14); the gold tuned into place — #c49464#c9913f, across every skill. Wave 2 follows close behind — visual-audit grows keyboard-a11y and LCP/CLS perf gates (38 checks now), verify.py learns a warn tier with SEO meta checks, --audit reports per-kind stats, --distill gets wired in (component 08) — and design-planner arrives as the 13th skill, component 01 filling out. The craft travels on the same day — ember-design and sage-design each take up diagram-craft: 8 templates apiece, an 8-figure gallery each, warm browns with a single gold focus on one side, green with indigo ink on the other; the monochrome gate now watches anthropic, ember and sage, while apple keeps its grayscale exemption. 51 项 visual-audit.mjs + 94 条 known-bugs。覆盖 59/59 canonical page-type —— 整张覆盖表,每张配绑定 .mdmulti-critic —— 4 位专家 subagent 并行(版式 / 文案 / 插画 / 品牌)+ 单 design-criticlearning-loop —— design-learner 把 critic miss 转成新机器 check。23 个 skill:9 个设计 generator、7 个系统 / 内容工具、2 个 workflow generator(gated-dual-clonedoc-review-loop)、2 个独立 evaluator(design-reviewgated-dual-clone-audit)、1 个 planner(design-planner)、1 个自进化 loop(design-evolve)、1 个 skills 自动更新器(skills-sync)。2026-04-27:anthropic-design v2(scenario recipes + ux-writing + recipe components 入 anthropic.css)、bin/design-review --audit 批量模式(file / dir / URL)。2026-05-22anthropic-design v3 —— md 渲染 4 件套(md-mirror / md-rewrite-links / md-pack / cross-link-pack),指向一个 md 目录,得到一套风格化 HTML,cp -r 到哪链接都不断。2026-06-11diagram-craft v3 —— 内核七图型 + 约束三层化(审美不可变 / 质量机器检查 / 结构自由定制);known-bugs 1.28–1.31 长成 visual-audit.mjs 5 个新检查;sprint-contract 新增 §1b 图密度;模板库 anthropic 14 件 / apple 14 件;双 diagram gallery(23 图 + 14 图);金主色调准 —— #c49464#c9913f,全 skill 同步。Wave 2 紧随其后 —— visual-audit 长出 keyboard 可达性与 LCP/CLS 性能检查(现在 38 项 check),verify.py 学会 warn 档并带上 SEO meta 检查,--audit 按 kind 出统计,--distill 接好了线(组件 08)—— design-planner 作为第 13 个 skill 到来,组件 01 渐渐补全。同日工艺继续延伸 —— ember-designsage-design 各自领到 diagram-craft:各 8 件模板、各 8 图画廊,一边暖棕配金单点,一边抹茶绿配靛蓝墨;monochrome 检查看住 anthropic、ember 与 sage,apple 守着灰阶身份得以豁免。

now · critic validated现在 · critic 已验证

4 real pages reviewed 评审了 4 张真实页

2026-04-22: all 4 HARNESS-ROADMAP variants went through bin/design-review --critic. anthropic 84 → 93 in three rounds. apple 88, ember 84, sage 78 → pass after fixes. Each verdict cites line numbers, class names, paste-ready CSS — specific and actionable. Each fix cycle pins an §H / §J / §K regression into the next visual-audit run. 2026-04-22:4 张 HARNESS-ROADMAP 都跑了 bin/design-review --critic。anthropic 三轮 84 → 93;apple 88,ember 84,sage 78 → 修完通过。每条 verdict 直接给行号、类名、可粘贴 CSS —— 具体可操作。每次修复都把 §H / §J / §K 回归锁进下一次 visual-audit。

shipped · 2026-06-14/15已发 · 2026-06-14/15

loop · grower · evolve — all built loop · grower · evolve 全建好

(06) /design-loop shipped its first real end-to-end run — an anthropic security page at critic 92, now the corpus's first page. (08) library-grower distills 5+ same-type pages into a candidate. (09) design-evolve has the full ratchet wired but has not run an evolution round yet. And the whole harness now installs globally under ~/.claude/ and runs from any directory (2026-06-15). (06) /design-loop 完成首次端到端实战 —— anthropic security 页 critic 92,已是 corpus 第一页。(08) library-grower 把 5+ 同类页蒸馏成候选。(09) design-evolve 棘轮已端到端接线,但还没跑过一轮进化。整套 harness 现在全局装到 ~/.claude/、任意目录可跑(2026-06-15)。

Rules for changing this document修改本文档的规矩

How this doc stays honest. 这份文档怎么保持诚实。

  • Adding a component新增组件 Must state the specific model weakness it defends. No weakness, no seat. 必须写清它挡的是模型的哪一个具体弱点。挡不住弱点,没位置。
  • Removing a component删除组件 Must present evidence the weakness no longer applies. Vibes aren't evidence. 必须给出证据,证明那条弱点已经不成立。感觉不算证据。
  • After every real page每跑完一张真实页 Update the Status today section. New bug class → known-bugs.md + a check. 同步"今天"这一节。新 bug 类 → known-bugs.md + 加 check。
  • If the harness drifts如果 harness 跑偏 Flattens the nine voices, produces noise, gives the generator a way to game the test — write a new Historical lessons entry. 抹平 9 种声音、产生噪音、给 generator 作弊空间 —— 追加"历史教训"。
Historical lessons历史教训

Things we were wrong about. 我们曾经搞错的事。

2026-04-20

Simplicity first is a heuristic, not a hard rule. "simplicity first" 是经验法则,不是硬规则。

While designing this roadmap, I (Claude) once argued for cutting 8 components to 5, citing the article's simplicity first principle. The user correctly pushed back. The article's actual hard rule isn't fewer is better — it's every component must correspond to a specific model weakness. The retreat was unwarranted. Lesson: the only valid argument for cutting is evidence that the weakness doesn't exist. 写这份 roadmap 时,我(Claude)曾经引用 "simplicity first" 把 8 组件砍到 5。用户正确地反问。文章真正的硬规矩不是"越少越好",是"每个组件对应一个具体弱点"。那次退让没有依据。教训:裁组件的唯一有效理由是证据表明弱点不存在

Think a component is wrong?
Missing? Misranked?
觉得某个组件不对?
漏了?顺序错了?

Argue with evidence. We'd rather restructure the harness than keep a component that isn't really doing anything. 用证据讲话。我们宁愿重构 harness,也不愿让一个没挡任何东西的组件继续占位。

Open a discussion开一个 discussion File an issue提一个 issue Back to sky-skills返回 sky-skills