Q2 2026 · Infrastructure report基础设施报告

More traffic, less latency, one bad Tuesday. 流量更多,延迟更低,外加一个糟糕的周二。


Prism's quarterly engineering letter on the infrastructure behind 2.4 million events a second: where the traffic came from, how p99 fell to 38ms, what broke on June 9, and what we are building before the next doubling. Prism 每季一封的工程信,写给支撑每秒 240 万事件的那套基础设施:流量从哪里来,p99 怎么降到 38 毫秒,6 月 9 日坏掉了什么,以及在下一次翻倍之前我们在建什么。

2026-07-08 · Infrastructure team · v2.1 · 14 min read 2026-07-08 · 基础设施团队 · v2.1 · 14 分钟读完

00 · Executive summary执行摘要

The quarter in three numbers. 三个数字,读完这个季度。

2.4M peak events ingested per second每秒接入事件峰值 +41% vs Q1较 Q1 +41%
38ms p99 query latency at quarter close季末查询延迟 p99 −38% vs Q1较 Q1 −38%
99.99% delivery, incl. the June 9 incident送达率(含 6 月 9 日故障) SLO 99.95% · metSLO 99.95% · 达标

Q2 closed with peak ingest 41% above Q1 and p99 query latency 38% below it — the two lines crossed without a single capacity incident on the query path. Delivery held at 99.99% for the quarter, including the 23 minutes on June 9 when a transit fiber cut degraded eu-central; the buffer-and-replay design turned a regional outage into late events, not lost ones. Q2 收官时,峰值接入比 Q1 高 41%,查询延迟 p99 比 Q1 低 38%——两条曲线交错而过,查询路径上没有发生一起容量事故。整季送达率保持在 99.99%,包括 6 月 9 日 eu-central 因中转光纤被挖断而降级的那 23 分钟;缓冲重放的设计把一次区域故障变成了迟到的事件,而不是丢失的事件。

01 · Traffic growth流量增长

Ingest grew 41%. The pipes barely noticed. 接入量涨了 41%,管道几乎没察觉。

Peak ingest climbed from 1.72M to 2.43M events per second across thirteen weeks. About two thirds of the growth came from existing teams instrumenting more services — the cohort that joined in Q1 doubled its event volume by mid-June — and the rest from 214 new teams onboarded during the quarter. 十三周里,峰值接入从每秒 172 万爬到 243 万。增长约三分之二来自既有团队接入更多服务——Q1 加入的那批团队到 6 月中旬把事件量翻了一倍——其余来自本季度新接入的 214 个团队。

The edge held without new hardware: batching agents absorbed the growth while ingest lag stayed flat, and the only visible artifact of 41% more traffic is the area under this curve. The June 9 dip is the §03 incident — one week of flattened peaks, not a hole. 边缘层没有添置任何新硬件:批量 agent 把增长吸收进恒定的接入延迟里,41% 的流量增长在页面上唯一可见的痕迹,就是这条曲线下方的面积。6 月 9 日那次故障在 §03 详述——在这张图上只是一周被压平的峰值,而不是一个洞。

2.5M 2.0M 1.5M JUN 9 · INCIDENT 2.43M PEAK APR MAY JUN WEEKLY PEAK · 60S SUSTAINED
Weekly peak ingest, Q2 2026 — 1.72M to 2.43M events/s. The June 9 incident reads as one flattened week, not a hole: buffered events replayed within the same hour. Q2 2026 每周接入峰值——从每秒 172 万到 243 万。6 月 9 日的故障在图上只是被压平的一周,不是一个洞:缓冲的事件在同一小时内完成了重放。
02 · Latency延迟改善

p99 fell from 61ms to 38ms. p99 从 61 毫秒降到 38 毫秒。

Three changes shipped, in order of measured impact: the vectorized engine v2.4 rollout (−9ms, fleet-wide by April 18), the hot tier doubling to 36 hours of in-memory data (−8ms, May), and scatter-gather pruning that skips regions a query provably cannot touch (−6ms, June). 三项改动按实测收益排序:向量化引擎 v2.4 全量上线(−9 毫秒,4 月 18 日覆盖全集群)、热层扩容到 36 小时全内存数据(−8 毫秒,5 月)、以及跳过查询确定触不到的区域的 scatter-gather 剪枝(−6 毫秒,6 月)。

The spread between regions narrowed too: the slowest region today (ap-southeast, 44ms) beats the fastest region of Q1 (eu-central, 57ms). The H2 latency budget moves from "keep p99 under 45ms" to "keep the regional spread under 10ms". 区域之间的差距也在收窄:今天最慢的区域(ap-southeast,44 毫秒)已经快过 Q1 最快的区域(eu-central,57 毫秒)。H2 的延迟预算从「p99 压在 45 毫秒以内」改成「区域间差距压在 10 毫秒以内」。

Q1 P99 Q2 P99 60ms 40ms 20ms 36 39 35 38 44 41 us-east us-west eu-central eu-west ap-southeast sa-east
p99 query latency by region, Q1 (indigo) vs Q2 (cyan). Every region improved 35–40%; the spread tightened from 14ms to 9ms. 各区域查询延迟 p99,Q1(indigo)对比 Q2(cyan)。每个区域都改善了 35–40%;区域间差距从 14 毫秒收窄到 9 毫秒。
03 · Incident review故障复盘

June 9: a fiber cut, 23 minutes, zero events lost. 6 月 9 日:一次光纤被挖断,23 分钟,零事件丢失。

At 14:01 UTC a construction crew cut a transit fiber bundle outside Frankfurt, taking 40% of eu-central's upstream capacity with it. Delivery lag in the region rose above the 5-second alert line within 38 seconds; the page reached the on-call at 14:02. What followed is the timeline below — and the reason the quarterly delivery number still reads 99.99%. UTC 时间 14:01,一支施工队在法兰克福城外挖断了一束中转光纤,eu-central 40% 的上游容量随之消失。38 秒内,该区域的送达延迟越过 5 秒告警线;14:02 值班工程师收到呼叫。下面是接下来发生的事——也是季度送达率依然写着 99.99% 的原因。

23 MIN DEGRADED 14:01 · CUT 14:02 14:09 14:25 Delivery lag alert fires Reroute — edge buffers engaged Restored — lag under 250ms 38s after the cut · page sent 31M events held at the edge 23 min after detection 14:06 14:18 Transit fiber cut confirmed Backlog replay begins 40% of upstream capacity lost drain at 2.1× line rate ALL TIMES UTC · JUN 9 2026 ZERO EVENTS LOST · 31M REPLAYED BY 15:10
June 9 incident timeline. Detection to full restore in 23 minutes; edge buffers held 31M events and replayed them by 15:10. 6 月 9 日故障时间线。从发现到完全恢复 23 分钟;边缘缓冲存住 3100 万个事件,并在 15:10 前完成重放。
23 min degraded降级 23 分钟

What users saw用户看到了什么

For 23 minutes, events entering through eu-central's affected transit paths — about 4% of global traffic at that hour — were delivered late, with delivery lag peaking at 41 seconds. Queries stayed up: the read path failed over in 9 seconds. Zero events were lost — 31 million were buffered at the edge and replayed by 15:10. 23 分钟里,经由 eu-central 受影响中转线路接入的事件——约占当时全球流量的 4%——被延迟送达,送达延迟峰值 41 秒。查询没有任何中断:读路径在 9 秒内完成切换。零事件丢失——3100 万个事件在边缘缓冲,并在 15:10 前完成重放。

3 actions · H23 项改动 · H2

What we are changing我们要改什么

  • A second transit provider for eu-central, live July 22 — no single fiber path carries more than half the region's upstream again. 为 eu-central 引入第二家中转供应商,7 月 22 日上线——不再有任何一条光纤线路承载该区域一半以上的上游流量。
  • Reroute on alert, not on confirm. Failover now triggers from the lag alert itself — June 9 spent 7 of its 23 minutes on a human confirming what the graphs already said. 改为「告警即切换」,不再等人工确认。切换现在直接由延迟告警触发——6 月 9 日的 23 分钟里,有 7 分钟花在人工确认图表早已说明的事情上。
  • Quarterly cut drills: once a quarter we sever a transit link on purpose, in production — the only failover that works is the one that runs. 每季度一次断链演练:每个季度在生产环境里主动切断一条中转链路——只有真正跑过的切换,才算可用的切换。
04 · Capacity planning容量规划

Planning for the next doubling, not the next quarter. 为下一次翻倍做规划,而不是下一个季度。

Fleet ceiling today is 4.36M events/s — 1.8× the Q2 peak, which is one strong quarter of margin, not two. Three expansions land before October and lift the ceiling to 6.3M: a new ingest cell in eu-central (July, paired with the second transit provider), a new cell in ap-southeast (August, the fastest-growing region), and a third cell row in us-east (September). At 25% quarterly growth, that holds a ≥1.3× margin through Q1 2027. 目前集群的容量上限是每秒 436 万事件——是 Q2 峰值的 1.8 倍,这只够一个高速增长季度的余量,而不是两个。10 月之前会落地三项扩容,把上限抬到 630 万:eu-central 的新接入单元(7 月,与第二家中转供应商一起上线)、ap-southeast 的新单元(8 月,增长最快的区域)、以及 us-east 的第三排单元(9 月)。按每季 25% 的增长计算,这能把 ≥1.3 倍的余量保持到 2027 年 Q1。

Region区域 Q2 peak · ev/sQ2 峰值 · ev/s Ceiling · ev/s容量上限 · ev/s Headroom余量 H2 planH2 计划
us-east690K1.20M1.7× expand cell · Sep扩建单元 · 9 月
us-west410K800K2.0× —
eu-central520K760K1.5× new cell + 2nd transit · Jul新单元 + 第二中转 · 7 月
eu-west350K720K2.1× —
ap-southeast300K480K1.6× new cell · Aug新单元 · 8 月
sa-east160K400K2.5× —
Fleet全集群 2.43M4.36M1.8× → 6.3M by Oct10 月前 → 630 万
Capacity by region, end of Q2. Ceilings come from quarterly load tests, not spec math — each cell is pushed to failure in a drill. Q2 期末各区域容量。上限来自每季度的压测,不是纸面推算——每个单元都会在演练中被压到失效点。
Methodology统计口径

Latency is measured at the query API edge and includes queue time; percentiles are computed over all production queries in the stated window, excluding synthetic monitors. Ingest peaks are 60-second sustained rates, not instantaneous. An event counts as delivered when every subscribed sink acknowledges it; events later than 5 seconds count as late but delivered. Q1 baselines use identical definitions; raw notebooks are linked from the internal edition of this report. 延迟在查询 API 边缘测量,包含排队时间;分位数基于所述时间窗内全部生产查询计算,剔除拨测流量。接入峰值为 60 秒持续速率,不是瞬时值。送达的定义是事件被所有订阅下游确认接收;晚于 5 秒的事件计为迟到但已送达。Q1 基线使用完全相同的口径;原始 notebook 链接在本报告的内部版本中。