Table 1: Training settings and computational complexity of selected sparse attention methods. N denotes the sequence length. Complexities include selection and attention computation. ✓ denotes trainable, and × denotes a training-free method. For LLSA [42], which is designed for diffusion models, the listed cost is per attention pass rather than autoregressive prefill.
逐行解读:① HiP(Lee et al., ICLR 2025)是免训练的层次化剪枝服务框架,复杂度同样是 $O(N\log N)$ / $O(\log N)$——注意它是训练无关的,即不改动权重;② HISA(Xu et al., 2026)是免训练的分层索引,预填充仍是 $O(N^2)$;③ MoBA(Lu et al., NeurIPS 2025)与 NSA(Yuan et al., ACL 2025)、HiLS(Hu et al., 2026)都是可训练的,但预填充复杂度仍是 $O(N^2)$,瓶颈正是选择阶段;④ LLSA(Zhou et al., CVPR 2026)已经做到 $O(N\log N)$,但它是为扩散 Transformer 设计的,表里解码一栏是"—";⑤ BSA 在这里指用均值池化摘要的单层选择基线,也就是 PISA 的直接对照组。报告 p.2
Figure 1: Comparison of two different Top-K selection methods for sparse attention. Top: BSA computes a score for each key block before performing Top-K selection. Bottom: PISA constructs a fine-to-coarse hierarchy of key blocks through mean pooling and expands only the selected candidates from coarse to fine during selection. Intermediate blocks are scored using LSE over their child summaries, while leaf blocks are scored using exact LSE over their original keys.
看图要点:上图(BSA)里 $\ell=1$ 那一整条键带全部被标成"待打分候选"(蓝网纹),右侧写着"为所有块计算 Mean Score"——成本正比于块数。下图(PISA)同一个 $\ell=1$ 层里只有两个被橙色虚框圈出的块在打分,其余是"跳过计算"(白框);因为选择是从最上面的大块(橙虚线框,横跨整条键带)开始收窄的,越往上候选越少、越往下才展开。右侧两组说明对应两种打分:中间层用 LSE over child-block(对孩子摘要做 LSE),叶层用 LSE over original keys(对原始键做 LSE)。图例中"Selected by Top-K"是橙色虚框(被保留),"Selected KV"是橙色实心小格(最终真正参与注意力的键),两者一个是选择决策、一个是计算对象。报告 p.4
Table 2: Language modeling and downstream evaluation after 100B tokens of pretraining at 4K. Baselines include NSA [36] and HiLS [14]. Loss is the final training loss, and accuracies are reported as percentages. Gray columns show group averages; boldface marks the best sparse loss and group averages at each scale.
怎么读这张表:横向分三块——困惑度(WikiText / LAMBADA 及其均值,越低越好)、八项多选题及其"Acc-8"均值、六项 containment 检索任务(SWDE / SQuAD / FDA / TQA / NQ / DROP)及其均值。"Containment"的判定标准是预测里包含任一标准答案(不区分大小写的字面子串)即可。三档规模各有一段:PISA 的 loss 在每一档都低于 BSA、PISA-1 与 PISA-2;containment 均值一栏,PISA 在 41.77 / 49.83 / 52.14 三个规模上都是稀疏方法中的最高值(加粗),但仍低于 Full Attention 的 45.15 / 52.13 / 53.01。报告 p.9
Table 6: Training loss and downstream performance for the 1.47B and 2.67B models after 10B tokens of continued pretraining at 16K. Sparse methods, including NSA [36], use C = 64 and K = 32. Loss is the final training loss. Gray columns show group averages; the containment average covers SWDE, SQuAD Completion, and FDA. The block budget does not apply to Full Attention.
两个必须注意的口径差:① 这张表的 containment 均值只覆盖 3 个任务(SWDE / SQuAD / FDA),而 Table 2 是 6 个任务(多了 TQA / NQ / DROP),所以"Table 2 的 52.14"与"Table 6 的 66.40"不能直接比大小,它们不是同一个指标。② 继续预训练后 PISA 的 containment 均值(1.47B 为 65.33、2.67B 为 66.40)低于 BSA(66.55 / 68.15),论文正文也如实写了这一点。同时 PISA 的 loss 在 1.47B(1.7010)与 2.67B(1.6297)上都低于 BSA(1.7039 / 1.6324)与 PISA-1(1.7042 / 1.6331),论文的措辞即为此处。报告 p.20
Table 3: RULER needle-in-a-haystack results after 10B tokens of continued pretraining at 16K for the 2.67B models. Sparse methods use C = 64 and K = 32. Scores are accuracies in percent. The gray Avg column averages scores over the four task families at 1K, 2K, 4K, 8K, and 16K, before rounding. Boldface marks the best sparse result in each column.
读法与两个反直觉之处:每个家族下有 1K/2K/4K/8K/16K 五列,最右的 Avg 是四个家族 × 五个长度的总平均。① 总平均一栏加粗的是 PISA-2 的 62.92,不是完整版 PISA 的 62.80——即在这个指标上"均值 + 半个方差"的二阶近似反而略高一点,差异很小但没有站在完整 LSE 一边。② 单键(niah_single)一栏的加粗几乎全被 NSA 拿走(16K 处 91.67,PISA 为 71.27),PISA 家族的优势集中在多键、多 query、多值这三个家族。③ 所有稀疏方法在 16K 的多键/多 query/多值上都出现明显衰减,最好的稀疏结果分别只有 15.20(多键,NSA)/ 25.20(多 query,PISA-2)/ 28.90(多值,NSA),与 Full Attention 的 23.13 / 34.85 / 57.20 相比,多 query 与多值两栏差距明显。报告 p.10
Figure 2: Block-selection quality on 100 FDA prompts using identical Full-Attention queries and keys. Curves show mean Recall@8 and attention mass ratio by query position. All selectors and the reference share K = 8 and the forced-block policy; Recall@8 includes forced blocks.
曲线解读:左图 Recall@8,右图注意力质量比,横轴是 query 位置(500–2000 token)。橙色实线(完整 PISA)在两张图里都全程位于最上方;蓝点划线是 PISA-2、灰色点线是 PISA-1、黄色虚线是 BSA。三条 LSE/均值线的排序在两张图里一致:PISA > PISA-2 > PISA-1 ≈ BSA(右图里 PISA-1 与 BSA 几乎重合,BSA 甚至略高)。两条曲线都随 query 位置单调下降:序列越长、可选的块越多,任何单层或层次选择器都更难命中全注意力最看重的块。注意 Recall@8 里含三个强制块(首块/前块/当前块),所以曲线起点很高(近 95%),下降部分才反映选择能力。报告 p.10Table 7: Block-selection quality on 100 FDA prompts using shared Full-Attention queries and keys, with C = 64 and K = 8. Scores are percentages, averaged over query rows with more than eight eligible blocks. Recall@8 includes the three forced blocks. Boldface marks the best result in each column.
四个选择器在三个指标上的聚合结果:PISA 拿到全部三项加粗——Recall@8 90.95%(对比 BSA 85.91%、PISA-1 84.75%、PISA-2 88.24%)、保留注意力质量 86.47%(BSA 85.67%)、注意力质量比 99.46%(BSA 98.42%)。这里可以读出两层信息:其一,完整 LSE 明显优于一阶近似 PISA-1(90.95 vs 84.75),说明"保留孩子之间的方差"这个二阶信息确实是选择质量的主要来源;其二,注意力质量比一栏所有方法都在 98% 以上,说明被选中的块几乎能覆盖参考集所覆盖的注意力质量,各选择器在这一指标上的差距被压缩了。报告 p.21
Figure 3: Prefill block-selection latency from 4K to 256K with C = 64 and K = 8. Timings include mean-summary construction and all selection stages, but exclude attention over the selected blocks. The key-block reuse implementation of PISA uses Q_tile = 4; annotations show speedups over BSA. Complete settings are provided in Appendix B.
三条曲线与交叉点:黄色虚线是 BSA,灰色点线是 PISA 的 per-query 融合实现,橙色实线是键块复用实现($Q_{\text{tile}}=4$,即论文的 Q4)。① 交叉点是这张图最重要的信息:4K–16K 区间 BSA 反而更快(16K 处 BSA 1.461 ms vs Q4 1.581 ms),从 32K 起 Q4 反超并一路拉开,64K/128K/256K 分别快 2.86×/5.31×/9.95×。② BSA 在长序列上的曲线呈明显上翘(256K 达 312.96 ms),正是 $O(N^2/C)$ 项在起作用。③ 键块复用相对 per-query 实现的增益在 64K–256K 稳定在 1.30×–1.35×,幅度不大但方向一致——这正是 §5.1 里 $G_Q+C/Q_{\text{tile}}=32<64$ 那个不等式带来的收益。报告 p.11Table 4: Fixed-length block-selection latency from 4K to 256K. Q1, Q2, and Q4 reuse one leaf-key tile across up to one, two, and four query references, respectively. Boldface marks the lowest latency at each sequence length.
逐列对照:Q1/Q2/Q4 是键块复用程度不同的三个实现(每个叶键块分别服务 1/2/4 个 query)。加粗位置标出了每行的最优:4K、8K、16K 三行是 BSA 最快(0.167007 / 0.462596 / 1.461185 ms);32K 到 256K 四行是 Q4 最快(3.184846 / 6.704480 / 14.452411 / 31.440747 ms)。另外注意 4K 处 Q4(0.701181 ms)比 per-query(0.573228 ms)慢——论文正文明确解释了原因:序列短时候选块少,"把 query 分组以复用键块"的组织开销超过了复用带来的收益,从 8K 起才转为正收益。报告 p.18