Longformer(Long Document Transformer)

Longformer(Long Document Transformer)

Longformer(Long Document Transformer)是由 Allen Institute for AI(AI2)在 2020 年提出的一种 Transformer 变体,专门用来高效处理长文本。它的设计核心是 线性复杂度的注意力机制,以此克服标准 Transformer 处理长文本时的计算瓶颈。

在讲它怎么改之前,我们先做个实验,分别输入不同的文本长度,大致感受下模型的计算量。实验用的是 ptflops 库里的 get_model_complexity_info 方法。

from transformers import BertModel
from transformers import BertConfig
from ptflops import get_model_complexity_info


def test():
    # 初始化 Bert 模型
    model = BertModel(config=BertConfig())
    # 获得 Bert 自注意力计算对象
    self_attn = model.encoder.layer[0].attention.self
    print(self_attn)
    print('-' * 50)

    # 计算输入长度为 512 时,模型的计算量
    flops, params = get_model_complexity_info(model=self_attn, input_res=(512, 768))
    print(flops, params)
    print('-' * 50)

    # 计算输入长度为 1024 时,模型的计算量
    flops, params = get_model_complexity_info(model=self_attn, input_res=(1024, 768))
    print(flops, params)


if __name__ == '__main__':
    test()

程序执行结果:

BertSelfAttention(
  (query): Linear(in_features=768, out_features=768, bias=True)
  (key): Linear(in_features=768, out_features=768, bias=True)
  (value): Linear(in_features=768, out_features=768, bias=True)
  (dropout): Dropout(p=0.1, inplace=False)
)
--------------------------------------------------
BertSelfAttention(
  1.77 M, 100.000% Params, 905.97 MMac, 100.000% MACs, 
  (query): Linear(590.59 k, 33.333% Params, 301.99 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (key): Linear(590.59 k, 33.333% Params, 301.99 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (value): Linear(590.59 k, 33.333% Params, 301.99 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (dropout): Dropout(0, 0.000% Params, 0.0 Mac, 0.000% MACs, p=0.1, inplace=False)
)
905.97 MMac 1.77 M
--------------------------------------------------
BertSelfAttention(
  1.77 M, 100.000% Params, 1.81 GMac, 100.000% MACs, 
  (query): Linear(590.59 k, 33.333% Params, 603.98 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (key): Linear(590.59 k, 33.333% Params, 603.98 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (value): Linear(590.59 k, 33.333% Params, 603.98 MMac, 33.333% MACs, in_features=768, out_features=768, bias=True)
  (dropout): Dropout(0, 0.000% Params, 0.0 Mac, 0.000% MACs, p=0.1, inplace=False)
)
1.81 GMac 1.77 M

从结果可以看到,单条样本输入长度为 512 时,计算量是 905 MMac,长度改成 1024 后,计算量变成 1.81 GMac。MAC 和 FLOPs 一样,都是计量模型浮点运算量的单位。MAC 指 Multiply–Accumulate Operations,即把一次浮点乘加算作一个 MAC,简单理解就是 1 MAC 相当于 2 FLOPs。MMAC 就是 100 万次 MAC,GMAC 就是 10 亿次 MAC。

输入长度翻了一倍,计算量也差不多翻了一倍,可见自注意力的计算量对输入长度非常敏感。为了高效处理长文本序列,Longformer 改进了自注意力的计算方法。

Paper:https://arxiv.org/pdf/2004.05150.pdf

另外一点思考:原始的自注意力机制计算时,每个 Token 都会关注到输入的所有其他 Token。是否有必要全部关注也是一个待思考的问题。

1. Sliding Window Attention

Longformer 提出了局部自注意力,也就是算注意力时只关注窗口内的 token,而不是所有 token。这显然能降低计算复杂度。论文里把这种方式叫做 sliding window attention,即基于滑动窗口的注意力。

这里的 window size 该怎么设置呢?论文里给了一句话:

Depending on the application, it might be helpful to use different values of w for each layer to balance between efficiency and model representation capacity.

简单说,就是让不同的编码器层用不同的 window size。比如从一个较小的窗口逐渐增大到较大的窗口,相当于一层一层地扩大感受范围。

2. Dilated Sliding Window Attention

为了在不增加计算量的前提下扩大 token 的关注范围,论文又用了扩张的滑动窗口,如下图所示。

从图里一眼就能看明白,扩张其实就是跳跃着关注 token,这样覆盖范围确实变大了。论文针对这一点又说了一段话:

In multi-headed attention, each attention head computes a different attention score. We found set- tings with different dilation configurations per head improves performance by allowing some heads without dilation to focus on local context, while others with dilation focus on longer context.

简单讲,我们不是有多头注意力吗?那就让不同的 head 用不同的扩张配置。比如第一个头每隔 2 个 token 关注一次,第二个头每隔 8 个,第三个头每隔 6 个,以此类推。这样搭配能提高模型性能。

3. Global Attention + Sliding Window

到这里,Longformer 的局部注意力好像就讲完了。但我们马上会发现一个问题。原来用 BERT 时,靠 [CLS] 来表征整个输入序列,如果对 [CLS] 也用局部注意力,它可能就没法很好地代表整段序列了。所以对这类特殊 token,我们仍然用全局注意力,也就是原来那种关注所有 token 的算法。

于是规则就成了:[CLS] 用全局注意力,关注所有 token,其他 token 除了看窗口内的,也都要关注 [CLS]。

we make this attention operation symmetric: that is, a token with a global attention attends to all tokens across the sequence, and all tokens in the sequence attend to it.

注意,上面只举了 [CLS] 用全局注意力的例子,但实际输入里可能还有别的位置也需要全局注意力,我们可以预先把这些位置标记出来。比如分类任务只对 [CLS] 用全局注意力就够了,而 QA 任务可能需要对整个 question 的 token 都开全局注意力。