PyTorch——自注意力（self-attention）机制实现（代码详解）

2023年6月17日上午1:15 • 人工智能 • 阅读 62

参考链接

https://www.bilibili.com/video/BV1JE411g7XF?p=54
https://arxiv.org/abs/1706.03762
https://blog.csdn.net/qq_36653505/article/details/83375160

简述自注意力机制（self-attention）

self-attention可以视为一个特征提取层，给定输入特征a 1 , a 2 , ⋅ ⋅ ⋅ a n a^{1},a^{2},\cdot \cdot \cdot a^{n}a 1 ,a 2 ,⋅⋅⋅a n，经过self-attention layer，融合每个输入特征，得到新的特征b 1 , b 2 , ⋅ ⋅ ⋅ b n b^{1},b^{2},\cdot \cdot \cdot b^{n}b 1 ,b 2 ,⋅⋅⋅b n。具体如下：

设输入特征为I I I，分别将其乘以三个矩阵W q W^{q}W q、W k W^{k}W k和W v W^{v}W v得到Q Q Q（query）、K K K（key）和V V V（value）三个矩阵；接下来使用矩阵Q Q Q和K K K的乘积得到注意力矩阵A A A，归一化得到A ^ \hat{A}A ^；最后，将归一化后的注意力矩阵A ^ \hat{A}A ^乘上V V V，得到最后的输出特征O O O。

; 多头自注意力机制（multi-head self-attention）

上述的self-attention中，每个输入特征a i a^{i}a i乘上矩阵W q W^{q}W q、W k W^{k}W k和W v W^{v}W v后，分别得到一个向量q i q^{i}q i、k i k^{i}k i和v i v^{i}v i，称为单头自注意力机制。如果将这些向量q i q^{i}q i、k i k^{i}k i和v i v^{i}v i分裂为n n n个就得到n n n头自注意力机制了。公认多头自注意力机制的效果好于单头的，因为前者可以捕获更多维度的信息。示意图如下：

代码实现

设超参数num_attention_heads为自注意力机制的头数，如此，计算出每个头的维度attention_head_size。

self.num_attention_heads = num_attention_heads
self.attention_head_size = int(hidden_size / num_attention_heads)
self.all_head_size = hidden_size

定义W q W^{q}W q、W k W^{k}W k和W v W^{v}W v三个矩阵。

self.query = nn.Linear(input_size, self.all_head_size)
self.key = nn.Linear(input_size, self.all_head_size)
self.value = nn.Linear(input_size, self.all_head_size)

下面开始逐步计算，需要主要的是计算过程中张量维度的变化。
将输入特征乘以三个矩阵W q W^{q}W q、W k W^{k}W k和W v W^{v}W v，输出的张量此时还没有区分出多个头。维度变化为：input_tensor ( b a t c h , n , i n p u t _ s i z e ) \left ( batch,n,input_size\right )(b a t c h ,n ,i n p u t _s i z e )到mixed_query_layer ( b a t c h , n , a l l _ h e a d _ s i z e ) \left ( batch,n,all_head_size\right )(b a t c h ,n ,a l l _h e a d _s i z e )

mixed_query_layer = self.query(input_tensor)
mixed_key_layer = self.key(input_tensor)
mixed_value_layer = self.value(input_tensor)

切分为num_attention_heads个头，并变换维度。维度变化为：mixed_query_layer ( b a t c h , n , a l l _ h e a d _ s i z e ) \left ( batch,n,all_head_size\right )(b a t c h ,n ,a l l _h e a d _s i z e )到query_layer ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,num_attention_heads,n,attention_head_size\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,a t t e n t i o n _h e a d _s i z e )

def transpose_for_scores(self, x):
   new_x_shape = x.size()[:-1] + (self.num_attention_heads, self.attention_head_size)
   x = x.view(*new_x_shape)
   return x.permute(0, 2, 1, 3)

query_layer = self.transpose_for_scores(mixed_query_layer)
key_layer = self.transpose_for_scores(mixed_key_layer)
value_layer = self.transpose_for_scores(mixed_value_layer)

矩阵Q Q Q和K K K相乘，得到注意力矩阵，并除以向量的维度的开方，防止注意力分数随维度增大而增大。维度变化为：query_layer ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,num_attention_heads,n,attention_head_size\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,a t t e n t i o n _h e a d _s i z e )到attention_scores ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , n ) \left ( batch,num_attention_heads,n,n\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,n )

attention_scores = torch.matmul(query_layer, key_layer.transpose(-1, -2))

attention_scores = attention_scores / math.sqrt(self.attention_head_size)

注意力矩阵归一化。维度变化为：attention_scores ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , n ) \left ( batch,num_attention_heads,n,n\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,n )到attention_probs ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , n ) \left ( batch,num_attention_heads,n,n\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,n )

attention_probs = nn.Softmax(dim=-1)(attention_scores)

将注意力矩阵乘以矩阵V V V。维度变化为：ttention_probs ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , n ) \left ( batch,num_attention_heads,n,n\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,n )乘以value_layer ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,num_attention_heads,n,attention_head_size\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,a t t e n t i o n _h e a d _s i z e )到context_layer ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,num_attention_heads,n,attention_head_size\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,a t t e n t i o n _h e a d _s i z e )。

context_layer = torch.matmul(attention_probs, value_layer)

变换context_layer维度，为了后面将各头得到的结果拼接。这里的contiguous()是将tensor的内存变成连续的，为后面的view()做准备。维度变化为：context_layer ( b a t c h , n u m _ a t t e n t i o n _ h e a d s , n , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,num_attention_heads,n,attention_head_size\right )(b a t c h ,n u m _a t t e n t i o n _h e a d s ,n ,a t t e n t i o n _h e a d _s i z e )到context_layer ( b a t c h , n , n u m _ a t t e n t i o n _ h e a d s , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,n,num_attention_heads,attention_head_size\right )(b a t c h ,n ,n u m _a t t e n t i o n _h e a d s ,a t t e n t i o n _h e a d _s i z e )

context_layer = context_layer.permute(0, 2, 1, 3).contiguous()

将各头的结果拼接起来。维度变化为：context_layer ( b a t c h , n , n u m _ a t t e n t i o n _ h e a d s , a t t e n t i o n _ h e a d _ s i z e ) \left ( batch,n,num_attention_heads,attention_head_size\right )(b a t c h ,n ,n u m _a t t e n t i o n _h e a d s ,a t t e n t i o n _h e a d _s i z e )到context_layer ( b a t c h , n , a l l _ h e a d _ s i z e ) \left ( batch,n,all_head_size\right )(b a t c h ,n ,a l l _h e a d _s i z e )

new_context_layer_shape = context_layer.size()[:-2] + (self.all_head_size,)
context_layer = context_layer.view(*new_context_layer_shape)

完整代码

class LayerNorm(nn.Module):
    def __init__(self, hidden_size, eps=1e-12):
        """Construct a layernorm module in the TF style (epsilon inside the square root).

"""
        super(LayerNorm, self).__init__()
        self.weight = nn.Parameter(torch.ones(hidden_size))
        self.bias = nn.Parameter(torch.zeros(hidden_size))
        self.variance_epsilon = eps

    def forward(self, x):
        u = x.mean(-1, keepdim=True)
        s = (x - u).pow(2).mean(-1, keepdim=True)
        x = (x - u) / torch.sqrt(s + self.variance_epsilon)
        return self.weight * x + self.bias

class SelfAttention(nn.Module):
    def __init__(self, num_attention_heads, input_size, hidden_size, hidden_dropout_prob):
        super(SelfAttention, self).__init__()
        if hidden_size % num_attention_heads != 0:
            raise ValueError(
                "The hidden size (%d) is not a multiple of the number of attention "
                "heads (%d)" % (hidden_size, num_attention_heads))
        self.num_attention_heads = num_attention_heads
        self.attention_head_size = int(hidden_size / num_attention_heads)
        self.all_head_size = hidden_size

        self.query = nn.Linear(input_size, self.all_head_size)
        self.key = nn.Linear(input_size, self.all_head_size)
        self.value = nn.Linear(input_size, self.all_head_size)

        self.attn_dropout = nn.Dropout(attention_probs_dropout_prob)

        self.dense = nn.Linear(hidden_size, hidden_size)
        self.LayerNorm = LayerNorm(hidden_size, eps=1e-12)
        self.out_dropout = nn.Dropout(hidden_dropout_prob)

    def transpose_for_scores(self, x):
        new_x_shape = x.size()[:-1] + (self.num_attention_heads, self.attention_head_size)
        x = x.view(*new_x_shape)
        return x.permute(0, 2, 1, 3)

    def forward(self, input_tensor):
        mixed_query_layer = self.query(input_tensor)
        mixed_key_layer = self.key(input_tensor)
        mixed_value_layer = self.value(input_tensor)

        query_layer = self.transpose_for_scores(mixed_query_layer)
        key_layer = self.transpose_for_scores(mixed_key_layer)
        value_layer = self.transpose_for_scores(mixed_value_layer)

        attention_scores = torch.matmul(query_layer, key_layer.transpose(-1, -2))

        attention_scores = attention_scores / math.sqrt(self.attention_head_size)

        attention_probs = nn.Softmax(dim=-1)(attention_scores)

        attention_probs = self.attn_dropout(attention_probs)
        context_layer = torch.matmul(attention_probs, value_layer)
        context_layer = context_layer.permute(0, 2, 1, 3).contiguous()
        new_context_layer_shape = context_layer.size()[:-2] + (self.all_head_size,)
        context_layer = context_layer.view(*new_context_layer_shape)
        hidden_states = self.dense(context_layer)
        hidden_states = self.out_dropout(hidden_states)
        hidden_states = self.LayerNorm(hidden_states + input_tensor)

        return hidden_states

Original: https://blog.csdn.net/beilizhang/article/details/115282604
Author: cqu_shuai
Title: PyTorch——自注意力（self-attention）机制实现（代码详解）

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/627709/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

git branch 分支管理

在多人协作的情况下,master通常是稳定的分支.可以再建一些”develop”,”testing”等名称的分支.主管master的…

人工智能 2023年6月4日
0082
SPSS单因素方差分析教程

文章目录 * – 写在前面 – 什么是单因素方差分析 – 单因素方差分析的原理 – + 单因素方差分析的零假设 + 单因素方差分析的…

人工智能 2023年7月14日
0070
jQuery提供的获取元素位置的接口方法

HTML元素的位置相关的css属性有top、left、bottom、right。要灵活使用这些属性，需要了解css的定位模型position：正常文档流，相对定位，绝对定位。了解…

人工智能 2023年6月27日
0093
机器学习笔记 – 在逻辑回归中使用分类权重处理不平衡数据

逻辑回归是用于分类任务的监督机器学习技术之一。大多数情况下，分类数据集会出现类别不平衡，某个类别的样本较多，而某些类别的样本数量非常少。使用不平衡的数据集进行模型构建会导致错误的…

人工智能 2023年7月1日
00119
【推荐算法】协同过滤算法代码（pyspark | ALS）

【推荐算法】协同过滤算法介绍_MachineCYL的博客-CSDN博客上文介绍了协同过滤算法的原理，接下来我介绍一下协同过滤算法的代码实现。下面我就开始介绍用pyspark中的…

人工智能 2023年6月16日
0091
【数据攻略】字节面试真题（含答案）+100道面试题库

整理了一套字节的面试真题，还有100道PDF版的面试题库一、SQL题面试真题1：抖音电商平台，现有一张订单表（order_info），有以下字段： order_id good…

人工智能 2023年7月16日
0066
MATLAB自动驾驶工具箱使用

打开工具箱 MATLAB R2017a及以后的版本才有自动驾驶工具箱。在MATLAB的APPS中选择AUTOMOTIVE下面的Driving Scenario Designer …

人工智能 2023年6月2日
0093
【图像分类】实战——使用ResNet实现猫狗分类（pytorch）

目录摘要导入项目使用的库设置全局参数图像预处理读取数据设置模型设置训练和验证验证完整代码：摘要 ResNet（Residual Neural Network）由…

人工智能 2023年7月22日
0076
【数学建模之Python】9.AttributeError: module ‘pandas‘ has no attribute ‘Panel‘

你们的每个赞都能让我开心好几天✿✿ヽ(°▽°)ノ✿ 其实这类报错仔细看看就能够明白为什么，如果自己写的程序是没问题的话，那报错就是库的问题简而言之，库的版本太旧了，有的函数、名词已…

人工智能 2023年7月8日
0065
SSD项目代码解析(pytroch)（一）

一、前言 SSD（Single Shot MultiBox Detector)是一种单阶段实时目标检测模型，在问世之初，取得了非常好的性能和实时检测能力，一度是最受欢迎的目标检测架…

人工智能 2023年7月9日
0062
python 数据处理学习pandas之DataFrame

请原谅没有一次写完,本文是自己学习过程中的记录,完善pandas的学习知识,对于现有网上资料的缺少和利用python进行数据分析这本书部分知识的过时,只好以记录的形势来写这篇文章….

人工智能 2023年6月2日
0091
Kaggle 机器学习实战朴素贝叶斯（原理+西瓜数据集实战）

Kaggle 机器学习实战朴素贝叶斯（原理+西瓜数据集实战）朴素贝叶斯概念（这一部分来自于国科大网安学院的PPT以及周志华的机器学习，需要的可在文章末尾加公号 AC粥回复 …

人工智能 2023年7月28日
0059
python-生成数据

文章目录 1 绘制简单折线图 * 1.1 绘制简单的折线图 1.2 修改图表 1.3 校正图形 1.4 使用内置样式 2 绘制散点图 * 2.1 使用scatter()绘制散点图并…

人工智能 2023年7月16日
0069
盘点五大类 DeFi 数据分析工具

Feb. 2022，Grace 伴随着 DeFi 的繁荣，加密数据分析的市场也方兴未艾。已实现对一个 DeFi 项目的初步解析。笔者在使用诸多分析工具后，整理了比较好用的，且市面上…

人工智能 2023年6月11日
0079
向量叉乘的几何意义及其模的计算

目的：在传统的向量叉乘计算中，常常遇到叉乘。定义为向量。其这个向量方向满足右手定则。它的模大小，一般被忽略。因此推测一下。向量叉乘定义：外积（英语： Cross product）…

人工智能 2023年6月16日
00109
NLP-文本处理：指代消解（Coreference Resolution）【回指消解（名词＜–＞代词）、共指消解（名词1＜–＞名词2）】【识别指向同一实体的不同表述】【难度较大，准确率不会太高】

共指消解（coreference resolution）技术同NER、RE。作为自然语言历届基础技术被广泛的应用于：文本摘要、机器翻译、自动问答和知识图谱等领域。共指消解的提出是…

人工智能 2023年6月10日
0074

2024 年 5 月
一	二	三	四	五	六	日
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31