YOLOv5-6.1添加注意力机制（SE、CBAM、ECA、CA）

2023年7月27日上午3:21 • 人工智能 • 阅读 161

主要步骤：
（1）在 models/common.py中注册注意力模块
（2）在 models/yolo.py中的 parse_model函数中添加注意力模块
（3）修改配置文件 yolov5s.yaml
（4）运行 yolo.py进行验证
各个注意力机制模块的添加方法类似，各注意力模块的修改参照SE。
本文添加注意力完整代码：https://github.com/double-vin/yolov5_attention

Squeeze-and-Excitation Networks
https://github.com/hujie-frank/SENet

; 1.1 SE

在 models/common.py中注册SE模块

class SE(nn.Module):
    def __init__(self, c1, c2, ratio=16):
        super(SE, self).__init__()

        self.avgpool = nn.AdaptiveAvgPool2d(1)
        self.l1 = nn.Linear(c1, c1 // ratio, bias=False)
        self.relu = nn.ReLU(inplace=True)
        self.l2 = nn.Linear(c1 // ratio, c1, bias=False)
        self.sig = nn.Sigmoid()

    def forward(self, x):
        b, c, _, _ = x.size()
        y = self.avgpool(x).view(b, c)
        y = self.l1(y)
        y = self.relu(y)
        y = self.l2(y)
        y = self.sig(y)
        y = y.view(b, c, 1, 1)
        return x * y.expand_as(x)

在 models/yolo.py中的 parse_model函数中添加SE模块
修改配置文件 yolov5s.yaml。
添加注意力的两种方法：一是在backbone的最后一层添加注意力；二是将backbone中的C3全部替换。
这里使用第一种，第二种见下文中的 C3SE

注意：SE添加至第9层，第9层之后所有的编号都要+1，则：
1> 两个Concat的from系数分别由[-1, 14]，[-1, 10]改为[-1, 15]，[-1, 11]
2> Detect的from系数由[17, 20, 23]改为[18,21,24]
验证：运行 yolo.py

1.2 C3-SE

在 models/common.py中注册C3SE模块：

class SEBottleneck(nn.Module):

    def __init__(self, c1, c2, shortcut=True, g=1, e=0.5, ratio=16):
        super().__init__()
        c_ = int(c2 * e)
        self.cv1 = Conv(c1, c_, 1, 1)
        self.cv2 = Conv(c_, c2, 3, 1, g=g)
        self.add = shortcut and c1 == c2

        self.avgpool = nn.AdaptiveAvgPool2d(1)
        self.l1 = nn.Linear(c1, c1 // ratio, bias=False)
        self.relu = nn.ReLU(inplace=True)
        self.l2 = nn.Linear(c1 // ratio, c1, bias=False)
        self.sig = nn.Sigmoid()

    def forward(self, x):
        x1 = self.cv2(self.cv1(x))
        b, c, _, _ = x.size()
        y = self.avgpool(x1).view(b, c)
        y = self.l1(y)
        y = self.relu(y)
        y = self.l2(y)
        y = self.sig(y)
        y = y.view(b, c, 1, 1)
        out = x1 * y.expand_as(x1)

        return x + out if self.add else out

class C3SE(C3):

    def __init__(self, c1, c2, n=1, shortcut=True, g=1, e=0.5):
        super().__init__(c1, c2, n, shortcut, g, e)
        c_ = int(c2 * e)
        self.m = nn.Sequential(*(SEBottleneck(c_, c_, shortcut) for _ in range(n)))

在 models/yolo.py中的 parse_model函数中添加C3SE模块
修改配置文件 yolov5s.yaml。
验证：运行 yolo.py
CBAM

《CBAM: Convolutional Block Attention Module》

; 2.1 CBAM

class ChannelAttention(nn.Module):
    def __init__(self, in_planes, ratio=16):
        super(ChannelAttention, self).__init__()
        self.avg_pool = nn.AdaptiveAvgPool2d(1)
        self.max_pool = nn.AdaptiveMaxPool2d(1)
        self.f1 = nn.Conv2d(in_planes, in_planes // ratio, 1, bias=False)
        self.relu = nn.ReLU()
        self.f2 = nn.Conv2d(in_planes // ratio, in_planes, 1, bias=False)
        self.sigmoid = nn.Sigmoid()

    def forward(self, x):
        avg_out = self.f2(self.relu(self.f1(self.avg_pool(x))))
        max_out = self.f2(self.relu(self.f1(self.max_pool(x))))
        out = self.sigmoid(avg_out + max_out)
        return out

class SpatialAttention(nn.Module):
    def __init__(self, kernel_size=7):
        super(SpatialAttention, self).__init__()
        assert kernel_size in (3, 7), 'kernel size must be 3 or 7'
        padding = 3 if kernel_size == 7 else 1

        self.conv = nn.Conv2d(2, 1, kernel_size, padding=padding, bias=False)
        self.sigmoid = nn.Sigmoid()

    def forward(self, x):

        avg_out = torch.mean(x, dim=1, keepdim=True)
        max_out, _ = torch.max(x, dim=1, keepdim=True)
        x = torch.cat([avg_out, max_out], dim=1)

        x = self.conv(x)

        return self.sigmoid(x)

class CBAM(nn.Module):

    def __init__(self, c1, c2, ratio=16, kernel_size=7):
        super(CBAM, self).__init__()
        self.channel_attention = ChannelAttention(c1, ratio)
        self.spatial_attention = SpatialAttention(kernel_size)

    def forward(self, x):
        out = self.channel_attention(x) * x

        out = self.spatial_attention(out) * out
        return out

2.2 C3-CBAM

class CBAMBottleneck(nn.Module):

    def __init__(self, c1, c2, shortcut=True, g=1, e=0.5,ratio=16,kernel_size=7):
        super(CBAMBottleneck,self).__init__()
        c_ = int(c2 * e)
        self.cv1 = Conv(c1, c_, 1, 1)
        self.cv2 = Conv(c_, c2, 3, 1, g=g)
        self.add = shortcut and c1 == c2
        self.channel_attention = ChannelAttention(c2, ratio)
        self.spatial_attention = SpatialAttention(kernel_size)

    def forward(self, x):
        x1 = self.cv2(self.cv1(x))
        out = self.channel_attention(x1) * x1

        out = self.spatial_attention(out) * out
        return x + out if self.add else out

class C3CBAM(C3):

    def __init__(self, c1, c2, n=1, shortcut=True, g=1, e=0.5):
        super().__init__(c1, c2, n, shortcut, g, e)
        c_ = int(c2 * e)
        self.m = nn.Sequential(*(CBAMBottleneck(c_, c_, shortcut) for _ in range(n)))

《ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks》
https://github.com/BangguWu/ECANet

; 3.1 ECA

class ECA(nn.Module):
    """Constructs a ECA module.

    Args:
        channel: Number of channels of the input feature map
        k_size: Adaptive selection of kernel size
"""

    def __init__(self, c1, c2, k_size=3):
        super(ECA, self).__init__()
        self.avg_pool = nn.AdaptiveAvgPool2d(1)
        self.conv = nn.Conv1d(1, 1, kernel_size=k_size, padding=(k_size - 1) // 2, bias=False)
        self.sigmoid = nn.Sigmoid()

    def forward(self, x):

        y = self.avg_pool(x)

        y = self.conv(y.squeeze(-1).transpose(-1, -2)).transpose(-1, -2).unsqueeze(-1)

        y = self.sigmoid(y)

        return x * y.expand_as(x)

3.2 C3-ECA

class ECABottleneck(nn.Module):

    def __init__(self, c1, c2, shortcut=True, g=1, e=0.5, ratio=16, k_size=3):
        super().__init__()
        c_ = int(c2 * e)
        self.cv1 = Conv(c1, c_, 1, 1)
        self.cv2 = Conv(c_, c2, 3, 1, g=g)
        self.add = shortcut and c1 == c2

        self.avg_pool = nn.AdaptiveAvgPool2d(1)
        self.conv = nn.Conv1d(1, 1, kernel_size=k_size, padding=(k_size - 1) // 2, bias=False)
        self.sigmoid = nn.Sigmoid()

    def forward(self, x):
        x1 = self.cv2(self.cv1(x))

        y = self.avg_pool(x1)
        y = self.conv(y.squeeze(-1).transpose(-1, -2)).transpose(-1, -2).unsqueeze(-1)
        y = self.sigmoid(y)
        out = x1 * y.expand_as(x1)

        return x + out if self.add else out

class C3ECA(C3):

    def __init__(self, c1, c2, n=1, shortcut=True, g=1, e=0.5):
        super().__init__(c1, c2, n, shortcut, g, e)
        c_ = int(c2 * e)
        self.m = nn.Sequential(*(ECABottleneck(c_, c_, shortcut) for _ in range(n)))

Coordinate Attention for Efficient Mobile Network Design
https://github.com/Andrew-Qibin/CoordAttention

; 4.1 CA

class h_sigmoid(nn.Module):
    def __init__(self, inplace=True):
        super(h_sigmoid, self).__init__()
        self.relu = nn.ReLU6(inplace=inplace)

    def forward(self, x):
        return self.relu(x + 3) / 6

class h_swish(nn.Module):
    def __init__(self, inplace=True):
        super(h_swish, self).__init__()
        self.sigmoid = h_sigmoid(inplace=inplace)

    def forward(self, x):
        return x * self.sigmoid(x)

class CoordAtt(nn.Module):
    def __init__(self, inp, oup, reduction=32):
        super(CoordAtt, self).__init__()
        self.pool_h = nn.AdaptiveAvgPool2d((None, 1))
        self.pool_w = nn.AdaptiveAvgPool2d((1, None))
        mip = max(8, inp // reduction)
        self.conv1 = nn.Conv2d(inp, mip, kernel_size=1, stride=1, padding=0)
        self.bn1 = nn.BatchNorm2d(mip)
        self.act = h_swish()
        self.conv_h = nn.Conv2d(mip, oup, kernel_size=1, stride=1, padding=0)
        self.conv_w = nn.Conv2d(mip, oup, kernel_size=1, stride=1, padding=0)

    def forward(self, x):
        identity = x
        n, c, h, w = x.size()

        x_h = self.pool_h(x)

        x_w = self.pool_w(x).permute(0, 1, 3, 2)
        y = torch.cat([x_h, x_w], dim=2)

        y = self.conv1(y)
        y = self.bn1(y)
        y = self.act(y)
        x_h, x_w = torch.split(y, [h, w], dim=2)
        x_w = x_w.permute(0, 1, 3, 2)
        a_h = self.conv_h(x_h).sigmoid()
        a_w = self.conv_w(x_w).sigmoid()
        out = identity * a_w * a_h
        return out

4.2 C3-CA

class CABottleneck(nn.Module):

    def __init__(self, c1, c2, shortcut=True, g=1, e=0.5, ratio=32):
        super().__init__()
        c_ = int(c2 * e)
        self.cv1 = Conv(c1, c_, 1, 1)
        self.cv2 = Conv(c_, c2, 3, 1, g=g)
        self.add = shortcut and c1 == c2

        self.pool_h = nn.AdaptiveAvgPool2d((None, 1))
        self.pool_w = nn.AdaptiveAvgPool2d((1, None))
        mip = max(8, c1 // ratio)
        self.conv1 = nn.Conv2d(c1, mip, kernel_size=1, stride=1, padding=0)
        self.bn1 = nn.BatchNorm2d(mip)
        self.act = h_swish()
        self.conv_h = nn.Conv2d(mip, c2, kernel_size=1, stride=1, padding=0)
        self.conv_w = nn.Conv2d(mip, c2, kernel_size=1, stride=1, padding=0)

    def forward(self, x):
        x1=self.cv2(self.cv1(x))
        n, c, h, w = x.size()

        x_h = self.pool_h(x1)

        x_w = self.pool_w(x1).permute(0, 1, 3, 2)
        y = torch.cat([x_h, x_w], dim=2)

        y = self.conv1(y)
        y = self.bn1(y)
        y = self.act(y)
        x_h, x_w = torch.split(y, [h, w], dim=2)
        x_w = x_w.permute(0, 1, 3, 2)
        a_h = self.conv_h(x_h).sigmoid()
        a_w = self.conv_w(x_w).sigmoid()
        out = x1 * a_w * a_h

        return x + out if self.add else out

class C3CA(C3):

    def __init__(self, c1, c2, n=1, shortcut=True, g=1, e=0.5):
        super().__init__(c1, c2, n, shortcut, g, e)
        c_ = int(c2 * e)
        self.m = nn.Sequential(*(CABottleneck(c_, c_,shortcut) for _ in range(n)))

Tips：添加注意力的位置不局限，可以尝试各种排列组合
参考：
多种注意力介绍
添加注意力视频讲解
添加CBAM

Original: https://blog.csdn.net/weixin_50008473/article/details/124590939
Author: June vinvin
Title: YOLOv5-6.1添加注意力机制（SE、CBAM、ECA、CA）

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/717758/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

【实战】——以波士顿房价为例进行数据的相关分析和回归分析

目录前言一、相关分析 * 1、概念 2、数据来源及处理 3、分析 – 3.1、协方差 3.2、相关系数二、回归分析 * 1、概念 2、一元线性回归 3、多元回归 …

人工智能 2023年6月19日
0087
python summary结果提取_使用summaryou时，将回归结果导出为csv文件

可以将括号中的标准错误更改为t-statistics，但前提是要修改statsmodel库中的文件summary2.py。在您只需将该文件中的函数_col_params()替换为…

人工智能 2023年6月18日
0078
python&tensorflow2.0各种数组详解及相互转化

1 数组详解 2 转化详解在数据预处理中，经常需要各种数据结构相互转化，元组、列表、numpy.array、字典、张量 tensor、dataframe。本文中代码较多，文字较…

人工智能 2023年5月26日
0092
Ubuntu 20.04 编译ORB_SLAM2源码（普通模式） + 点云地图构建 + 增加颜色信息

主要记录一下自己跑得时候遇见的问题的整合，自己搭建环境的时候差不多浏览了不下一百多个网站，整和一下资源提高大家的效率。参考链接里面的大佬都写的非常详细，可以看看他们的文章 *Wri…

人工智能 2023年7月18日
0073
Modeling Conversation Structure and Temporal Dynamics for Jointly Predicting Rumor Stance and Veracity（ACL-19）

记录一下，论文建模对话结构和时序动态来联合预测谣言立场和真实性及其代码复现。 1 引言之前的研究发现，公众对谣言消息的立场是识别流行的谣言的关键信号，这也能表明它们的真实性。因此…

人工智能 2023年6月4日
0086
【深度学习】Retina Net 计算机视觉目标检测 Focal Loss

论文： https://arxiv.org/abs/1708.02002 文章目录 Retina Net Focal Loss Retina Net损失函数代码 Retina …

人工智能 2023年7月12日
0086
3. 微服务之nacos服务注册发现

3.1 nacos服务搭建 nacos 相对于 eureka 来说功能更加强大，在搭建服务中心的时候也不同于eureka引入模块依赖运行就可以，需要独立进行安装，启动后如图：默认…

人工智能 2023年6月29日
0074
stm32的语音识别_基于STM32实现孤立词语音识别系统

当接触或点击屏幕时，触摸控制器可读取触摸点位置，如此可通过屏幕直接接受用户的操作。相比较机械式按钮，触摸屏在操作上更加直观生动。综合考虑，本设计中采用2.5寸240×320分辨率的…

人工智能 2023年5月27日
0068
Yolov5：强大到你难以想象──新冠疫情下的口罩检测

初识 Yolov5是看到一个视频可以检测街道上所有的行人，并实时框选出来。之后学习了CNN卷积神经网络，在完成一个项目需求时，发现卷积神经网络在切割图像方面仍然不太好用。于是我想到…

人工智能 2023年6月19日
00103
整理了几个100%提高Python代码质量的技巧，直呼过瘾

B站|公众号：啥都会一点的研究生 ; 相关阅读整理了几个100%会踩的Python细节坑，提前防止脑血栓整理了十个100%提高效率的Python编程技巧，更上一层楼Python-…

人工智能 2023年7月5日
0060
一道经典的Python数据分析笔试题

最近无意看到一份关于数据分析的Python笔试题，做起来还是很有意思的，特意自己动手做了一下，和大家分享一下，希望大家也可以跟着练习。题目如下：首先，模拟数据： importp…

人工智能 2023年7月17日
0063
小波变换进行图像变换Matlab实现

小波变换是傅里叶变换的发展和扩充，在一定程度上克服了傅里叶变换的弱点与局限性。小波分析与Fourier变换相比，小波变换是空间域和频率域的局部变换，因而能有效地从信号中提取信息。 …

人工智能 2023年6月18日
0085
C语言学习之路（基础篇）—— 函数

说明：该篇博客是博主一字一码编写的，实属不易，请尊重原创，谢谢大家！概述 1) 函数是什么函数就是一段封装好的，可以重复使用的代码，它使得我们的程序更加模块化，不需要编写大量重…

人工智能 2023年6月26日
00116
窗口函数深度探索（一）：底层原理

前言在日常SQL数据分析中，经常会遇到需要在每组内排名，面对这类需求就需要使用sql的高级功能窗口函数了。一言以蔽之：在进行分组聚合以后，我们还想操作集合之前的数据就需要用…

人工智能 2023年7月17日
0073
论文阅读：Efficient Estimation of Word Representations in Vector Space

目录前言 Abstract 1.Introduction * 1.1 Goals of the Paper 1.2 Previous Work 2. Model Architec…

人工智能 2023年5月30日
0077
Android Studio实现一个简单的健身系统

文章目录一、系统背景二、系统概述三、开发环境四、系统结构五、详细设计 * 5.1、RecycleView 5.2、ViewPager 5.3、OkHttp 六、运行演示 …

人工智能 2023年5月30日
0075

2024 年 5 月
一	二	三	四	五	六	日
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31