深度强化学习–Deep Q Network(using tensorflow)

2023年5月25日上午3:50 • 人工智能 • 阅读 88

文章目录

前言
一、什么是Q-learning
二、什么是Deep Q Network
三、DQN的两大利器
*
1.Experience replay
2.Fixed Q-targets
四、DQN算法
五、DQN的实现(using tensorflow)
*
1.run_this.py
2.RL_brain.py
1)DQN类的整体框架
2)建立两个神经网络(Q估计和Q现实)
3)初始值
4)存储记忆
5)选行为
6)学习
7)看学习效果

前言

总结记录深度强化学习算法Deep Q Network
参考自莫凡python

深度强化学习--Deep Q Network(using tensorflow)

; 一、什么是Q-learning

Q-Learning
α-学习效率 γ-衰减率 ε-贪婪度
r-环境所对应回报
s-状态
s_-下一状态
a-行为
Q(s,a)-Q表(某状态下采取某行为的价值)

ε:每次执行时，会有ε的概率选择Q表当中的最优项，1-ε的概率选择随机项(用于学习)。
γ:越高，效果越有远见，不会只看到近期的回报。

Q(s,a) += α[r + γQ(s_,a).max – Q(s,a)]

二、什么是Deep Q Network

之前传统的强化学习算法，比如： Q-learning, Sarsa等方法都是用表格的方式去存储每一个状态的state,和在这个state每个action所拥有的Q值。局限性在于如果 实际问题中的state过多，计算机内存无法满足存储要求。
于是 Deep Q Network将传统的 强化学习算法与 神经网络结合.我们可以将 状态和动作当成神经网络的输入, 然后经过 神经网络分析后得到动作的 Q 值, 这样我们就没必要在表格中记录 Q 值, 而是直接使用 神经网络生成 Q 值. 还有一种形式的是这样, 我们也能 只输入状态值, 输出所有的 动作值, 然后按照 Q learning 的原则, 直接选择拥有最大值的动作当做下一步要做的动作.

三、DQN的两大利器

; 1.Experience replay

Experience replay是指DQN可以 建立记忆库用于存储 之前学习过的一些经历，Q-learning是一种 off-policy的算法，可以学习过去的或者是别人的经历。所以每次 DQN 更新的时候, 我们都可以 随机抽取一些之前的经历进行学习. 随机抽取这种做法 打乱了经历之间的相关性, 也使得神经网络更新更有效率.

2.Fixed Q-targets

Fixed Q-targets 也是一种 打乱相关性的机理, 如果使用 fixed Q-targets, 我们就会在 DQN 中使用到两个 结构相同但参数不同的神经网络, 预测 Q 估计 的神经网络具备 最新的参数, 而预测 Q 现实 的神经网络使用的参数则是 很久以前的.

Q估计神经网络参数是随着学习的进行 不断更新的，而 Q现实的参数是被冻结的，我们可以设定参数，固定在X次更新之后，将 Q估计中的参数覆盖到 Q现实中的参数。

; 四、DQN算法

在Q-learning算法的基础上增加了以下特点：
记忆库 (用于重复学习)
神经网络计算 Q 值
暂时冻结 q_target 参数 (切断相关性)

五、DQN的实现(using tensorflow)

1.run_this.py

首先import所需模块

from maze_env import Maze
from RL_brain import DeepQNetwork

DQN与环境的交互

def run_maze():
    step = 0
    for episode in range(300):

        observation = env.reset()

        while True:

            env.render()

            action = RL.choose_action(observation)

            observation_, reward, done = env.step(action)

            RL.store_transition(observation, action, reward, observation_)

            if (step > 200) and (step % 5 == 0):
                RL.learn()

            observation = observation_

            if done:
                break
            step += 1

    print('game over')
    env.destroy()

if __name__ == "__main__":
    env = Maze()
    RL = DeepQNetwork(env.n_actions, env.n_features,
                      learning_rate=0.01,
                      reward_decay=0.9,
                      e_greedy=0.9,
                      replace_target_iter=200,
                      memory_size=2000,

                      )
    env.after(100, run_maze)
    env.mainloop()
    RL.plot_cost()

2.RL_brain.py

1)DQN类的整体框架

class DeepQNetwork:

    def _build_net(self):

    def __init__(self):

    def store_transition(self, s, a, r, s_):

    def choose_action(self, observation):

    def learn(self):

    def plot_cost(self):

2)建立两个神经网络(Q估计和Q现实)

class DeepQNetwork:
    def _build_net(self):

        self.s = tf.placeholder(tf.float32, [None, self.n_features], name='s')
        self.q_target = tf.placeholder(tf.float32, [None, self.n_actions], name='Q_target')
        with tf.variable_scope('eval_net'):

            c_names, n_l1, w_initializer, b_initializer = \
                ['eval_net_params', tf.GraphKeys.GLOBAL_VARIABLES], 10, \
                tf.random_normal_initializer(0., 0.3), tf.constant_initializer(0.1)

            with tf.variable_scope('l1'):
                w1 = tf.get_variable('w1', [self.n_features, n_l1], initializer=w_initializer, collections=c_names)
                b1 = tf.get_variable('b1', [1, n_l1], initializer=b_initializer, collections=c_names)
                l1 = tf.nn.relu(tf.matmul(self.s, w1) + b1)

            with tf.variable_scope('l2'):
                w2 = tf.get_variable('w2', [n_l1, self.n_actions], initializer=w_initializer, collections=c_names)
                b2 = tf.get_variable('b2', [1, self.n_actions], initializer=b_initializer, collections=c_names)
                self.q_eval = tf.matmul(l1, w2) + b2

        with tf.variable_scope('loss'):
            self.loss = tf.reduce_mean(tf.squared_difference(self.q_target, self.q_eval))
        with tf.variable_scope('train'):
            self._train_op = tf.train.RMSPropOptimizer(self.lr).minimize(self.loss)

        self.s_ = tf.placeholder(tf.float32, [None, self.n_features], name='s_')
        with tf.variable_scope('target_net'):

            c_names = ['target_net_params', tf.GraphKeys.GLOBAL_VARIABLES]

            with tf.variable_scope('l1'):
                w1 = tf.get_variable('w1', [self.n_features, n_l1], initializer=w_initializer, collections=c_names)
                b1 = tf.get_variable('b1', [1, n_l1], initializer=b_initializer, collections=c_names)
                l1 = tf.nn.relu(tf.matmul(self.s_, w1) + b1)

            with tf.variable_scope('l2'):
                w2 = tf.get_variable('w2', [n_l1, self.n_actions], initializer=w_initializer, collections=c_names)
                b2 = tf.get_variable('b2', [1, self.n_actions], initializer=b_initializer, collections=c_names)
                self.q_next = tf.matmul(l1, w2) + b2

tensorboard神经网络可视化

3)初始值

class DeepQNetwork:
    def __init__(
            self,
            n_actions,
            n_features,
            learning_rate=0.01,
            reward_decay=0.9,
            e_greedy=0.9,
            replace_target_iter=300,
            memory_size=500,
            batch_size=32,
            e_greedy_increment=None,
            output_graph=False,
    ):
        self.n_actions = n_actions
        self.n_features = n_features
        self.lr = learning_rate
        self.gamma = reward_decay
        self.epsilon_max = e_greedy
        self.replace_target_iter = replace_target_iter
        self.memory_size = memory_size
        self.batch_size = batch_size
        self.epsilon_increment = e_greedy_increment
        self.epsilon = 0 if e_greedy_increment is not None else self.epsilon_max

        self.learn_step_counter = 0

        self.memory = np.zeros((self.memory_size, n_features*2+2))

        self._build_net()

        t_params = tf.get_collection('target_net_params')
        e_params = tf.get_collection('eval_net_params')
        self.replace_target_op = [tf.assign(t, e) for t, e in zip(t_params, e_params)]

        self.sess = tf.Session()

        if output_graph:

            tf.summary.FileWriter("logs/", self.sess.graph)

        self.sess.run(tf.global_variables_initializer())
        self.cost_his = []

4)存储记忆

class DeepQNetwork:
    def __init__(self):
        ...

    def store_transition(self, s, a, r, s_):
        if not hasattr(self, 'memory_counter'):
            self.memory_counter = 0

        transition = np.hstack((s, [a, r], s_))

        index = self.memory_counter % self.memory_size
        self.memory[index, :] = transition

        self.memory_counter += 1

5)选行为

class DeepQNetwork:
    def __init__(self):
        ...

    def store_transition(self, s, a, r, s_):
        ...

    def choose_action(self, observation):

        observation = observation[np.newaxis, :]

        if np.random.uniform() < self.epsilon:

            actions_value = self.sess.run(self.q_eval, feed_dict={self.s: observation})
            action = np.argmax(actions_value)
        else:
            action = np.random.randint(0, self.n_actions)
        return action

6)学习

class DeepQNetwork:
    def __init__(self):
        ...

    def store_transition(self, s, a, r, s_):
        ...

    def choose_action(self, observation):
        ...

    def _replace_target_params(self):
        ...

    def learn(self):

        if self.learn_step_counter % self.replace_target_iter == 0:
            self.sess.run(self.replace_target_op)
            print('\ntarget_params_replaced\n')

        if self.memory_counter > self.memory_size:
            sample_index = np.random.choice(self.memory_size, size=self.batch_size)
        else:
            sample_index = np.random.choice(self.memory_counter, size=self.batch_size)
        batch_memory = self.memory[sample_index, :]

        q_next, q_eval = self.sess.run(
            [self.q_next, self.q_eval],
            feed_dict={
                self.s_: batch_memory[:, -self.n_features:],
                self.s: batch_memory[:, :self.n_features]
            })

        q_target = q_eval.copy()
        batch_index = np.arange(self.batch_size, dtype=np.int32)
        eval_act_index = batch_memory[:, self.n_features].astype(int)
        reward = batch_memory[:, self.n_features + 1]

        q_target[batch_index, eval_act_index] = reward + self.gamma * np.max(q_next, axis=1)

"""
        假如在这个 batch 中, 我们有2个提取的记忆, 根据每个记忆可以生产3个 action 的值:
        q_eval =
        [[1, 2, 3],
         [4, 5, 6]]

        q_target = q_eval =
        [[1, 2, 3],
         [4, 5, 6]]

        然后根据 memory 当中的具体 action 位置来修改 q_target 对应 action 上的值:
        比如在:
            记忆 0 的 q_target 计算值是 -1, 而且我用了 action 0;
            记忆 1 的 q_target 计算值是 -2, 而且我用了 action 2:
        q_target =
        [[-1, 2, 3],
         [4, 5, -2]]

        所以 (q_target - q_eval) 就变成了:
        [[(-1)-(1), 0, 0],
         [0, 0, (-2)-(6)]]

        最后我们将这个 (q_target - q_eval) 当成误差, 反向传递会神经网络.

        所有为 0 的 action 值是当时没有选择的 action, 之前有选择的 action 才有不为0的值.

        我们只反向传递之前选择的 action 的值,
"""

        _, self.cost = self.sess.run([self._train_op, self.loss],
                                     feed_dict={self.s: batch_memory[:, :self.n_features],
                                                self.q_target: q_target})
        self.cost_his.append(self.cost)

        self.epsilon = self.epsilon + self.epsilon_increment if self.epsilon < self.epsilon_max else self.epsilon_max
        self.learn_step_counter += 1

7)看学习效果

class DeepQNetwork:
    def __init__(self):
        ...

    def store_transition(self, s, a, r, s_):
        ...

    def choose_action(self, observation):
        ...

    def _replace_target_params(self):
        ...

    def learn(self):
        ...

    def plot_cost(self):
        import matplotlib.pyplot as plt
        plt.plot(np.arange(len(self.cost_his)), self.cost_his)
        plt.ylabel('Cost')
        plt.xlabel('training steps')
        plt.show()

Original: https://blog.csdn.net/qq_52654678/article/details/121895814
Author: 旺仔不涨价
Title: 深度强化学习–Deep Q Network(using tensorflow)

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/511984/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

百度校招社招-知识图谱部门直推机会多多

啊哦~你想找的内容离你而去了哦内容不存在，可能为如下原因导致： ① 内容还在审核中 ② 内容以前存在，但是由于不符合新的规定而被删除 ③ 内容地址错误 ④ 作者删除了内容。可…

人工智能 2023年6月1日
0064
Python回归预测建模实战-支持向量机预测房价（附源码和实现效果）

机器学习在预测方面的应用，根据预测值变量的类型可以分为分类问题（预测值是离散型）和回归问题（预测值是连续型），前面我们介绍了机器学习建模处理了分类问题（具体见之前的文章），接下…

人工智能 2023年6月17日
0099
解决Anaconda3 solving environment 巨慢的方法

解决Anaconda3 solving environment 巨慢的方法，亲测有效！！！最近在做毕设辽，准备做一个基于深度学习的MOT项目，python开发，coding期间由…

人工智能 2023年6月16日
0081
如何在docker上面安装常用的容器环境

本文主要记录了如何在docker上面安装最长的容器环境，包括redis、mongodb、es以及mysql 安装redis 1. 从docker hub上(阿里云加速器)拉取red…

人工智能 2023年6月4日
0099
R语言逻辑运算符（Logical Operators，大于、小于、等于、不等于、与或非、是否为真）、R语言逻辑运算符（Logical Operators）实战示例

抵扣说明： 1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。2.余额无法直接购买下载，可以购买VIP、C币套餐、付费专栏及课程。 Original: https:…

人工智能 2023年6月19日
00105
Yolov5更换backbone，与模型压缩（剪枝，量化，蒸馏）

~~~欢迎各位交流、star、fork、issues~~~ 项目介绍：本仓库是基于官方yolov5源码的基础上，进行的改进。目前支持更换yolov5的backbone主干网络为…

人工智能 2023年6月23日
0071
Metashape（Photoscan）【制作DOM和DEM】超级详细的步骤，文末有安装包

Metashape（Photoscan）【制作DOM和DEM】超级详细的步骤 1. Metashape软件操作简介 * 1.1.Metashape页面简介 1.2.Metashap…

人工智能 2023年6月17日
0073
数据挖掘原理与实践第四章作业

P147 4.2 假设数据挖掘的任务是将如下的8个点（用 (x,y) 代表位置）聚类为三个簇：A1 (2,10)，A2(2,5)，A3(8,4)，B1(5,8)，B2(7,5)，B…

人工智能 2023年7月18日
0061
PX4装机教程（八）常用外接传感器

抵扣说明： 1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。2.余额无法直接购买下载，可以购买VIP、C币套餐、付费专栏及课程。 Original: https:…

人工智能 2023年6月25日
0096
256qam是什么意思_基带、射频，到底是干什么用的？

现在都流行”端到端”，我们就以手机通话为例，观察信号从手机到基站的整个过程，来看看基带和射频到底是干什么用的。当手机通话接通时，人的声音会通过手机麦克风拾…

人工智能 2023年5月27日
0084
机器知道哪吒是部电影吗？解读阿里巴巴概念图谱AliCG

概念是人类认知世界的基石。比如对于”哪吒好看吗？”，”哪吒铭文搭配建议”两句话，人可以结合概念知识理解第一个哪吒是一部电影，第二个哪…

人工智能 2023年6月10日
00101
条件DDPM：Diffusion model的第三个巅峰之作

抵扣说明： 1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。2.余额无法直接购买下载，可以购买VIP、C币套餐、付费专栏及课程。 Original: https:…

人工智能 2023年6月17日
0073
【youcans 的 OpenCV 例程200篇】172.SLIC 超像素区域分割算法比较

OpenCV 例程200篇总目录-202205更新【youcans 的 OpenCV 例程200篇】172.SLIC 超像素区域分割算法比较 5. 区域分割之聚类方法 5.3 …

人工智能 2023年6月3日
0082
conda安装指定版本TensorFlow

文章目录 * – 一、系统环境 – 二、安装步骤一、系统环境操作系统：Windows7 64位，Python环境：Python3.7；conda 4.1…

人工智能 2023年5月23日
0076
ISPRS2022/遥感影像云检测：Cloud detection with boundary nets基于边界网的云检测

ISPRS2022/云检测：Cloud detection with boundary nets基于边界网的云检测 0.摘要 1.概述 * 1.1. 云检测方法综述 1.2. 挑战…

人工智能 2023年6月20日
00137
Topic 11. SCI中多元变量筛选—单/多因素表

单因素和多因素这里有时也会很困惑，在分析中也同样会遇到很多问题，比如当做单因素分析时，得到的 P 值显著，但是在做多因素时却不显著？又比如多因素分析时，选择变量的个数不同，得到的 …

人工智能 2023年6月17日
00110

2024 年 5 月
一	二	三	四	五	六	日
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31